A Framing Example

We began the course with a concrete example. A professor gives their (undergrad) research assistant data files with students' homework scores and usage logs of an AI tutor. The professor asks the assistant to produce a program that analyzes the data and determines whether students who used the AI tutor did better on the homework. The assistant gives an AI agent the tables and the professor's request. The AI-generated program reports that students who used the tutor performed 10% better than those who didn't. The assistant copied the program's output back to the professor. (Hey, ain't research easy?!)

The professor, preparing to protect their grant funding for the project, asks the assistant to explain the program and why they trust the result. The assistant honestly reports that they can't.

In class: We ran this process live, then ran a slightly different program that Kathi (the real 0111E instructor) had produced with this same process the day before. We indeed saw that the output looked plausible and that the agent was capable of this task. We also had a volunteer "assistant" report out loud that they had no idea how the program worked.

Evaluation

We assume that you will be uneasy about this interaction, especially had you been the assistant who either had to fake it or admit having not been involved in the analysis. Practically speaking, the professor could have done this AI interaction themselves, saving them the time and funds spent to hire the pass-through assistant. So part of the question here becomes what makes an assistant worth hiring?

Let's dig in: what actually went wrong in this scenario?

Stop and think

Come up with at least two concrete ideas for what the assistant should or should not have done in this situation.

We asked for two ideas so you couldn't stop with "the assistant used AI". Truth is, AI agents are powerful and useful tools, when used in appropriate ways. The point of this question is to articulate boundaries around those useful ways. Plausible ideas include:

  1. Accepting what the generated program produced on the full dataset without running the program on a small sample dataset for which we knew what the answer should be
  2. Not telling the agent what is meant by the term "does better" from the professor's initial email
  3. Not reading the program code line for line to confirm that it does what it's supposed to do
  4. Sharing the raw data with an agent: student work is supposed to be kept private and protected
  5. Trusting that all of the data were valid and should be included in the analysis

In class: Indeed, when we opened the files to look, we discovered missing entries, an ignored column regarding whether students had consented to participation, students who worked with the tutor after submitting the assignment, and tutor logs from student IDs that weren't in the class roster, among others.

Which of these are important in practice? How can we think about them systematically?

The capabilities of AI agents

By now, most of us have heard terms such as "AI slop", referring to the low-quality work that agents often produce. We hear seemingly contradictory statements such as "AI is making programmers vastly more productive" and "AI can't be trusted". Truth is, professional software developers who use AI tools effectively do so in particular ways. They state boundaries on what agents should do or how they do it, at the level of specific problems. These boundaries limit the programs that agents can construct, and give the human concrete details against which to evaluate the work.

In the context of 0111E, we will teach you to work with four such boundaries (what software engineers call "specifications"):

Data concepts: What information is important to the problem, what shape does it have, and where does it come from?

Plans: Given a problem to solve by writing a program, what high-level steps should be part of the solution or program and how do they interact?

Constraints: What assumptions do we need to make about the data or the analysis process? How will we check or enforce them?

Testing Plans: What scenarios should we test our program on before declaring it ready to use on larger, "real", data?

Our goal is to teach you three broad sets of skills:

Stop and think

Go back to the proposed list of "what should be done differently", where does each seem to fit into these specifications and learning objectives?

Throughout the course, you should expect to see us labeling activities as "concept", "plan", "constraints", "test", and "audit". You will learn how to engage in each of these activities both manually by humans and collaboratively with AI tools. If you feel lost about the point of an activity, come back to this list and try to identify which kind of specification and which task we're working on, and on the roles of humans (you) and AI on each one. If you're still not sure, just ask!