Example workshop · For software engineers
Does your terminal config work?
Test the instructions, tools and context your coding agent uses in the terminal or editor. In two hours, you’ll run a real task, change one part of your setup and check what difference it makes.
- 2 hours
- For teams already using coding agents
- Your repository or a prepared sample
I’ve run this work with engineering teams at a major Baltic energy company. The workshop comes with runnable examples and a participant guide.
What you take away ↗Give one config change a fair test.
An AGENTS.md file can sound convincing. You still need to see what the agent does with it.
We work with the setup you already have: repository instructions, skills, MCP tools or supporting documentation. You choose one behaviour that matters and write a small check for it.
Did the agent request the tests? Did a tool fail? We inspect the recorded calls, then run the repository’s tests and review the code separately. Asking for a test doesn’t mean it passed.
The two-hour session
Bring a task you recognise.
00–24 minutes
Choose the task and connect tracing
We start with the team’s hopes and concerns, choose a small task and connect Copilot to Phoenix.
24–65 minutes
Run it and inspect the evidence
Find the tool calls. Try the supplied checks for test-command requests and recorded tool errors.
65–87 minutes
Write your own check
Adapt a small TypeScript evaluator, question it with a partner and run the repository checks. There’s an optional LLM-judge demonstration.
87–111 minutes
Change one thing and repeat
Keep the task, starting code, model and checks the same. Change one instruction, open a fresh session and compare both runs.
111–120 minutes
Decide what to keep
Keep, revise or revert the change. Choose what you’ll test next week and when you’ll share the result.
A check you can run again.
You leave with your evaluator, the results from both runs, your configuration diff and a decision about the change. One comparison gives you a starting point for testing other tasks.
The workshop kit
The participant guide includes a worksheet and the commands used in the session. You also have the example evaluators and a sample repository to repeat the exercise.

I’m Phil.
I’ll work through the experiment with your team. If the change makes things worse, we’ll look at why.
Before we book.
What do we need before the session?
For this version: Copilot in VS Code or Copilot CLI, Docker, and Node 24.11 or newer within Node 24. Bring a configured repository and a task the agent can tackle in 5–10 minutes. I have a sample if you need one. We’ll check the setup beforehand.
We use another coding agent. Can we do this?
Tell me which one. The existing exercises use Copilot and Phoenix; we’ll agree any changes to the tools and exercises when we plan your session.
Can we use our own code?
Yes, where your organisation allows it. Traces can include prompts, source code and tool output. Use approved material and keep credentials, personal data and production actions out of the exercise. The prepared sample is also available.
How much does it cost?
I’ll quote after we’ve agreed the scope, team size, preparation and delivery format. This two-hour session is an example we can adapt around your team.
What’s in your team’s config?
Tell me which agents you use and what you’d like them to do better. We’ll plan a session around that.
Plan a workshop