Now — What I'm Working On
A living page, not a static one. This is where I keep the honest, up-to-date record of what I'm building, learning, reading, and betting on right now — updated as the work actually moves.
last updated: August 2026
What I'm building, learning, reading & betting on.
LLM & Agentic AI Engineer at Property Turkey
I build agent systems end to end — orchestrating language models, wiring tools, and composing memory into agents that actually get work done.
- Agent orchestration — routing, planning, and multi-step task decomposition.
- Tool use — giving agents real handles on the world, with verification.
- Memory — persistence and retrieval across turns and long-running goals.
- Human-in-the-loop feedback loops — the loop closes when a human can steer.
Direction: exploring vision-based agents that act in real environments — with a verifier that checks the work actually happened.
AI Sensile
The AI operating layer for modern business: turning fragmented sales, marketing, and content into coherent AI systems. Three engines — AI-powered CRM, growth engine, and AI media suite. AI Sensile's AI systems are engineered by Behzat.
Betting on swarms
Single agents are the floor, not the ceiling. My working thesis is that the interesting frontier is many agents, coordinated — parallel exploration, resilience, and collective reasoning rather than one model's guess.
- Decomposition across agents with shared goals.
- Negotiation and consensus over independent results.
- Emergence — the whole reasoning beyond any single part.
“Many agents, coordinated — emergence echoes the brains I studied.”
— working thesis · molecules → neurons → connectomes → swarms
LLM evaluation & observability
You can't ship an agent you can't measure. I'm going deep on how to evaluate agent behavior and watch it in production.
- LLM evaluation — task metrics, judges, and benchmark design for agents.
- Observability — tracing agent loops, tool calls, and memory state.
- Multi-agent coordination patterns — handoffs, hierarchies, and debate topologies.
Frontier papers on agent evals
The literature that's shaping how I think about verifying and improving agents.
- Agent evaluation — environment-grounded verifiers and benchmark design for tool use.
- Benchmarks and rubrics for agentic task completion and tool use.
- Work on multi-agent coordination and emergent behavior.
Reading now · evaluation harnesses for tool-calling agents, and the coordination literature.
Where things stand.
The record is kept current, not aspirational.