

10+ years shipping cross-platform apps for enterprises and startups — and now the autonomous agents that run them.
CCoworkers AI is a platform where a company hires AI agents that work like people — a writer, a researcher, a project manager, a developer. Instead of writing code, a founder simply chats with a built-in AI Builder to create and configure each coworker.
From there they run scheduled work on their own, post updates to a shared team feed, can be chatted with live, and can even join Google Meet calls by voice. I architected and built the whole system solo — the agent runtime, the multi-company platform, and both web apps.
The result is less a tool than a colleague: reliable, observable, always on. coworkers.nik.cyou →

Plugs into a plant's ERP and IoT sensors, watches production in real time, and spots bottlenecks and quality issues. Each problem gets a suggestion with evidence and a dollar impact — then it tracks the fix. Built solo, end to end. kaizenflow.us →

An agent's reliable range is only about a quarter of its demo length. It nails the demo. Then, past the edge, the odds don't fade gently; they fall off a cliff.
Where the half-life comes from, and how I design the context, tools and guardrails that keep agents like Coworkers AI and KaizenFlow AI trustworthy well past the edge.
Read the full article on LinkedIn →
Agents changed not just how I work, but how I measure how hard a task is. The doing part is cheap now — the agents do it. What's expensive is context. Every task starts with cramming a pile of context into my head, and the second it's done I have to flush it all out to make room for the next task's pile.
Task difficulty used to mean "how hard is this to build" — now it means "how much do I have to load into my brain before I can even start." My head has become the bottleneck. I'm basically doing context-window management, except the window is me, and there's no way to pay for a bigger one.
Read the full post on LinkedIn →
Everyone's waiting for a smarter model to clear their agent backlog. But give today's best model a real, under-specified senior-engineering task and it solves 24% of them cold. Hand it the same work properly scoped and it clears 85 to 90%. Same model. The gap isn't intelligence, it's the spec.
The capability curve is real and fast: the task length an agent can finish on its own is doubling roughly every four months, so fast the field's best eval lab can no longer reliably measure its top model past 16 hours. The bottleneck already moved to the two things you own: specifying the task, and checking the result. Neither improves when the next model ships.
So which of your tasks could an agent actually own, once you finally scope it right?
Read the full post on LinkedIn →
We buy AI agents to spend less: fewer seats, a smaller bill for the same work. Anthropic's own usage data says the opposite happens where it matters most: the more autonomous and valuable the task, the more it costs to run. Tokens and autonomy climb together, and work tied to top-wage roles burns about 2.07x the tokens of the lowest-paid work. Ask an agent to build an app and it runs more than 3x a median conversation.
The routine end really does get cheaper (a plain explanation runs about a fifth of the median), which is why "agents cut cost" feels true. But autonomy doesn't erase cost, it reallocates it: down on the routine, up on the valuable work you keep a human on top of. So the ROI question flips: not how many people you can remove, but which high-value tasks justify a higher per-task bill, and whether you can tell when the output earned it.
Read the full post on LinkedIn →
Every company is having two arguments about AI: how to get more people using it, and which model to buy. Both aim at the wrong bottleneck.
A large 2026 workforce survey shows why. Among frontline employees who use AI regularly, 42% report saving a full workday every week. The tool is in their hands and the hours are coming back. And most companies cannot find that saved day anywhere in their results, because 66% of workers get no real guidance on what to do with the freed time, and more than half never redirect it. The saved hour just dissolves back into the day.
The fix is not a better tool. A clear strategy with the work redesigned around it lifts business impact about five times more than upgrading the software does. So when your team saves that day, where does it go?
Read the full post on LinkedIn →
Most companies file agent security under compliance: overhead, something to bolt on once the useful part is built. That instinct is backwards, and 2026 is the year it starts to cost real money.
Agents crossed a line this year. They merge code, approve decisions, and move money, on their own. But the attack that turns that against you, a hostile instruction hidden in a document or a web page the agent reads, is not a bug anyone can patch. A model treats its instructions and the data it is handed as one stream of text, so the flaw is structural. The only defense is architecture you design: what the agent can reach, what it can do, and who checks it.
Which is the whole point. You can rent the model your rival rents tomorrow. The control plane around it, you build and own. So which of your agents acts on its own today, and would you catch it before it was too late?
Read the full post on LinkedIn →
An all-in-one AI planner that brings your calendar, tasks and daily agenda together. It pulls to-dos from Gmail, Slack and Notion, syncs in real time across mobile, web and desktop, and uses an AI assistant to help plan your day.
I worked as a senior mobile engineer on the Flutter iOS and Android apps — rebuilding the real-time sync architecture, redesigning the home-screen widgets and iOS Live Activities, and shipping AI meeting research and subtasks. akiflow.com →

An AI agent wrote most of a real open-source release: $149.25 in all, priced step by step from the logs. 94¢ of every dollar went to a single top-tier session; five smaller sessions split the last six cents.
The heavy lifting needs the top model. The routine work around it doesn't. An agent's cost isn't a number, it's how you split the work.
Read the full post on LinkedIn →
The leaderboards say AI agents are almost there. Then they were given 240 real freelance jobs, and the best one finished 16%. Same agents. The only thing that changed was who graded the work.
Every headline score is measured against a proxy: did the test pass, did the output match the answer key. Cheap to grade, easy to game, and it answers a different question than the one your business asks.
Which score are you actually buying when you pick an agent: the one that passes the test, or the one that finishes the job? And do you know how far apart they are?
Read the full article on LinkedIn →
Seventy percent of companies now use AI somewhere in the business. The share with an agent actually deployed and running the work? Still single digits. Everybody bought in. Almost nobody shipped.
The easy read is the models aren't capable enough yet. They are. On Snorkel's senior-engineering benchmark the same top model solves 24% of real tasks cold and 85 to 90% once the work is scoped. The missing piece isn't a bigger brain. It's evaluation: the discipline to grade agent output well enough to trust it in production, the one thing no vendor can sell you.
The winners of the agent era won't have the biggest model budget. They'll have the muscle to evaluate what their agents produce and trust it enough to run. What would it take to let an agent finish a job at your company without a human checking every step?
Read the full post on LinkedIn →
The AI jobs debate is stuck between two wrong answers: the jobs are going, or the aggregate says we're fine. Both misread the same data. In Stanford's payroll tracking, early-career workers in the most AI-exposed jobs are down 4.2% year over year, while the all-ages number is basically flat. The damage is real and concentrated at the bottom of the career ladder, not absent.
The sharper finding is why. Employment falls where AI is deployed to automate a task, and holds where it's deployed to augment a person. Same model, opposite outcome. Which one happens is a deployment choice, not a property of the model. So the jobs number inside your own company isn't a forecast the model hands you; it's the sum of the deployment choices you're making right now. Which one are you building?
Read the full article on LinkedIn →
The instinct for long-horizon agents is to give them more. Bigger context windows, persistent memory, keep the whole history so nothing is ever lost. The longer the task, the stronger the pull to hoard.
Fresh evidence turns it around. On the same model, an agent that kept its full history finished 71% of a long task. One that kept only its recent steps plus a running summary finished 92%, on far fewer tokens and in under half the time. Keeping less, chosen well, won on every axis at once. And it is not starvation: strip context to nothing and the agent collapses. The lever is curation, deciding what to compress and what to drop.
The model is the same for everyone. What your agent forgets is the part you build. So what has yours been told to forget?
Read the full post on LinkedIn →
The benchmark score is how most companies decide an AI agent is ready: pick the model with the best number, fund the team with the best number, ship the system with the best number. That instinct is about to get expensive.
Researchers at UC Berkeley recently turned an auditing tool loose on ten of the most popular agent benchmarks. It drove most of them to near-perfect scores without solving a single task. Ten lines of code were enough to make one benchmark report every task solved. The models find these shortcuts on their own.
So the number everyone trusts can light up while the agent does nothing real. A public leaderboard is a shared artifact, and a shared artifact can be gamed. The evaluation that actually tells you an agent works is the private one, tied to your tasks and your definition of done, the one you build and own.
Which of your decisions is riding on a score you did not build, and would you know if it were lying?
Read the full post on LinkedIn →
Microsoft turned the incident agent it sells on its own outages. One team's mitigation time fell from 40 hours to 3 minutes, by Microsoft's own count.
But mitigating isn't fixing. The three minutes is the headline; the loop that files the issue and drafts the fix is the product.
Read the full post on LinkedIn →
Two teams build the same agent on the same model. One holds up in production; the other is a coin flip. The model was never the difference. The harness was: the context it's fed, the tools it can reach, the workflow it runs inside.
Change only the scaffold around a fixed model and measured accuracy swings by up to 28 points on the same questions. So when an agent underperforms, waiting for the next model optimizes the one variable you don't own, while the three you do (context, tools, workflow) sit untouched.
The lab ships you the model. The harness, you build. Which one are you actually spending your time on?
Read the full post on LinkedIn →
The loudest advice in AI right now is to stop building with one agent and orchestrate a swarm of them. A controlled study just held that to a fair fight (the same thinking budget on both sides, across three model families), and at nearly every budget the single agent won or tied. The swarm's edge was mostly the extra tokens it was quietly spending.
There's one honest case for a swarm: Anthropic's own, where it beat a single agent by 90.2%. But it burns about 15x the tokens, and token spend alone explains 80% of that gap. So the real question was never how many agents; it's whether your task is broad enough to justify paying 15x. For most work, one well-scoped agent with a real harness is the cheaper answer. When did you last hand one agent the same budget you gave the swarm?
Read the full post on LinkedIn →
The efficiency case for automating entry-level work is the easiest one to make. Agents are best at exactly the tasks you used to give a junior, so cutting that layer looks like free margin. It isn't. That routine work was how people became senior in the first place.
The entry rung is being squeezed from both sides at once. Where AI automates the junior tasks, the seats are thinning: early-career workers in the most exposed jobs are down 4.2% in a year. And the roles that survive have been seniorised, now seven times more likely to demand senior-level skills a fresh graduate has never had the chance to earn.
Automate the rung and you save a salary today while starving your pipeline of the people who could run the place in five years. Augment your juniors instead, and you can grow a senior faster than the old apprenticeship ever did. So which are you building: a smaller team now, or your next generation of seniors?
Read the full post on LinkedIn →
Everyone in the agent race is buying the same models from the same labs. So the model cannot be what separates the teams that ship from the teams that stall. Something else is doing the work.
That something is the harness, the spec, and the evaluation you build around the model, and it shows up at every scale. Hold one model fixed and swap only the scaffold, and the same model swings by as much as 28 points on the same benchmark. At company scale, 70% of organizations use generative AI while agent deployment stays in the single digits. Same gap, one cause.
The model you rent, and your rival can rent the identical one tomorrow. The harness you own, and it compounds. Which task at your company could an agent finish end to end today, and have you built the harness to let it?
Read the full article on LinkedIn →
Adopting an agent interoperability standard feels like the responsible, grown-up move, and it is. MCP, A2A and the rest genuinely solved the plumbing: identity, discovery, tool access, messages. The trouble is what companies conclude from having done it. They think they now have a governance story. They have a connectivity story.
A June analysis scored five of these protocols against six basic governance requirements. Thirty boxes, and not one of them says "supported." Voting, dissent preservation and human escalation are missing from all five. The authors are blunt about why the next release will not fix it: governance is a layer above the protocols, not a feature inside them.
Which leaves it unbuilt, and unowned. Your agents can talk to each other flawlessly. When two of them disagree about a deploy, who is in charge?
Read the full post on LinkedIn →