Skip to content
Bot jobsJob breakdowns

OpenAI just launched Dot. Here’s how I’d actually use Dot, Grok Bot and Muse

A lot has happened in AI over the last few weeks. Actually, that sentence feels stupid now because you could probably write it every Monday until 2030. Big year. OpenAI has dots. xAI has Grok Bot.

Clifton MhlangaImported from X10 min read
Cliff305x article
See this runHouse 428 · 00577

Article

Job breakdowns

A lot has happened in AI over the last few weeks. Actually, that sentence feels stupid now because you could probably write it every Monday until 2030.

Big year.

OpenAI has dots. xAI has Grok Bot. Meta has Muse. Anthropic has folded Cowork into Claude. Then you have companies like Fellou approaching the same problem through the browser.

Everyone, apparently, would now like to give you some version of an AI employee.

Cool. But I think we are already asking the wrong question.

Most of the conversation immediately becomes: which one is best? Which model is smartest? Which agent has the best benchmark? Which one can run the longest without falling over?

I don’t think that is particularly useful.

The question I keep coming back to is much simpler:

What responsibility should I give each one?

Because after spending time with these products, reading how the companies themselves describe them and starting to build some of these systems for myself, I think there is a much easier way to look at the market.

Grok Bot is your organisation. Your dot gives your context agency. Muse is your personal operations person.

That is obviously an oversimplification. It is also probably the most useful framework I have found so far.

We have been heading here for years

Think about how most of us started using AI. You opened ChatGPT, typed something into a box and got an answer back.

That was basically it.

Then everyone discovered prompts.

Suddenly LinkedIn became people telling you to comment “PROMPT” so they could send you the exact 742 word framework they use to write emails.

Very strong period for civilisation.

But we figured out pretty quickly that a brilliant prompt with no understanding of your world was still limited.

So the conversation moved to context.

We started giving models our files, company information, examples of how we write and eventually access to things like email, calendars, CRMs and all the other places where our work actually lives.

The output improved because the model finally understood the world the prompt was sitting inside.

That is essentially what we now call context engineering.

And once you combine instructions, context, memory, tools and access to a computer, there is an obvious next question.

Why am I still sitting here asking you to do every individual thing?

Why can’t I give you responsibility?

That is where I think we are now.

The interesting transition is not really from chatbots to agents. It is from:

“Do this for me.”

to:

“This is yours.”

Small difference in language.

Massive difference in product design.

Grok Bot: build an organisation

This is probably the easiest one for me to explain because I think of Grok Bot like building a company.

You are the CEO.

Your first Bot might be your chief of staff. You give that Bot a job, context, standards and responsibility.

Then, where it makes sense, you start creating specialists around it.

Maybe one owns research. Another owns content. Another handles quality control. Another looks after product or operations.

The important part is that these are not simply separate chat windows.

They can maintain different responsibilities over time, work simultaneously, communicate and hand work between themselves.

That starts looking less like software and more like an organisation.

Say I was running a YouTube channel.

I could create a content lead whose responsibility is to understand what I talk about, what has performed well historically, what my audience cares about and, equally importantly, the stuff I simply refuse to make.

That agent might use a research specialist to look across YouTube, X and the broader market for topics worth exploring.

Another specialist could evaluate those ideas against the standard I have defined.

Maybe another writes an initial draft. Another reviews it.

By the time I get involved, I am no longer opening ChatGPT and asking:

“What should I make this week?”

I am looking at three ideas that have already survived the system.

That feels much closer to hiring.

And I think ownership is the useful word when thinking about Grok Bot.

You are not simply creating an automation that fires every morning.

You are creating a persistent role and giving it responsibility.

Routines are simply one way that responsibility gets triggered.

There is also an important distinction here between automation and agency.

An automation says:

“Follow these predefined steps.”

An agent is closer to:

“Here is the outcome I want. Work out how to get there.”

They are not the same thing.

Your dot: give your context agency

This is where OpenAI’s new dots get really interesting.

Because I would not start by building a bunch of complicated workflows inside one.

I think that misses the thing that makes a dot different.

If you have been using ChatGPT for a long time, it probably already knows quite a lot about you.

How you work. What you are building. Things you care about. Projects you have discussed. Problems you have been wrestling with. Things you keep saying you are going to do and, somehow, still have not done.

No comment.

A dot takes that existing idea of context and adds something really important to it:

Agency.

It has memory. It can use connected information. It has its own computer. It can keep making progress across projects.

More importantly, OpenAI is explicitly building it to proactively review information and surface useful work without you having to begin every interaction with another prompt.

That last part is the thing I find fascinating.

Imagine I have spent months talking to ChatGPT about launching a product.

We have discussed the problem, positioning, website, competitors, customers, funding, pricing, onboarding and probably fourteen different shades of blue for the bloody interface.

Everything except actually launching it.

A useful dot should eventually be able to see that gap.

Not because I wake up and ask:

“Why haven’t I launched this yet?”

But because it understands the goal, the work I have done and what remains unresolved.

Maybe it starts researching something I have been stuck on.

Maybe it pulls together the decisions I keep revisiting.

Maybe it surfaces the two things that are actually blocking progress and prepares the rest.

That is very different from a normal chatbot.

Your context now has agency.

And that is why I think calling a dot another AI employee slightly undersells it.

Grok Bot begins with the job.

A dot begins with you.

That means my dot could potentially understand my sales work, a product I am building, this newsletter, my calendar, things I have been researching and whatever else I have allowed into that context without me having to recreate myself inside a brand new system every time.

There are boundaries, obviously.

A dot is not going to wake up tomorrow morning, incorporate your company, fire your accountant and spend $70,000 on Meta ads.

Probably for the best.

Its unsolicited proactive behaviour is deliberately more restricted. It can notice something, research it, prepare work and surface a recommendation.

Depending on the task and permissions you have given it, actions can then happen from there.

But even with those boundaries, I think the model is incredibly interesting because the question changes from:

“What should I ask ChatGPT?”

to:

“Given everything you know about what I am trying to do, what should we be doing?”

Different game.

Muse: make the annoying stuff disappear

Then there is Muse.

I have been thinking about Muse much more as my personal operations person.

This is where I want the boring stuff to go.

Research a purchase. Help me organise a trip. Fill in a form. Handle a booking. Stay on top of some personal admin. Compare a bunch of options.

Do the thing that I was convinced would take five minutes and somehow has me sitting there 45 minutes later with sixteen browser tabs open.

Muse has its own persistent computer, can keep working after you close the app and is designed heavily around everyday life.

There is also a free tier, which makes it very easy to throw a bunch of these smaller personal jobs at it and see where it is useful.

Could I use Muse for work?

Of course.

Could I ask a dot to research a restaurant?

Probably.

Could I create a Grok Bot whose sole purpose in life is finding me a MacBook?

Absolutely.

Very strong employee.

That is not really the point.

I do not think every AI system I use needs to become my one AI system.

In fact, I suspect the opposite is going to happen.

We are going to become much better at allocating work.

I would not hire my CFO to book a dinner reservation simply because technically they know how to use the internet.

Wrong abstraction.

Muse can absorb a bunch of the everyday operational stuff I do not particularly want occupying my head.

That remains useful even if it never becomes the most intelligent model on Earth.

And what about Claude?

I can already hear someone typing.

“But what about Claude?”

Yes.

Please relax.

I still use Claude.

Claude is excellent, particularly for substantial project work and software development, and Anthropic is absolutely building agentic workflows.

In fact, Anthropic has now folded Cowork into Claude itself, so you no longer necessarily choose between chat and Cowork before you start.

Claude can take on long running work, use connected tools and execute recurring tasks.

So I would not say Anthropic is missing agents.

Not remotely.

What I think it currently lacks is the same product abstraction.

Grok Bot feels like:

Here is your employee.

A dot feels like:

Here is your AI counterpart.

Muse feels like:

Here is someone to help run your life.

Claude feels more like:

Give me the project and I’ll go do the work.

That might change.

Probably will.

Anthropic, I am still waiting for my little permanent Claude person.

I know you are busy.

Please stop building your life around one model

There is one thing I think we get completely carried away with in AI.

Benchmarks.

I cannot believe how much content I consume about benchmarks.

Model A beat Model B on this test. Astra can do this. Claude can do that. Grok has moved ahead on something else.

Then three weeks later another model ships and apparently everything we knew has become obsolete again.

I understand why we care.

These models are incredible, and different models genuinely are better at different things.

But I think there is a danger in designing your entire system around whichever company happens to be winning this week’s benchmark screenshot.

Models improve. Occasionally they regress. Pricing changes. Products disappear.

A model that is incredible at one task might be strangely average at another.

Something you had written off six months ago might suddenly produce better work for a particular workflow because its capabilities, context handling or tools have changed.

Again, this sounds suspiciously like hiring humans.

The smartest person on paper is not automatically the best person for every job.

Same mechanism.

This is why I am increasingly obsessed with platform agnostic systems.

I want the intelligence layer to be replaceable.

If my dot stops working for something I care about, I should be able to move that work elsewhere.

If Grok suddenly becomes exceptional at a workflow I am currently running in Claude, I want to test it.

If some company that does not exist today ships something ridiculous six months from now, I do not want to rebuild my entire operating system just so I can use it.

And the funny thing is that we now know how to do this.

We have spent the last few years learning the pieces.

We know how to create evals.

We know how to write skills and instruction files.

We know how to use Markdown files to describe how something should operate.

We know how to give models reference material, examples and structured context.

We know how to organise folders so the system is understandable beyond the particular AI product sitting on top of it.

So instead of burying your entire operating system inside one agent’s custom instructions, build a portable context layer.

The model becomes a component. Not the system.

The exact folder names do not particularly matter.

What matters is that the system describes what good looks like, what context matters, which rules should be followed, how quality gets evaluated and what outcome you actually want.

Then the model becomes replaceable.

Maybe the most expensive frontier model turns out to be unnecessary for part of the workflow.

Great.

Move that job somewhere cheaper.

Maybe a smaller model consistently produces better output for one very specific task.

Use it.

Maybe your research agent performs better on Grok, your coding system works better with Claude, your personal context lives in a dot and your life admin disappears into Muse.

Fine.

The goal is not brand loyalty.

The goal is a system that works.

This is not one AI to rule them all

That is probably where I land on all of this.

I do not think the future looks like everyone picking one model and using it for absolutely everything.

Maybe I am wrong, but right now that feels a bit like asking whether Slack, Salesforce or Excel is the better product.

Depends what the hell you are trying to do.

I would use Grok Bot when a responsibility needs an owner.

I would use a dot where understanding me and my broader goals is the advantage.

I would use Muse when I want a bunch of personal life administration to disappear.

I would use Claude when I need substantial work executed, particularly around development and knowledge work.

I would use something like Fellou when the job itself is heavily browser based.

But I would build the underlying system so I can change my mind.

The model matters.

Of course it does.

But I think we are going to spend less time obsessing over tiny benchmark differences and more time thinking about system design.

What context does this agent need?

What tools should it have?

What does it own?

What is it allowed to do?

When should it ask me?

When should another agent take over?

And if something better appears tomorrow, can I move the workflow without starting from scratch?

That is the skill.

Strangely, it feels like the last few years have been training us for it.

First we learned how to talk to AI.

Then how to instruct it.

Then how to give it context.

Now we have to learn how to delegate to it.

And eventually, I think we learn how to design systems where the intelligence itself is interchangeable.

That might be the bigger shift.

Because the next phase of AI is not really about getting a better answer.

It is about deciding what you no longer need to do yourself, then building the system so you are not trapped by whoever happens to provide the model underneath it.

Anyway, that is how I am currently thinking about dots, Grok Bot and Muse.

Not competitors fighting for one slot, but different people on the team with different jobs and different strengths.

And all of it is going to change again.

Probably sooner than I would like.

Still working on it.

Cliff

Published on grokbot.sh. Cite the public log, not a prompt pack.

Command Menu