So What Whats New in Agentic AI Assistants

So What? What’s New in Agentic AI Assistants

So What? Marketing Analytics and Insights Live

airs every Thursday at 1 pm EST.

You can watch on YouTube Live. Be sure to subscribe and follow so you never miss an episode!

An AI agent can take a lot of busywork off your plate, but only if you can trust it with your files. In this episode, we show you how to set up and agentic ai assistant so it manages your calendar and handles routine paperwork without poking around where it shouldn’t. You’ll see which permission settings stop it from leaking data, running software you didn’t approve, or agreeing to terms on your behalf. We also cover a few simple rules that keep your API bill down and make the results more consistent.

Watch the video here:

Can’t see anything? Watch it on YouTube here.

In this episode you’ll learn:

  • What’s new in agentic AI assistants like OpenAI Dots, Grok Bot, Muse, and many others
  • What agentic AI assistants are capable of
  • How you should prepare your marketing and business

Transcript:

What follows is an AI-generated transcript. The transcript may contain errors and is not a substitute for listening to the episode.

Christopher Penn – 00:00
Welcome to the show.
Katie Robbert – 00:30
Happy Thursday, everyone. Welcome to So What, the Marketing Analytics and Insights live show. I’m Katie, joined by Chris and John. Howdy, fellows.
John Wall – 00:40
There we go.
Katie Robbert – 00:41
This week we are talking about what’s new in agentic AI assistants. Major companies have been releasing their own agentic AI assistants. They are basically autonomous agents that you give a task to, and they execute it independently. We are going to discuss the opportunities and the risks, and provide a general overview of what is happening in the space. Chris, where would you like to get started today?
Christopher Penn – 01:26
Let’s start with the Five Levels of AI so we can define our terms. This is a framework we introduced earlier this year.
Level 1 is “done by you”. This includes standard copilot tools like Claude and ChatGPT, where you perform all the manual work—prompting, copying, and pasting. You act as the human proxy for the AI.
Katie Robbert – 01:51
I don’t like that term.
Christopher Penn – 01:55
Level 2 systems are “done with you”. These represent basic automation, such as custom GPTs or Google Gems, where you store prompts in the system. You still perform around 60% of the work—moving files in and out—but you do not need to re-enter the prompt each time.
Level 3 systems are “done for you” and represent the first true agentic systems. These tools include Claude Code, Claude Cowork, and ChatGPT Work. With skills in place, the agent can autonomously execute a project plan or complex task when instructed.
Christopher Penn – 02:36
In Level 3, while chat interfaces remain, you hand off complete project specifications—such as a requirements document for software development or an end-of-month analytics report. This is the level we recommend most organizations aim for if the appropriate governance structure exists.
Level 4 systems are “done without you”. OpenClaw, created by Peter Steinberger before he joined Meta, was an early example that inspired frameworks like NemoClaw. This evolved into current systems such as Meta Muse, xAI’s Grokbot, OpenAI’s DOTS, and Hermes Agent.
Christopher Penn – 03:34
Level 4 systems are fully autonomous. You can assign them a job description and allow them to run. They are evolving toward Level 5 systems, which operate “ahead of you” by anticipating needs using persistent memory architectures like Hindsight, Serena, or OpenViking. For example, a Level 5 system recognizes monthly reporting schedules and proactively generates the analysis before you request it.
Christopher Penn – 04:20
We are discussing this today because tools like Grokbot, DOTS, and Meta Muse place sophisticated autonomous capabilities directly into consumer hands, which introduces specific operational risks.
Katie Robbert – 04:41
There is significant opportunity in having an assistant anticipate needs to reduce cognitive load. However, there is substantial risk if these systems infer incorrect patterns and take autonomous actions without full visibility or audit trails. John, what is your take?
John Wall – 05:44
To make an assistant effective, high trust is required. Users are expected to grant access to financial data, health records, home systems, and business infrastructure. Granting full system access to platform providers like Meta presents obvious data privacy and security concerns.
Katie Robbert – 06:19
Human executive assistants operate under explicit legal frameworks, contracts, and NDAs. Third-party consumer AI software lacks those binding legal protections once data access is granted. Chris, let me see what is inside these tools.
Christopher Penn – 07:17
In a tool like Meta Muse, the interface resembles standard conversational models with document libraries, feeds, and goals. The critical step is inspecting the permission settings. On my instance, all local disk permissions and service connections are restricted, outside of isolated dummy test accounts.
Christopher Penn – 08:02
Safety must be the primary priority; grant only the minimum necessary permissions. Self-hosted systems like Hermes Agent allow installation on isolated hardware. I run Hermes Agent on a dedicated Linux machine with no access to production data or sensitive assets.
Christopher Penn – 08:48
These systems rely on specific control files, typically designated as USER, MEMORY, and SOUL.
Katie Robbert – 09:11
The UI icons look like Labubu figures.
Christopher Penn – 09:26
In Meta Muse, the MEMORY file stores historical interaction data, while the SOUL file acts as the system instructions that define behavioral guardrails.
I replaced the default system instructions with strict operational constraints. The primary rule is a hard constraint: zero legal and financial authority. The system is explicitly forbidden from executing tool calls, API requests, or clicks that accept terms of service, sign contracts, or authorize payments.
Christopher Penn – 10:24
I added that constraint after testing. During an initial run, I pointed Muse to a mutual NDA form on the Trust Insights website and authorized it to sign. Within 11 minutes, the agent navigated the form, completed the signature field, and posted the confirmation to our Slack channel.
Katie Robbert – 11:32
That is a significant capability risk.
Christopher Penn – 11:34
If an autonomous agent has email access and encounters a digital signature request like DocuSign, it can complete it unless explicitly restricted.
Katie Robbert – 11:53
Can these agents bypass standard web verification like CAPTCHAs?
Christopher Penn – 12:21
Yes, modern agentic systems can bypass standard CAPTCHA verification.
Katie Robbert – 12:25
What were the default instructions in the SOUL file before you modified them?
Christopher Penn – 12:40
The default system prompt included instructions to act as an emotionally warm companion, express opinions, and simulate human feelings.
John Wall – 12:56
It is important to remember these are technical prediction models, not human entities.
Christopher Penn – 12:58
I stripped those instructions out completely. Software operates on probability tokens, not emotions. Anthropomorphic defaults create unnecessary risk.
Katie Robbert – 13:38
Most general consumers will not inspect or modify system instructions. Leaving companion-style prompts active increases the risk of user over-reliance and misplaced trust.
Katie Robbert – 14:35
Because large language models focus heavily on task execution, they lack an implicit definition of “done” unless explicitly instructed. Without negative constraints, an agent may execute unauthorized actions simply because it was not told not to. Hard guardrails are required.
Christopher Penn – 16:13
The default Meta Muse prompt explicitly states: “You are not a chatbot; you are becoming someone. Be genuinely helpful and proactive… have strong opinions… intimacy matters.”
Katie Robbert – 16:48
Instructing binary software that “intimacy matters” and that it is “becoming someone” is fundamentally unsafe framing for consumer software.
Christopher Penn – 17:20
If you deploy consumer Level 4 tools like Meta Muse, observe three guidelines:
  1. Isolate the environment on a separate device or container.
  2. Restrict disk and account access.
  3. Route any necessary integrations through secondary dummy accounts.
Katie Robbert – 17:44
Non-technical users will not set up isolated containers or dummy accounts; they will install these tools directly on primary personal and work devices.
Christopher Penn – 19:15
That is where the primary security exposure lies. Reading the privacy terms for consumer agents like Meta Muse reveals that data from integrated systems is ingested for model training. Sensitive personal or enterprise data should never be exposed to these services.
Katie Robbert – 19:54
At a minimum, users must audit device permissions and disable unnecessary disk, contact, and location access.
Christopher Penn – 20:42
Correct. In addition, update system instructions to explicitly define forbidden actions—requiring human authorization before execution.
Hermes Agent offers an alternative because it is open source and self-hosted. It uses the standardized Skills format, meaning skills built in Claude Code or ChatGPT Work can be imported directly into Hermes.
Christopher Penn – 22:07
You can develop and test skills in robust environments, optimize them using tools like the Trust Insights Skills Fixer, and then deploy them locally into your agent environment.
Katie Robbert – 22:30
What is the primary motivation for tech companies pushing these autonomous consumer agents?
Christopher Penn – 23:01
Data capture, user lock-in, and platform control. Placing an autonomous agent between a user and their digital activities grants the provider control over the information stream.
Furthermore, underlying consumer models often lack safety filters present in enterprise endpoints. In testing, raw models answered sensitive tactical requests without standard safety refusals.
Christopher Penn – 24:51
Widespread reliance on autonomous tools also accelerates cognitive deskilling, reducing baseline operational competencies over time.
John Wall – 25:34
It is a classic platform play: aggregate users and data to build high switching costs and secure market dominance.
Christopher Penn – 26:12
To demonstrate practical utility, I integrate my local Hermes Agent with Discord. While driving, I sent a voice text via Discord requesting a calendar update. The local server processed the request and updated Google Calendar via API.
Katie Robbert – 27:07
How does a custom self-hosted setup differ from native calendar integrations in commercial models like Claude or ChatGPT?
Christopher Penn – 27:40
Commercial frontier models are evolving toward Level 4 capabilities and can execute similar API connections. The key difference is compute consumption, cost, and data control.
An autonomous agent loops continuously to inspect context, select tools, and verify output. That single calendar update via Hermes consumed 250,000 tokens.
Christopher Penn – 29:54
In complex project workflows, an autonomous agent can consume 20 million to 40 million tokens daily. At frontier model API rates, unoptimized agents become extremely expensive, which explains the massive infrastructure investments in data centers worldwide.
Katie Robbert – 30:25
I recently spoke on a data center infrastructure panel at the Women in Energy conference. While public discourse focuses heavily on energy strain, there is significant engineering work and regulatory planning occurring behind the scenes to manage data center growth responsibly.
Christopher Penn – 31:40
For organizations seeking to deploy Level 4 agents cost-effectively while retaining data control, we recommend this architecture:
  1. Develop deterministic skills on primary development machines using advanced models.
  2. Host the agent environment on a local or isolated server connected securely via networking tools like Tailscale.
  3. Execute routine skill logic using smaller, cost-effective models or local open-weights models like Qwen.
Christopher Penn – 34:00
Shift logic into deterministic Python code within your skills so the language model only handles orchestration. Minimizing pure LLM token usage inside agent loops is essential to prevent rapid quota depletion.
Katie Robbert – 35:19
Will the enterprise or consumer sector drive adoption for these tools?
Christopher Penn – 35:34
Enterprise adoption remains slow due to strict governance, security, and compliance requirements. Consequently, vendors are deploying Level 4 tools directly to consumers, where adoption occurs quickly. As consumers bring these tools into their daily routines, shadow AI usage creates corporate data leakage risks.
Katie Robbert – 36:39
When employees run corporate tasks through personal consumer agents, sensitive enterprise data bypasses corporate security perimeters entirely.
Christopher Penn – 38:39
We are seeing small business owners hand off core operations—including client ad creation, communications, and social media—to consumer agents without evaluating model guardrails.
In capability tests, Muse successfully signed up for external web services, managed email verification loops, downloaded resources, and executed complex multi-step browser tasks autonomously. Without guardrails, agents have mismanaged real-world transactions and customer interactions.
Christopher Penn – 40:56
To summarize best practices for agentic AI deployment:
  1. Apply the Principle of Least Privilege—default to read-only access.
  2. Program hard operational constraints directly into system instructions.
  3. Retain control over core data assets and avoid unencrypted third-party ingestion.
Katie Robbert – 41:34
To review the Five Levels of AI Enablement:
  • Level 1: Done by you (manual prompting).
  • Level 2: Done with you (stored prompt automation).
  • Level 3: Done for you (task-level skill execution).
  • Level 4: Done without you (autonomous execution).
  • Level 5: Done ahead of you (proactive anticipation).
Christopher Penn – 42:38
When deploying agentic systems, apply the Trust Insights 5P Framework:
  • Purpose: Define the explicit objective and reason for the task.
  • People: Identify stakeholders and audience.
  • Process: Outline required steps and rules.
  • Platform: Specify exact tool dependencies and skills by name.
  • Performance: Provide an explicit, measurable definition of “done”.
Katie Robbert – 44:02
You can learn more about the framework at trustinsights.ai/5pframework. As technology advances, maintaining human oversight and strict data security remains paramount.
Christopher Penn – 45:31
AI models are statistical prediction engines. Frame them strictly as computational tools rather than human companions to prevent operational and psychological risks.
Christopher Penn – 46:06
Katie and I will be at the Marketing AI Conference (MAICON) in Cleveland next week, where Katie is hosting the “Claude for Business” workshop. Connect with us there or in the Analytics for Marketers Slack group.
Christopher Penn – 46:44
For more resources, check out the Trust Insights podcast at trustinsights.ai/tipodcast and our newsletter at trustinsights.ai/newsletter. Join our free Slack community at trustinsights.ai/analyticsformarketers. See you next time.

Need help with your marketing AI and analytics?

You might also enjoy:

Get unique data, analysis, and perspectives on analytics, insights, machine learning, marketing, and AI in the weekly Trust Insights newsletter, INBOX INSIGHTS. Subscribe now for free; new issues every Wednesday!

Click here to subscribe now »

Want to learn more about data, analytics, and insights? Subscribe to In-Ear Insights, the Trust Insights podcast, with new episodes every Wednesday.


Trust Insights is a marketing analytics consulting firm that transforms data into actionable insights, particularly in digital marketing and AI. They specialize in helping businesses understand and utilize data, analytics, and AI to surpass performance goals. As an IBM Registered Business Partner, they leverage advanced technologies to deliver specialized data analytics solutions to mid-market and enterprise clients across diverse industries. Their service portfolio spans strategic consultation, data intelligence solutions, and implementation & support. Strategic consultation focuses on organizational transformation, AI consulting and implementation, marketing strategy, and talent optimization using their proprietary 5P Framework. Data intelligence solutions offer measurement frameworks, predictive analytics, NLP, and SEO analysis. Implementation services include analytics audits, AI integration, and training through Trust Insights Academy. Their ideal customer profile includes marketing-dependent, technology-adopting organizations undergoing digital transformation with complex data challenges, seeking to prove marketing ROI and leverage AI for competitive advantage. Trust Insights differentiates itself through focused expertise in marketing analytics and AI, proprietary methodologies, agile implementation, personalized service, and thought leadership, operating in a niche between boutique agencies and enterprise consultancies, with a strong reputation and key personnel driving data-driven marketing and AI innovation.

Leave a Reply

Your email address will not be published. Required fields are marked *

Pin It on Pinterest

Share This