Skip to content

What an AI agent audit log must capture

An AI agent audit log must capture who acted, which client called, which account was reached, what was tried and how it ended. Argument contents stay out.

Nachi Raman Updated 9 min read
A single audit log row expanded into labeled fields: actor, actor kind, connection, tool, classification, outcome

What must an AI agent audit log capture?

An AI agent audit log must capture who acted, and which client the call came through. It must also capture which account the call actually reached, what was attempted, whether it was allowed, and how it ended. Elaichi writes one entry for each connected-tool call that reaches execution (a call refused earlier writes none). Each entry records the operation and tool, the connection, the classification and the approval decision, then the result, with an error code when the call fails. It logs the one path argument that names the object, as the target id, and nothing else about the arguments.

The test is a Friday afternoon. A Salesforce opportunity changes owner, and nobody says they did it. Three people had an AI assistant connected to that CRM that week, and two of them hold two Salesforce accounts each. A useful log closes that question in one read.

The standards ask for the same shape. NIST SP 800-53 control AU-3 lists what an audit record should establish. That is the type of event, when and where it occurred, its source, its outcome, and the identity of anyone or anything associated with it. OWASP's Logging Cheat Sheet frames it as when, where, who and what, with a result status saying whether the action succeeded.

What is being logged here is MCP traffic. MCP (Model Context Protocol) is the standard way an AI assistant calls tools in other apps. The MCP specification's tools page asks clients to log tool usage for audit purposes. Elaichi keeps the record at the organization-wide endpoint every client calls, so it does not depend on which client a person used. Audit events and application logs share one record shape, so one query answers "what happened" without anyone correlating two systems by eye.

Who acted, and which client did the call come through?

The row carries the actor, the surface and the client. A call from Claude, ChatGPT or Cursor is recorded under the person who signed in, with surface mcp and the OAuth client named, and those three clients are marked verified. A client that signs in through a loopback address, such as Claude Code, shows the name it registered with, marked unverified. The actor_kind values include user, system, staff, scim, api_token and ai_assistant, and ai_assistant marks the Elaichi Agent only. All of it is written at the point of action. None of it is guessed afterwards from a user agent string.

That distinction is the point. A value parsed out of headers is an opinion formed at read time, and it breaks the moment a client changes its user agent. A recorded field is evidence. Six months later a reviewer reads the surface and the client on each row, with no heuristic in the middle. OWASP's cheat sheet leaves room for exactly this, describing the who of an event as a human or machine user.

Actor names resolve server-side. A member who has since left renders as "Former member" rather than dropping out of the trail, so a departure cannot quietly edit history.

Elaichi staff impersonation is recorded and attributed in the customer's own audit log. Support appears as staff, not as the customer, which keeps the chain of custody intact when somebody steps in to reproduce a problem.

Which account did the call actually reach?

The connection field names the account the call reached, taken from the execution rather than from the intent. Calls that run and calls that fail at execution both carry it.

This is the field people underestimate until they need it. A person may hold a personal Notion workspace and a corporate one. "Which of my two Notion workspaces did the agent write to?" is the first question after an unexpected change, and an intent-derived value cannot answer it. The caller asked for Notion, and the call landed on one connection. If those two ever differ, only the execution-derived value is true.

The same field separates a personal connection from one shared with a team or the whole organization. A write through the shared finance account is a different incident from the same write through somebody's personal login, and the row says which without a follow-up.

What was attempted, and was it a read, a write or a delete?

The row records the operation and the tool, plus the classification of that operation as a read, a write or a delete. The classification is what a reviewer scans first, because it separates an afternoon of reads from the three calls that changed something.

Tool names are labels. Whoever edits a connector's documentation can change a tool's advertised name, so when a governance rule is written, Elaichi records the tool's underlying operation from the catalog. Governance binds the operation, never the label, and the record reflects that binding. The reasoning is in why a block matches the name but an allow matches the operation.

Two kinds of call look as if they might escape the row, and neither does. In Elaichi, connected tools are never listed one by one, however few there are, so a model reaches them through search_tools and execute_tool. The execute_tool call is only a naming indirection. It unwraps to the same name and arguments and passes the same gates, so the row names the real tool rather than the wrapper. Synthetic tools, which chain steps over other tools, send every step through the same restriction checks as any other call. A run is audited as one row with the tool's name, status, duration and counts of steps, calls and retries. It holds no step results, and the step calls are not separate rows.

Was the call allowed, and how did it end?

The row records whether the call was approved and how it ended. A call refused at execution is still a row. A call refused earlier is not. OWASP's cheat sheet lists authorization failures among the events to log, so know where Elaichi's trail stops.

Elaichi writes one row for each connected-tool call that reaches execution and then runs, fails or is refused there. One example is a call for an account Elaichi cannot match. Another is a call after a restriction changed once the tool was listed. A call refused before that point writes no row. A call to a tool a restriction withholds is refused before execution. A missing MCP scope, an unknown tool name or a rate limit stops a call first. A call that waits for a person's approval writes no row until it runs. A connected tool reached through the toolbox execute operation is recorded as a toolbox execution row instead.

Timing matters when you tighten a rule. A change to a role or a restriction takes effect within about two minutes, through a short cache. A call that succeeds shortly after you tighten a rule is expected, not a bug.

Grant revocation is different, because its flag is re-read on every call with no cache. In Elaichi, removing or suspending a member revokes every live grant in the same transaction as the membership change. For removal, suspension and revocation, the next call is the accurate answer. The contractor offboarding walkthrough covers the connection preflight that runs alongside it.

Why does the log leave out argument contents?

Because the trail travels further than the call did. Elaichi logs the one path argument that names the object, as the target id, and nothing else about the arguments. So a row says which object the call acted on, when it names one, as the target id, and not what the call wrote to it. That is a real loss. You cannot replay a call from the trail, and "what exactly did it write" needs the third-party application's own record of the change.

The reason is where the trail goes. Audit records are visible to the organization, readable by the in-product assistant, and built to be forwarded to a SIEM. Argument contents are the part of a call most likely to carry customer data, credentials pasted by a user, or the text of a contract. A field copied to three places should not hold the payload. The target id is enough to scope an incident, and cheap to expose. Argument contents are neither.

NIST makes the same trade explicit. The discussion under AU-3 warns that audit records can reveal personally identifiable information, especially when the trail records inputs. Its enhancement AU-3(3) limits that information to the elements a privacy risk assessment names. OWASP's list of data to keep out of logs covers the same ground: access tokens, passwords, connection strings, encryption keys and sensitive personal data.

Which error message gets written down?

A failed call produces two error strings, and only one of them is ever written to the trail. The row carries an error code. The caller receives a message derived from the third party's response body, and that message is never written anywhere else.

The audit-side text is never derived from the request or the response. A remote error body in a record that the whole organization and the in-product assistant can read would be third-party payload leaving the system through the log pipe. It is the same leak as logging argument values, wearing a different hat.

Promoted metadata follows the same instinct. It is a fixed allowlist, not a flattening of whatever keys arrive. Metadata keys can be influenced by users, and unbounded flattening would let one organization's traffic grow the field namespace for a whole tenant. In practice, your ingestion schema stays predictable when somebody adds a field in a third-party application.

Where does the log live, and who can read it?

Each organization's audit history is scoped in the type system rather than by a WHERE clause in a query. The failure modes are not symmetric. A dropped WHERE clause leaks, and a wrong scope returns nothing. Customer-visible audit records and internal application logs sit apart, which matters more than usual because the in-product assistant can read the audit log.

Elaichi has three regions, EU, US and APAC, chosen when the organization is created. For EU and US, the organization's data store and the execution of its org-scoped requests and tool calls stay in that jurisdiction. APAC uses best-effort placement and is not a residency guarantee. Every organization's audit trail, whatever its region, is stored in one log instance in the EU. The security page carries the rest of the data-handling detail.

NIST's AU-9 asks for audit information to be protected from unauthorized modification and deletion, and Elaichi's trail is append-only. It reads newest-first, cursor-paginated, and filterable by free text, category, actor, action kind and time. The Auditor role is read-only and a free seat, so a compliance reviewer does not consume a license. Auditor also lacks tool:execute, the permission that gates the whole MCP endpoint, so a reviewer can read the trail and cannot make calls. Seat classes sit on the plans page.

Can you send the log to a SIEM?

Not on Gold. Beyond the in-app trail, export to your own Datadog comes with the Black plan, which is launching soon. A SIEM is the system that gathers security logs from across a company, so forwarding puts agent activity next to everything else your security team already watches.

Splunk HEC and Microsoft Sentinel destinations are accepted as destinations, but Elaichi delivers events only to Datadog. Forwarding is filterable by log type, where an absent filter forwards everything and an empty list forwards nothing. The audit trail inside Elaichi is on Gold.

What can the log not tell you?

Three limits, stated plainly.

The trail is eventually consistent. A row can take a moment to appear, so an empty screen one second after a call is not proof that nothing happened.

The log cannot show what the user typed. The prompt-injection write gate lives in the Elaichi agent window and does not apply to a raw tool call. An MCP server never sees the user's prompt, so on POST /mcp there is nothing for that gate to inspect. What does hold there: role-based permissions per operation, the forbidden classification, output redaction, OAuth scope limits and the audit record itself.

Deleting an organization deletes its credentials and tears down the workspace, but it leaves three stores behind: the audit history (each record ages out under the log server's 90-day retention, counted from when it was written), analytics events and the credential service's connector configuration rows. The deletion returns that residue by name rather than reporting a clean sweep. If your retention policy needs proof that log data is gone, have that conversation before you sign.

When do you not need a log this detailed?

If one person uses one assistant against one application with read-only access, that application's own activity history probably answers every question you will ask. Salesforce, Notion and Jira each keep some record of who changed what. Adding a control plane to observe a single user is overhead with no reader.

A detailed row starts to pay when three things are true at once: several people, several accounts of the same application, and at least one write. That is the point where the intended account and the reached account can differ, and where "who did this" stops having an obvious answer. The case for waiting is made at more length in when an MCP gateway is premature.

How this record maps onto an audit is in SOC 2 evidence for AI agents, and more on deciding what agents may reach sits in the governance posts. The applications these rows get written about are in the connector catalog, and the team-by-team rollouts are on use cases.

FAQ

Frequently asked questions

What should an AI agent audit log include?

At minimum: the actor, the surface and client, the account the call actually reached, the operation and tool attempted, its classification as a read, write or delete, whether the call was approved, how it ended, and an error code on failure. Elaichi's audit log covers these, because it writes one entry for each connected-tool call that reaches execution (a call refused earlier writes none).

Does Elaichi log the arguments an AI agent sent to a tool?

Elaichi logs the one path argument that names the object, as the target id, and nothing else about the arguments. The trade-off is that you cannot rebuild a payload from the audit trail, and have to consult the third-party application's own record for that. The reason is exposure: audit records are visible across the organization, readable by the in-product assistant and built to be forwarded to a SIEM, so values are the wrong thing to copy into them.

What is actor_kind in an audit log?

The actor_kind field records what sort of actor performed an action, with values including user, system, staff, scim, api_token and ai_assistant, which marks the Elaichi Agent. In Elaichi it is written at the point of action rather than inferred afterwards from a user agent string. A call from Claude, ChatGPT or Cursor is recorded under the person who signed in, with surface mcp and the OAuth client named.

How quickly does a restriction change show up in the audit log?

A role or restriction change in Elaichi takes effect within about two minutes, because it resolves through a short cache and edge propagation. Calls made in that window still follow the old rule and are logged as such. Grant revocation, member removal and suspension are different: they are effective on the next call, since the revocation flag is re-read on every call with no cache.

Can Elaichi send audit logs to a SIEM?

Not on Gold. Beyond the in-app trail, export to your own Datadog comes with the Black plan, which is launching soon. A Splunk HEC or Microsoft Sentinel destination is accepted but does not deliver events. Forwarding is filterable by log type: an absent filter forwards everything, and an empty list forwards nothing.

Put agents to work on your own systems

14 days on Gold, no credit card. Start with one app and one team.

Works with
Claude ChatGPT Cursor and any other MCP client, or the Elaichi Agent.
When the trial ends
Nothing is deleted. Connections, roles and the audit log stay where they are, so subscribing picks up exactly where you left off.