An April 2026 report raised a concern about code-execution risks in MCP servers. Since this article first appeared, the EU has revised the AI Act timetable. The two questions still matter, but they should not be collapsed into one deadline: which connected tools can an agent use, and what records will show what happened?
MCP, the Model Context Protocol, is one way an agent connects to tools and data. A client can request a tool call with arguments, and the server executes it. That makes the permissions granted to the server, and the inputs it accepts, important security questions.
What the disclosures are actually saying
OX Security documented command-execution flaws in applications that accepted unsafe MCP STDIO server configurations. In affected products, user-controlled configuration values could reach a shell command, sometimes despite command allowlists. The finding concerns how those applications create and run MCP servers. It is not a claim that every MCP tool call becomes a shell command.
The shape of an MCP assessment
A real MCP assessment answers four questions about a server before the agent is allowed to use it.
- Step 01What tools does this server expose, and what is the schema of each? An MCP server with eight tools has eight ways an agent can act through it. Enumeration before connection.
- Step 02What permissions does each tool require, and what does it return? A read-only directory listing tool is a different risk class than a tool that writes to a database or executes a shell command.
- Step 03How are tool arguments and server configuration validated? Check inputs that reach a database, file path, or command runner, including STDIO startup commands.
- Step 04What identity does the server run as, and what blast radius does that identity have? An MCP server running as root on the same host as your production database is not a tool; it is a privilege boundary that an agent now controls.
The reason these four questions matter together is that an agent in production answers them implicitly every time it makes a tool call. The assessment makes those answers explicit before the call happens.
The current EU timetable
The rules for Annex III high-risk AI systems now apply from 2 December 2027. Article 12 requires automatic event recording for systems classified as high risk. An agent is not automatically high risk because it uses MCP or works near a regulated sector; classification depends on the system's intended use. Teams should check the classification and plan their records accordingly.
Article 12 asks for logging that supports traceability and monitoring. It does not prescribe storing every prompt, tool argument, response, and model output. For an MCP-connected agent, a useful design is to link relevant tool actions to the applicable control decision and outcome while limiting sensitive data in the log.
The point is operational as much as regulatory. A reviewer needs to know what the agent attempted, which permission applied, and what the tool actually did. A current inventory and a clear record make that answer possible.
What an audit-defensible log actually looks like
- The log is structured, not free-text. Each event is a record with a timestamp, an actor, a source, an action, and a result.
- The record identifies the protected interaction or finding, the applicable control, and the outcome with supporting evidence.
- The log links relevant tool calls, control decisions, and results; sensitive prompts, arguments, and responses are captured only when justified.
- The log is exportable in a format a reviewer can use, such as JSON or CSV, with a clear mapping to any applicable requirement.
These properties are familiar in security engineering. The work is to apply them to each supported agent and tool path before expanding permissions, then test whether a reviewer can reconstruct the result.
HikmaAI assesses supported MCP paths, applies controls to protected interactions, and connects the resulting records for review. The value is in being able to show what was assessed, what was protected, and what happened within the scope actually covered.
Start with the tools your agents can reach. Test them, set the boundary, and make the resulting actions understandable to the next person who reviews them.


