Malware traffic analysis often involves spending a significant amount of time working through HTTP sessions. Fiddler sessions, response headers, request bodies, redirect chains. The patterns are there, but finding them requires manually scanning hundreds of sessions and knowing exactly what to look for. I wanted to improve that workflow by connecting Fiddler directly to a large language model so I could ask questions about captured traffic in plain English.
The result is Fiddler-MCP-Server, an open source tool that bridges Fiddler's traffic capture to an LLM using the Model Context Protocol. An analyst can capture a browsing session and then ask things like "which sessions have suspicious redirect chains" or "show me the headers for session 47 and explain the caching behaviour" and get a structured analysis back in seconds.
The problem with manual traffic analysis
When you are triaging a potential malware delivery chain, you are typically looking at dozens or hundreds of HTTP sessions. You need to identify the initial redirect, track cookie propagation, spot TDS fingerprinting in query parameters, find the final payload delivery, and correlate it all against known patterns. That process is entirely manual in Fiddler. You click through sessions one at a time, inspect headers, decode bodies, and build a mental model of the infection chain.
LLMs are good at exactly this kind of pattern recognition when given structured data. The challenge is getting Fiddler session data into a format the model can consume, and doing it in a way that feels like a natural conversation rather than a copy paste exercise.
Architecture: four components, one data path
The system has four components. Each one does exactly one job, and data flows through them in sequence.
Fiddler captures traffic and publishes each completed session as JSON to a local HTTP endpoint. A CustomRules script fires after every response, serialises the session data, and POSTs it to the staging server.
static function McpTryPost(oSession: Session): void {
try {
if ((oSession.oResponse == null) || (oSession.responseCode == 0)) return;
var json: String = McpBuildSimpleJson(oSession);
McpHttpPost(json);
} catch (e) {
FiddlerApplication.Log.LogString("MCP error: " + e.Message);
}
}
The key design decision here is pushing data on every response rather than batching. This means the staging server always has the most recent traffic available for queries, and the analyst does not need to manually export or refresh anything.
The staging server is a Flask application that buffers sessions in a ring buffer and exposes REST endpoints for different data views. Headers, response bodies, statistics, timelines. Each endpoint returns clean JSON that can be consumed by any HTTP client.
@self.app.route('/api/sessions/headers/<session_id>', methods=['GET'])
def get_session_headers(session_id):
with self.session_lock:
for session in reversed(self.live_sessions):
if str(session.get('id', '')) == str(session_id):
return jsonify({
"success": True,
"session_id": session_id,
"request_headers": session.get('requestHeaders', {}),
"response_headers": session.get('responseHeaders', {}),
"found": True
})
Keeping it as plain REST means the staging server is debuggable with curl and testable independently of the MCP layer. That separation saved significant time during development.
The MCP bridge translates LLM tool invocations into REST calls against the staging server. Each MCP tool maps to one REST endpoint. When the model decides it needs session headers, it calls the tool, which calls the endpoint, which returns the JSON.
@mcp.tool()
def fiddler_mcp__session_headers(
session_id: Annotated[str, Field(description="Session ID from live_sessions.")],
) -> Dict[str, Any]:
"""Fetch the HTTP headers for a captured session."""
return client.get_session_headers(session_id=session_id)
The current server exposes ten tools covering live sessions, traffic search, headers, response bodies, session comparison, statistics and timelines. Each remains deliberately narrow, with the MCP layer translating the request into the corresponding REST operation.
The Python client orchestrates the conversation. When a tool returns data, the client injects the JSON result into a new prompt and asks the model to analyse it in the context of the original question. Whatever the bridge returns becomes part of the prompt sent to the model.
tool_result = self.call_tool(tool_name, arguments)
self.conversation_history.append({
"role": "tool",
"tool": tool_name,
"content": json.dumps(tool_result, indent=2)
})
analysis_prompt = f"""The tool '{tool_name}' returned this result:
{json.dumps(tool_result, indent=2)}
Please analyze this result and answer: "{user_query}" """
analysis_response = self.model.generate_content(analysis_prompt)
What this enables
With the bridge running, an analyst can capture traffic from a suspicious site and immediately start asking questions. "List all sessions that returned JavaScript content." "Show me the full redirect chain from session 12 to the final landing page." "Are there any sessions with unusual Set-Cookie headers that might indicate TDS fingerprinting?"
The model has access to the same data the analyst would manually inspect, but it can process all sessions simultaneously and surface patterns that might take a human analyst several minutes to find through manual inspection.
The conversation history means follow-up questions work naturally. Ask about a specific session, then ask "compare that to session 23" without re-specifying context. The model retains the prior tool results and builds on them.
Design decisions that mattered
Separating the staging server from the MCP bridge turned out to be the most important architectural choice. During development I could test data flow by curling the REST endpoints directly, without needing the LLM in the loop. When something went wrong, I could immediately isolate whether the problem was in data capture, staging, or the MCP layer.
Using a ring buffer for session storage keeps memory bounded. In a malware analysis session you might capture thousands of requests across dozens of sites. The ring buffer drops the oldest sessions automatically, keeping the system responsive without manual cleanup.
Making tool names explicit and descriptive, like fiddler_mcp__session_headers rather than generic names, helps the model select the right tool. The model sees the tool list and descriptions, and clear naming reduces incorrect tool selection significantly.
The model still got tool calls wrong
Explicit tool names helped the model choose the right operation, but they did not eliminate bad calls. I started seeing requests for tools that looked perfectly reasonable but did not actually exist. The model might ask for get_sessions rather than fiddler_mcp__live_sessions, or get_body rather than fiddler_mcp__session_body. It might choose the right tool but send id when the schema expected session_id. None of those guesses are strange.
They are still wrong.
That becomes a different problem once the model is operating software rather than simply generating text. I did not want the reliability of the workflow to depend on whether the model remembered an exact function name or argument schema every time. Adding more instructions to the prompt helped, but it did not solve the underlying problem.
I moved tools execution into Python
The better fix was to stop treating the model as the execution layer. The current client acts as a Python harness between the LLM and the MCP server. It discovers the available tools, binds their schemas to the active model, and handles the actual MCP calls. The model does not talk directly to Fiddler.
The data path now looks like this:
Fiddler Classic
|
| HTTP session JSON
v
enhanced-bridge.py
|
| REST
v
5ire-bridge.py
|
| MCP tools/list
| MCP tools/call
v
Python client / harness
|
| native tool calling
v
LLM provider
The model can decide that it needs a session body or a traffic search. The harness decides whether that request maps to a real capability and whether the arguments are valid before anything is executed. The model asks. The harness decides. The tool acts.
The MCP layer remains independent of the model provider. The same client can currently bind the tools to Gemini, DeepSeek, or models accessed through OpenRouter. The staging server and bridge do not change when the model does.
Once that control existed in Python, I could deal with common model mistakes in code rather than trying to prompt them away. The client maintains a small set of aliases for calls I had actually observed during testing:
get_sessions
|
v
fiddler_mcp__live_sessions
get_body
|
v
fiddler_mcp__session_body
It also handles naming variations such as missing MCP prefixes and dot notation. Arguments go through the same process. The session body tool expects something like this:
{
"session_id": "262"
}
A model may instead produce:
{
"id": 262
}
The harness can normalise that before the request reaches the MCP server. The same mapping exists for searches. A host argument can be rewritten as host_pattern, and url as url_pattern. Malformed query forms are rejected with a correction rather than passed through blindly.
There is still a hard boundary underneath that convenience. Before execution, the client compares the requested tool against the tool catalogue returned by tools/list. If the resulting name does not exist, the call stops. The model can be flexible when reasoning over JavaScript, redirects, or suspicious HTTP behaviour. There is no advantage in allowing that flexibility when selecting an executable function.
The harness became part of the control boundary
Tool name correction was only the first use for that layer. Once the harness existed, it became a natural place for other workflow controls. The client tracks which session bodies have been analysed in the current investigation. If the model asks for the same body again, the request can be stopped unless the analyst explicitly asks for a refresh. It can avoid retrieving large media bodies that are unlikely to help with JavaScript or malware analysis.
Tool calls and their arguments can be logged. The tool surface itself is deliberately narrow. Searching session metadata is different from retrieving a response body. Reading traffic is different from clearing the capture buffer.
Breaking those capabilities into individual tools gives the harness somewhere to apply different controls later. This is where the project stopped being only an MCP integration exercise for me. It became an agent security exercise. The question was no longer just whether an LLM could use Fiddler. The better questions became: what can it read, what can it invoke, what does the harness validate, and what evidence is retained afterwards?
Building this changed how I think about agentic workflows. Prompting still matters. Clear tool names, descriptions, and schemas reduce ambiguity. They make it easier for the model to select the intended action.
Prompts influence behaviour. Code can enforce it.
Giving a model access to tools is relatively easy. The harder engineering problem sits between the model deciding that it wants to perform an action and the system actually performing it. For this project, that boundary is now the Python harness.
The LLM does the pattern recognition. The MCP tools retrieve the evidence. The harness controls execution. The analyst makes the final decision. That is a better mental model for me than treating the LLM itself as the agent.
What comes next
The original version of this project was primarily about connecting Fiddler traffic to an LLM. The current version has moved further towards a controlled investigation workflow. There are now ten MCP analysis tools and an investigation flow that can start with suspicious sessions, retrieve selected JavaScript or HTML bodies, pivot into related hosts, and reconstruct an infection chain. The underlying evidence remains visible to the analyst.
I want to keep developing that part rather than simply adding more tools. The interesting question for me is not how much authority I can give the model. It is how little authority it needs to make the analyst meaningfully faster.
The full source code is available on my GitHub.