python Case study
Microsoft Clarity Analytics Extractor
Clarity records what people actually do on a page: where they rage-click, where they scroll to and stop, where they come back and leave. That is the most direct evidence available about a UX problem, and it is nearly useless while it stays inside the recording viewer.
The business problem
Session recordings and heatmaps are compelling to watch and hard to act on. Watching twenty sessions leaves you with an impression, not a ranked list of the pages that have a problem. To turn a behaviour signal into a decision, I need it out of the viewer and in a form I can sort, compare across pages, and check again after a change.
What I delivered
- I built a pipeline that pulls Clarity behaviour signals into structured output instead of leaving them in the session viewer.
- I aggregate to the page, so the unit of analysis is a page with a problem rather than one visitor's session.
- I shaped the output for before-and-after comparison, because the point of a UX change is to move a signal.
- I matched the structure to the other reporting, so behaviour data sits alongside search and commerce data for the same page.
Technical approach
- Aggregate to the page. One session shows you a possibility; a page with the same signal across many sessions shows you a problem.
- I built the output for comparison over time, because the useful question is rarely what is happening. It is whether the change helped.
- I kept the shape consistent with the other extracts, so I can look at one page from several angles at once instead of in separate tools.
- I treat a behaviour signal as evidence pointing at a page, not as a conclusion about why. To understand the cause, I still go back to the recording.
Result and evidence
Behaviour data can now be ranked and rechecked. That is what makes it usable for prioritising UX work instead of for illustrating a hunch.
Commercial value
Most UX debates are won by whoever sounds most confident. Page-level behaviour signals turn the argument into a question that has an answer.
Readable implementation brief
implementation_brief {
project: "Microsoft Clarity Analytics Extractor"
input: "Clarity session behaviour signals"
unit: "page-level aggregate, not individual sessions"
purpose: "rank pages by problem; verify after a change"
alignment: "shares shape with the search and commerce
extracts so one page can be seen from all sides"
boundary: "signals point at a page; recordings explain why"
}What this project shows
The judgement call was to aggregate to the page instead of showing off individual recordings. Recordings persuade; aggregates let you prioritise, and prioritising is the actual job.
I lined the output up with the other extracts on purpose. Data in three incompatible shapes is three tools nobody ever compares.