Research

FingerprintJS Pro and Kaidn, side by side in 24 privacy browser sessions

Alex MugoFounder, Kaidn
11 min readRevised
We ran both engines in the same page loads across Tor, Mullvad, Brave and Camoufox. Every number, including the two rounds where our own detector was the one that was wrong.

We build a fraud scoring API. Part of what it scores is a browser fingerprint, so we wanted to know two things vendors in this market do not usually publish:

  1. Does our detector actually fire on the browsers it claims to catch, or is that just design intent nobody ever measured?
  2. How does it compare to the market leader when both engines see the exact same page load?

So we built a rig that runs FingerprintJS Pro and our own collector in the same page, in the same browser, on the same visit. Then we ran real privacy and anti-detect browsers through it as normal desktop apps rather than as headless automation.

This post is every number that came out, including the two rounds where we were the ones who got it wrong. It is longer than a marketing post because the numbers only mean something with the method attached.

First, what "the agent" is#

Skip this if you already know.

Browser fingerprinting vendors do not run inside your server. They ship a JavaScript agent: a script your site loads into the visitor's browser. It measures a few hundred properties of that browser (how it draws to a canvas, which fonts it has, its audio stack, screen size, graphics driver strings, time zone) and sends them to the vendor's API. The vendor turns that into an identifier and a risk signal.

That design has a dependency most risk dashboards never show you. The script has to load and run before anything else can happen. It is fetched from someone else's domain, and it runs in a browser the visitor controls. Either of those can stop it.

When it does not load, you do not get a cautious verdict, a raised score, or an error row. You get nothing at all, and nothing looks a great deal like a clean session.

That was the first thing we set out to measure.

The setup#

  • Both engines in one page. FingerprintJS Pro's agent and our collector run in the same page load, so neither gets a different environment to judge. About 160 browser features are captured per session alongside both engines' output.
  • Real desktop apps, not headless. Privacy browsers were launched as normal applications after a manual first run, because running them headless changes the exact signals under test.
  • Sessions counted one by one, never averaged. Anti-detect tools randomise on every launch, so an average across launches hides the thing you are looking for.
  • Two machines, one country. Every session in the table below was captured on one Mac. The wider corpus also carries an honest Windows PC, which is what candidate signals are cross-checked against before any of them can ship. Two computers is still a real limit, and it is repeated in the limitations section rather than buried there.

Tor needed Security Level Standard and network.proxy.allow_hijacking_localhost=false to reach a local capture page at all. Otherwise it routes 127.0.0.1 through SOCKS and the page never loads.

What each engine gives you, and why the two numbers are not the same number#

This is the part that makes the results readable, and it is where most comparisons quietly cheat.

FingerprintJS Pro's browser agent returns a visitor_id, an event_id, a suspect_score, and a sealed_result that was null in every session we ran. Setting extendedResult: true added nothing. Their richer signals, including their tampering and anti-detect machine-learning scores, need a server-side lookup keyed on that event_id with a secret key. We report the browser agent's suspect_score, because that is what the agent itself exposes.

Our collector (@kaidn/fp) returns an antidetect_score from 0 to 100, the list of signals that fired, and a rough confidence. That score is worked out in the browser, by rules, in the open. There is no model.

Two warnings before you read any table below:

  • These are two vendors' internal 0 to 100 scales and they are not interchangeable. Their line for "suspicious" is not published, and ours is not their scale. Compare the direction and the behaviour, never the absolute values.
  • Our antidetect_score is a collector output, not a /v1/score verdict. At the time of writing the API accepts the field and does not yet store or score it, so nothing in this post describes a production block decision. Saying otherwise would be exactly the overclaiming we are trying to measure.

How our score is built#

Worth explaining in full, because every figure in this post is one of four combinations of it.

Seven signals, each with a weight:

signalweightwhat it means
tampered45a built-in browser function was overwritten and the patch shows
contextMismatch40the browser describes itself differently in a background thread
osMismatch40something that reveals the real OS contradicts what the browser claims
engineMismatch35the real JavaScript engine contradicts the browser it claims to be
noiseInjected30the same canvas drawing reads back differently within one session
fontStandardized30the fonts match a known anti-detect profile
headless15automation tells: webdriver flag, headless user agent, no languages

They combine as a noisy-OR rather than a sum. Each signal that fires multiplies the chance the session is clean, and the score is what is left:

the noisy-OR, in full
P(clean) = Π (1 - weight)        score = round((1 - P(clean)) × 100)

So the score can never pass 100, and the second signal always adds less than the first. That is deliberate. Two signals agreeing is stronger than one, but not twice as strong.

Confidence is separate and blunter. A signal counts as strong if its weight is 30 or more. Two or more strong signals is high, one is medium, none is low. Remember that rule, because it is what makes our worst result in this post worse than it first looks.

That is where every number below comes from:

scorewhat firedarithmetic
0nothingno signal
15headless only1 - 0.85 = 0.15
30one 30-weight signal only1 - 0.70 = 0.30
41headless + noiseInjected1 - (0.85 × 0.70) = 0.41
51noiseInjected + fontStandardized1 - (0.70 × 0.70) = 0.51

The results#

Three runs per session unless stated. Every row says who was right, explicitly.

sessionrunsKaidnFingerprintoutcome
Chrome 151, driven headless315suspect 7Both correct. Ours is the headless signal firing on a genuinely headless browser, which is what it is for
Brave, real app30suspect 0Both correct. An honest browser scored honest by both
Tor Browser, honest user, real app351suspect 21We were wrong and they were less wrong. See below
Mullvad Browser, honest user, real app351⛔ script blocked 3/3We were wrong. They had no verdict at all
Camoufox pretending to be macOS, on macOS90 in 6, 30 in 3⛔ script blocked 9/9Nobody caught it. We missed it two thirds of the time
Camoufox pretending to be Windows, on macOS930 in 9⛔ script blocked 9/9We caught it, unopposed

Four findings come out of that table, and they do not all point the same way.

Finding 1: their script never ran on 21 of 24 sessions#

Mullvad 3 of 3, Camoufox 18 of 18. The Mullvad failures were fast, 316 to 833ms, which is what a blocked request looks like rather than a slow one.

The cause is not clever. Both of those browsers ship uBlock Origin, and uBlock subscribes to EasyList and EasyPrivacy by default. Both lists carry rules for Fingerprint's domains:

the two rules that do it
EasyPrivacy    ||fpjs.io^$third-party
EasyList       ||fpnpmcdn.net^

The first blocks their API domain whenever it loads from a site other than the one being visited, which for a vendor embedded on customer sites is every time. The second blocks their CDN outright.

We confirmed it two ways rather than one. From the lists: 164,000 lines of EasyPrivacy, EasyList and uBlock's own privacy and badware filters. By experiment, with the capture page on localhost so both vendors' domains are third-party exactly as they would be on a real customer site:

browserapi.kaidn.iofpnpmcdn.netapi.fpjs.io
Mullvad (uBlock)200⛔ NetworkError⛔ NetworkError
Camoufox (uBlock)200⛔ NetworkError⛔ NetworkError
Brave (Shields)200⛔ Failed to fetch⛔ Failed to fetch

This is not the fingerprinting being defeated. The fingerprinting never runs.

And it cannot be repaired afterwards. Their richer signals need a server-side lookup keyed on the event_id the agent produces. No agent run means no event_id, so a blocked session is not delayed data. It is missing data, permanently.

The honest reading of our own 200s: we are not on those lists because nobody has added us. Zero matches for kaidn across all 164,000 rules. Any fingerprinting endpoint gets listed once it is worth listing, and ours will be. Note also the $third-party scope on their rule. Serving the script from your own domain is the known way out, and it is available to them exactly as much as to us. This is obscurity, not architecture. A head start, not a moat.

Finding 2: we scored honest privacy users as anti-detect browsers, at high confidence#

This is the worst result in the study and it is ours.

Tor Browser and Mullvad Browser, real apps, honest users, scored 51 across all 15 runs. From the table above, 51 is noiseInjected plus fontStandardized. Both are 30-weight signals, so two strong signals agreed and the collector reported high confidence. It was confidently wrong.

Both reasons were wrong, each for its own reason:

  • fontStandardized fired because Tor and Mullvad ship a 150-entry font whitelist that includes Helvetica and leaves out the macOS system face. Gecko's macOS default for sans-serif is Helvetica, so system-ui falls back onto it and three generic font families end up identical. Our rule read that as an anti-detect profile. It is a privacy browser doing exactly what a privacy browser is supposed to do.
  • noiseInjected fired because privacy.resistFingerprinting randomises canvas readback, which is the entire point of the setting. Tor, Mullvad and LibreWolf all turn it on by default.

For a rewards operator, raising the risk on a Tor session may well be the right verdict. But the reason attached to it claimed an anti-detect browser was actively randomising its fingerprint to evade detection, which is a different accusation about a different person.

Explainability is the product. A right answer with wrong evidence is not a right answer.

What we changed. We recomputed the rule over the full raw corpus, 85 honest sessions across 18 configurations plus 10 Camoufox sessions, and replaced the naive version with a narrower rule that requires the specific system-ui behaviour only the anti-detect tool shows:

shipped rulereplacement
fires on honest sessions (n=85)150
catches Camoufox (n=10)710

Strictly better on both counts. Shipped and deployed, then re-measured in the real browsers rather than assumed: Tor and Mullvad both fell from 51 to 30, with the font signal silent 3 of 3 each.

The remaining 30 is still a false positive and it is still open. It is resistFingerprinting tripping noiseInjected: one strong signal, medium confidence. Brave is already suppressed in that detector. This setting needs either the same treatment or its own reason code, and until it gets one, an honest Tor user still collects a signal that says something untrue about them.

Finding 3: where both engines actually ran, Fingerprint was better#

Three sessions had both engines running. On the one that matters, they beat us.

On the honest Tor session they returned suspect 21 against our 51. Their threshold is not published, so we cannot say 21 was a pass. What we can say is that we crossed our own review line on an honest user, and they scored the same session at less than half our number.

The honest Chrome and Brave rows are ties, and the Chrome 15 is our automation signal correctly spotting a headless browser, not a false positive.

There is a fourth data point from earlier work that belongs here, because it is the most important one. On a Multilogin profile lying about its operating system, Fingerprint's own explicit os_mismatch method returned false, and their machine-learning model returned 0.9999. Their rule missed it. Their model caught it, at near-certainty.

Our engine is explicit rules of exactly the kind that returned false. Their anti-detect detection is a model trained over billions of labelled sessions. We cannot replicate that data, and pretending the gap is about cleverness rather than corpus size would be dishonest. What we can do is start building our own labelled corpus, which is why this rig exists at all.

Finding 4: spoofing your own OS beat both engines#

The strongest anti-detect setup we tested is the least exotic one. Camoufox pretending to be macOS while actually running on macOS.

There is no operating system to contradict, so every OS-truth signal has nothing to compare against.

We scored 0 in 6 of 9 runs. Fingerprint had no verdict, because their script was blocked. Neither engine caught it. Our new font rule did fire on 9 of 9, and it is the only thing in either engine that saw that session, but it was not shipped into a verdict at the time.

We are saying this plainly because the blocked-script result would otherwise suggest we came out of this study ahead. An anti-detect profile that does not lie about its OS defeats the current state of the art on both sides.

What this means if you are buying fraud detection#

Any vendor whose detection runs in the browser has the dependency described at the top of this post. Ask them:

  • How often does your script fail to load, and how do those sessions show up in my dashboard? Do they appear as rows scored unknown, or do they simply never arrive?
  • Is a failed load itself a signal? A fingerprinting script that will not load is unusual, and that is information even when the fingerprint is not. Most stacks throw it away.
  • When a signal fires, what evidence comes with it? Our Tor result is the argument for this one. A score you cannot audit is a score you cannot correct.
  • Which findings are rules and which are models? Both are legitimate. They fail differently, and you should know which failure you have bought.

Limitations, stated properly#

  • The consumer privacy browser finding rests on n=3. Of the 21 blocked sessions, 18 were Camoufox, an automation browser, and only 3 were Mullvad, a browser real people install. Three is below the n≥5 bar we hold ourselves to before quoting a rate. Treat the Mullvad result as a strong signal that needs replicating, not as a published rate.
  • The Tor agent ran 2 of 3, and the third failure was our own 8 second deadline on a cold load over the Tor network, not a block. We count it as a rig artefact, not a finding.
  • Brave is the useful control. It ships aggressive tracker blocking and Fingerprint runs fine in it, scoring an honest Brave session at 0. The blocking is specific to browsers that ship uBlock Origin with EasyPrivacy enabled, not to privacy browsers as a category.
  • One Mac for the head-to-head, two machines in the corpus, one country throughout. Every row in the results table came off the same Mac. The second machine is an honest Windows PC, and it is load-bearing rather than decorative: it is what demoted nine "it says Windows" features to false-positive risks instead of signals, and it is the box the MoreLogin profile claiming macOS was caught running on. Two computers is not a population, and network differences elsewhere could change how the script loads.
  • The positive class is one tool. Every Camoufox catch is Camoufox. A rule fitted to one tool on one machine is a lead, not a general detector, and we say so in the code.
  • Corpus sizes: 103 honest sessions, 28 anti-detect sessions, 24 privacy and anti-detect sessions in the head-to-head above.

Reproduce the blocklist half yourself#

One command, no access to our rig needed:

Terminal
curl -s https://easylist.to/easylist/easyprivacy.txt \
       https://easylist.to/easylist/easylist.txt \
  | grep -nE 'fpjs\.io|fpnpmcdn\.net'

You should get two rules back, one from each list. Swap in any other vendor's domains to check where they stand. The lists change, so line numbers will drift.


If you want the engine that produced our numbers above, it is @kaidn/fp on npm, and every signal it reports comes back with the evidence attached. The scoring API docs show what the server does with it.

Frequently asked questions

Do ad blockers stop device fingerprinting?

They stop a third-party script from loading, which for a hosted fingerprinting vendor amounts to the same thing. In this head-to-head, Fingerprint's script did not run on 21 of 24 privacy and anti-detect browser sessions, because both EasyPrivacy and EasyList carry rules for its domains and the browsers involved ship uBlock Origin by default. A blocked script does not return a weak verdict. It returns none at all.

Which filter rules block Fingerprint, specifically?

Two, one from each list. EasyPrivacy carries ||fpjs.io^$third-party, which blocks their API domain whenever it loads from a site other than the one being visited, and for a vendor embedded on customer sites that is every time. EasyList carries ||fpnpmcdn.net^, which blocks their CDN outright. You can verify both in one curl command against the published lists.

Is Kaidn better at detecting anti-detect browsers than Fingerprint?

No, and the honest reading of these results is more mixed than that. Where both engines actually ran, Fingerprint was better. Our advantage in this test was availability rather than accuracy: a first-party collector is not on the filter lists, so it returned a verdict where a third-party script returned nothing. Availability is a real advantage and it is not the same as being more accurate.

Can either engine catch an anti-detect profile that does not lie about its OS?

Neither did. Camoufox pretending to be macOS while running on macOS was missed by Kaidn in six of nine runs, and produced no verdict at all from Fingerprint in all nine, because the script was blocked. A profile that does not lie about its operating system removes most of the contradiction both engines look for. That is the state of the art, not a gap in one product.

What went wrong with the font signal?

It fired at high confidence on honest privacy users. Tor Browser and Mullvad Browser both standardise fonts deliberately, as an anti-fingerprinting measure, and our rule read that as an anti-detect profile: 51 out of 100 on real people doing nothing wrong. The fix narrowed the rule to the specific behaviour only the anti-detect tool shows, and it is deployed.

How big was the sample, and what does that limit?

103 honest sessions, 28 anti-detect sessions, 24 privacy and anti-detect sessions, in one country. Every session in the head-to-head table was captured on one Mac; the wider corpus behind it also includes an honest Windows PC, which is what the cross-machine false-positive checks are run against. That is enough to establish that the blocking happens, and that spoofing your own OS defeats both engines. Two computers is not enough to quote rates as population statistics. Sessions are counted one by one and never averaged, because anti-detect tools randomise on every launch.

browser fingerprintingmeasurementprivacy browsersanti-detect browsers

Score your own traffic

10,000 events a month on the free tier, no card. One POST to /v1/score and you get a verdict with the evidence behind it.

Read next