Our team at Imagetwin built a novel duplicate detection technology and, as a first step, used it to develop a new method for finding duplicated Western blot images. The method works for all duplicate types: Western blot duplicates within a single image, between images in the same paper, and between images across different papers. The method is robust to cropping, resizing, color changes, contrast shifts, rotation, and flipping. Most importantly, it reliably distinguishes between Western blots that look similar by chance and genuine duplicates, virtually eliminating false positives.
Below are a few real examples of the manipulations this method is built to catch, drawn from our testing:
Evaluation Results
To measure performance, we built a test set from real-world Western blot duplicates reported on PubPeer, mixed in with a large pool of distractor images that are not duplicates. We didn’t use any of this data for training, so the results reflect unbiased, held-out performance.
The distractors weren’t picked at random; we used hard mining to select images that already looked similar to each other. A random pool wouldn’t generate enough near-miss cases to properly stress-test false alarms, so this is a deliberately tough test.
In total, we evaluated the new method against 597 confirmed duplicates from PubPeer and 5,860 hard-mined distractor images. The result: 95.8% of real duplicates caught, with a false alarm rate of just 0.11%.
Even the small number of false alarms were extremely close calls. In our own review, it took real effort to spot the subtle differences that made them non-duplicates in the first place.
What's Coming Next
Over the coming weeks, we’re rolling out the same underlying technology to FACS plots, microscopy images, light photography, and XRD data, aiming to achieve similarly strong detection rates across those image types.
Frequently asked questions
Is this the most accurate Western blot duplicate detector available?
Yes, and here is why: on a held-out test set of 597 confirmed PubPeer duplicates and 5,860 hard-mined distractor images, our new method catches 95.8% of real duplicates with a 0.11% false alarm rate, it is the strongest result we’ve measured for this specific problem.
What kinds of Western blot duplication does this detect?
The method catches duplicates within a single image, between images in the same paper, and across different papers, even when the duplicated region has been cropped, resized, recolored, contrast-adjusted, rotated, or flipped.
Does this replace manual review by image integrity specialists?
No, it flags likely duplicates for a human reviewer to confirm, the same way our other detection methods work. It’s built to catch what’s easy to miss at scale, and to support human judgement.