HyperMix: open-source detection of engineered biosignatures in remote hyperspectral imagery

$205
Pledged
9%
Funded
$2,500
Goal
27
Days Left
  • $205
    pledged
  • 9%
    funded
  • 27
    days left

About This Project

Long-range biosensing is now possible: engineered microbes can be read from ~90 m by hyperspectral cameras. But the reporter signal is weak, buried in natural backgrounds (soil, leaves, water), and few molecules have measured spectra. HyperMix is an open-source, physics-informed deep-learning toolkit that jointly detects and unmixes engineered reporters in noisy remote cubes — with an open simulator, dataset and benchmark. MIT-licensed, built on clinical OCT methods.

Ask the Scientists

Join The Discussion

What is the context of this research?

Imagine spraying harmless engineered bacteria on a field and reading their chemistry from the sky, getting warned about pollution in the soil, missing nutrients in a crop, or cracks forming in concrete, all without collecting a single sample by hand.

That future just got closer. In a 2026 Nature Biotechnology study, scientists engineered bacteria that emit a specific color, then spotted it from 90 meters away with a hyperspectral camera, a sensor that sees hundreds of shades of light instead of just red, green, and blue.

The catch: the signal is very faint and hides inside the colors of soil, leaves, and water, which change everywhere you look. Today's standard detection methods work well only when the signal is clean and predictable, and it rarely is. My hypothesis is that a smarter, physics-guided algorithm can find these faint signals where the old methods fail. This project builds an open, free tool to test exactly that.


What is the significance of this project?

Better detection algorithms are pure software, so the return per dollar is high: every cheaper camera and every dataset the field funds becomes more useful when software can pull signal from clutter. This work is field-building infrastructure: an open toolkit, an open scene simulator, and an open spectral dataset that others can build on. It needs no wet lab, no new hardware, and no engineered-organism release, so it carries no regulatory friction and ships fast.

Its concrete output is the evidence itself: a public benchmark that quantifies, with confidence intervals, how well each detection method recovers faint engineered reporters against real backgrounds at controlled signal-to-noise, alongside an open dataset and a shared leaderboard. That gives hyperspectral biology a common yardstick it lacks today, so future sensors, reporters, and algorithms can be compared fairly and reproducibly rather than claim by claim.

What are the goals of the project?

This project will deliver an open, reproducible toolkit that detects engineered hyperspectral reporters against unknown backgrounds in noisy remote images. First, I will build a physics-based simulator that composites reporter signatures over real backgrounds with realistic illumination, atmosphere, and noise, producing images with known ground-truth labels. I will then reproduce matched-filter and unmixing baselines as the performance floor, and train a physics-informed network that jointly detects and unmixes, handles unknown backgrounds, and outputs detection maps with calibrated uncertainty. Finally, I will release everything under the MIT open-source license: a python library, Colab notebooks, an open spectral dataset, a public benchmark, and a standard scene format. Success means the benchmark measures whether and where the learned detector improves on classical baselines, and the pipeline runs end-to-end on a real, open hyperspectral cube.

Budget

Please wait...

This minimum target funds the core, runnable deliverable on free GPU tiers (Kaggle/Colab). Cloud GPU compute covers training the physics-informed detection model beyond what free tiers allow. Reference data curation pays for the measured spectra and validation cubes needed to build and benchmark the method. Open hosting, docs and DOI archiving (Zenodo/Hugging Face) make the library, dataset and benchmark permanently public and citable. The fee line covers Experiment's platform and payment-processing charges so the net reaches the work. Larger ambitions — the full detector, an extended open spectral dataset and a public leaderboard — are framed as stretch goals to be crowdfunded on top of this minimum.

Endorsed by

João Victor is one of the most brilliant Data Scientists that I had the opportunity to work with. I'm completely sure he can delivery this solution and impact all the future research in this field. His previous experience in image detection and processing is a perfect background for this project.
João Victor's work is built around extracting clinically meaningful signal from noisy, multi-vendor imaging data with little ground truth, precisely the core challenge here, just with a hyperspectral cube instead of an OCT scan. I'd bet on him getting a real, reusable toolkit out the other end of this.

Project Timeline

This project runs for five months, within the program's one-year limit. In month one I assemble the open datasets and build the physics-based simulator with classical baselines. In months two and three I train and validate the physics-informed detector, adding unknown-background handling and uncertainty maps. In months four and five I package the library and Colab notebooks, publish the open dataset, and release the benchmark, scene format, and a DOI archive.


Jul 30, 2026

Project Launched

Oct 31, 2026

Open datasets gathered; physics-based scene simulator and classical baselines built

Dec 31, 2026

Physics-informed detector trained and validated on synthetic and real cubes

Feb 28, 2027

Public release: pip library, Colab notebooks, open dataset, benchmark and DOI

Meet the Team

Joao Victor Dias
Joao Victor Dias
Statistician

Joao Victor Dias

Joao Victor Dias is a statistician working at the intersection of medicine and applied machine learning, focused on medical and ophthalmic imaging. He builds deep-learning pipelines for image reconstruction and segmentation, including cross-vendor retinal OCT segmentation for a MICCAI 2026 challenge track, where the core problem is recovering a faint, clinically meaningful signal from noisy, blurred, multi-channel data acquired on heterogeneous devices, often with little or no ground truth. That is exactly the problem hyperspectral biology faces at a distance: detecting a weak engineered reporter against an unknown natural background, in low-SNR cubes from cheap or far-away sensors, with only a handful of measured reference spectra. He has shipped reproducible models on free and low-cost GPUs (Kaggle, Colab), which keeps this project affordable and easy for others to reuse. This is an individual, non-commercial project: all code, the scene simulator, the spectral dataset and the benchmark will be released under the MIT license and archived with a DOI, so any researcher, university lab, community lab, or dedicated amateur, can build on them.


Project Backers

  • 2Backers
  • 9%Funded
  • $205Total Donations
  • $102.50Average Donation
Please wait...

See Your Scientific Impact

You can help a unique discovery by joining 2 other backers.
Fund This Project