About This Project
Hyperspectral biology depends on understanding how molecules absorb light, yet no spectral atlas of nature's metabolites exists. Using the NCI natural products library and automated LC-DAD-MS/MS, we will pair each compound's 250-800 nm absorbance spectrum with its molecular structure. A pilot of 10-20 extracts will validate the pipeline, quantify spectral yield per extract, and openly release thousands of spectra, creating a resource for predicting spectra from molecular structure.
Ask the Scientists
Join The DiscussionWhat is the context of this research?
Hyperspectral biology reads the light that molecules absorb, but the reference data it depends on barely exists. Absorbance spectra are recorded constantly, most LC runs pass eluent through a diode-array detector, yet they are used to confirm a compound's identity and then discarded. What survives is partial, non-standardized, scattered, proprietary, and rarely linked to structure. Existing libraries (PhotochemCAD and commercial UV-Vis collections) hold of thousands of spectra, but few are for natural metabolites, contain full spectra, are structure-linked, uniformly acquired and hyperspectral-biology-ready.
The bottleneck is not instrumentation. Inline Liquid Chromatography with diode-array and mass detection is standard, off-the-shelf equipment. Nobody has run it systematically across chemical space. The UCSB Mammalian Synthetic Biology Foundry holds one of the few complete copies of the NCI natural-products library (about 230,000 extracts) and the automation to change that.
What is the significance of this project?
A natural metabolite absorbance atlas would let us train models that can predict, from structure alone, what wavelengths any molecule biology can make will absorb. That will enable better molecular identification in hyperspectral imaging and turn hyperspectral reporter engineering from a search, hunting for a useful chromophore one lucky find at a time, into design at a chosen wavelength (Chemla et al., Nat Biotech 2025). It also supplies the reference spectra needed to deconvolute complex biological spectra into their constituent molecules.
The dataset, not the device, is the contribution of our proposal. Natural metabolites, terpenes, polyketides, alkaloids, pigments, occupy chromophore space that dye-derived training sets and TD-DFT benchmarks largely miss, and they are measured in real solvent environments rather than computed in vacuum. Because the atlas will be openly released and uniformly acquired, it becomes shared infrastructure for the emergence of hyperspectral biology.
What are the goals of the project?
1) Develop the pipeline. Run authentic standards through inline LC-DAD-MS/MS. Confirm wavelength and photometric accuracy against reference standards, set peak-purity and absorbance thresholds, and characterize the spectral shift introduced by gradient elution so it can be corrected rather than ignored.
2) Validate it on the NCI collection. We will run 10-20 extracts, chosen to span different organisms and chemistries rather than picked at random. The key result is how many spectra per extract we can actually link to a known structure and with what accuracy.
3) Establish the database. Publish every spectrum openly with its full acquisition metadata, in a permanent repository. This becomes the seed dataset for a large scaleup effort to measure hundreds of thousands of extracts.
Budget
Because the HPLC, diode-array detector, and mass spectrometer, and person-hours, are cost-shared in kind by the UCSB Mammalian Synthetic Biology Foundry, ExFab, and other UCSB facilities, the entire microgrant will fund the materials and methods needed to establish this workflow and capture the first set of molecules. Reference standards set the validation ground truth and chromatography and mass-spec consumables cover the pilot runs.
Project Timeline
Milestones (6 mo). (1) M0-2: Validate pipeline on authentic standards spiked into extract matrix; 90% reproduce reference λmax within 2 nm and match band shape. (2) M2-4: Pilot 10-20 NCI extracts spanning organisms and chemistries; report structure-linked spectra per extract. (3) M4-6: Release all spectra openly, with metadata, in a permanent repository. (4) M6: Convert yield into a cost model and a go/no-go for mapping the full collection.
Dec 31, 2026
Validate pipeline on authentic standards
Dec 31, 2026
Release all spectra openly, with metadata, in a permanent repository
Dec 31, 2026
Pilot 10-20 NCI extracts spanning organisms and chemistries
Meet the Team
Affiliates
Maxwell Wilson
Dr. Max Wilson is a molecular biologist and engineer who develops optogenetic tools to control cellular behavior with light, integrating synthetic biology, robotics, and artificial intelligence to accelerate discovery for aging and age-related disease. He directs the Robotic Synthetic Biology Foundry at UC Santa Barbara, where he automates the scientific process through AI-guided experimental design and high-throughput robotic screening.
Wilson earned his Ph.D. in Molecular Biology from Princeton University, combining experimental biology with applied mathematics to model how cells allocate metabolic resources under stress. He showed that the spatial organization of enzymes controls metabolic flux and that engineered protein assemblies can reprogram cellular behavior. He also built one of the first imaging-based AI platforms for antibiotic discovery.
As a postdoctoral researcher at Princeton, Wilson turned to optogenetics, recognizing light as an ideal tool to manipulate cells with precision. He engineered the first optogenetic systems to control key signaling pathways with light, establishing a foundation for dynamic, light-based cellular engineering. His lab now extends these tools to decode how cells respond to stress and reveal strategies that make them more resilient, a principle central to both antiviral defense and longevity.
At UC Santa Barbara, Wilson studies how cells process information and make fate decisions. When the COVID-19 pandemic struck, he redirected his team toward an urgent public health need, co-developing UCSB's CRISPR-based diagnostic platform (CREST) and helping build regional testing infrastructure.
Wilson's work bridges academia and biotechnology, theory and experiment. By developing optogenetic technologies, AI-driven discovery platforms, and robotic systems within an academic setting, he expands what is experimentally possible while keeping discovery grounded in the search for therapies that restore resilience in aging cells.
Lab Notes
Nothing posted yet.
Additional Information
Approach:
Using the NCI extract library and the automated instruments at the UCSB Mammalian Synthetic Biology Foundry, we will build a high-throughput workflow that pulls each extract apart and reads its components in three steps:
- Separate. Run one extract through liquid chromatography; the column resolves its molecules in time, eluting them one after another.
- Read continuously. An inline diode-array detector records a full UV-Vis absorbance spectrum (250-800 nm) up to tens of times per second, capturing a clean spectrum for each compound as it passes.
- Identify. An inline mass spectrometer records accurate mass and an MS/MS fragmentation fingerprint for each peak. High-confidence identities come from MS library matching (mzCloud, GNPS); the remainder are retained as accurate-mass features at a labeled, lower-confidence tier, identifiable later by NMR if needed.
The NCI library. The Foundry holds one of the few complete copies of the NCI Program for Natural Product Discovery prefractionated library: over 230,000 extract and 326,000 fractions generated from a repository of natural-product extracts assembled since 1986 from more than 25 countries, spanning thousands of plant, marine, fungal, and algal genera, and among the largest and most chemically diverse natural-product collections in existence (dctd.cancer.gov/programs/dtp/organization/npb/npnpd). Equally important, the library sits alongside the automation that ran a 370,830-compound iPSC screen (Wong … and Wilson, et al., Cell 2025).
Anticipated challenges:
Co-elution. Not every LC peak is a single compound. Each extract will be run on several columns with different chemistries, producing different elution profiles. Where compounds still overlap, finely-sampled inline spectra plus curve resolution (MCR-ALS) can often deconvolute them, with the mass spectrometer confirming purity. We note MCR-ALS degrades where congeners share both chromophore and elution behavior; such peaks are flagged, not forced.
Spectral fidelity. We calibrate wavelength against holmium oxide and photometric accuracy against a dichromate standard, rechecking each run. A spectral peak-purity threshold and an upper-absorbance cutoff reject over-concentrated, distorted readings, so only high-confidence measurements enter the atlas.
Solvent environment. A molecule's spectrum depends on its solvent, and under gradient elution every compound elutes into a different mobile-phase composition. Storing metadata (pH, solvent, temperature at elution) makes records auditable but not comparable. We will therefore measure the shift empirically, a standards panel run across the gradient range, so elution conditions can be modeled as a covariate rather than treated as noise.
Detection mismatch. DAD detects at low-ng on-column while MS reaches pg–fg, so we do not expect a spectrum for every MS feature. The atlas deliberately targets the abundant, strongly absorbing subset that carries the chromophores hyperspectral biology cares about.
Annotation rate. Confident structural annotation in untargeted natural-product metabolomics is typically well under 10% of features. We report structure-linked spectra, not total spectra, as our headline metric.
Spectral range. 250–800 nm captures most biologically relevant absorbance and is the off-the-shelf DAD range. NIR-absorbing metabolites are rare but are the highest-value targets for field- and remote-readable reporters; once the methodology is established, an extended 250-1000 nm detector is a straightforward upgrade. The atlas's near-term value for NIR is predictive, learning the structure–spectrum map well enough to design into that window, rather than discovery.
Project Backers
- 0Backers
- 0%Funded
- $0Total Donations
- $0Average Donation

