50 years of cassava diversity: what went into the bank, and what came out

Two new papers in Plants, published two months apart, tell the story of CIAT’s cassava genebank 1 from opposite ends: how the collection was assembled and conserved, and what breeders have actually done with it. Read together, they amount to a remarkably candid 50-year audit of a slow-motion agricultural asset. I’ll just give you the main beats here. It’s really worth reading both papers in full.

Part I is the origin story. Botanist Victor Manuel Patiño’s 1969–70 expeditions alone brought in roughly a third of today’s 5,000 or so cassava landraces, transported as stem cuttings across Colombia, Ecuador, Venezuela and beyond, often in difficult conditions and under changing quarantine rules.

The collection has its blind spots: Brazil and the Guianas remain poorly represented 2, passport information can be patchy, and the overwhelmingly male composition of historical collecting teams probably meant that some of the varietal knowledge held by women farmers went unrecorded.

Conservation itself has been a moving target. It has evolved from field genebanks to in vitro slow-growth tissue culture storage after a frogskin-disease scare forced the field collection to close in 2003, and now cryopreservation at CIAT’s Future Seeds facility. Meanwhile, DNA fingerprinting keeps revealing an awkward truth familiar to genebank curators everywhere: the names people give varieties and the genetic identities of the material do not always agree. A lot of effort over the years has also gone into testing for pathogens to ensure that distribution is safe.

Part II asks what all this diversity is actually for? The answer has also changed considerably over five decades. Early researchers chased traits such as high protein and low cyanogenic content before turning towards yield and starch percentage in the Green Revolution era.

The payoffs have been tangible: landraces have contributed traits such as resistance to pests and diseases, adaptation to acid soils and highland cold, and quality traits for fresh-market cooking versus industrial starch. The route from accession to released variety is rarely direct: a landrace gets screened for a trait, crossed, then recombined and selected over several more generations before anything reaches a farmer’s field. The results include varieties like Nataima-3, bred for whitefly resistance thanks to an Ecuadorian landrace. At the same time, the authors argue that cassava breeding now needs to move beyond broad phenotypic selection towards more systematic use of inbred lines, because the crop’s high heterozygosity makes the introduction of specific traits particularly difficult.

The two papers therefore tell a big story about genebanks. Collecting diversity is only the beginning; its value is realized decades later, when a breeder encounters a problem that nobody could have anticipated when the stuff was first collected. The cassava collection is a long-term portfolio of biological options, one whose contents have taken half a century to assemble, whose inventory is still being corrected, and whose most valuable assets may be the ones that have not yet been used.

A new project is looking to genetically engineer cassava to photosynthesize more efficiently at higher temperatures. I do wonder whether someone has already checked whether any of those 5,000 landraces might help with that. Maybe they can’t, but it’s worth having a look. And you never know, something else of interest might jump out. That’s the beauty of large international collections of crop diversity such as CIAT’s.

Conserving the tangle of grapevines

I think we may have already pointed to Conservation gap analysis for wild grapevines (Vitis L.) of the Americas, the latest in a series of papers by our friend Colin Khoury and a rotating assortment of colleagues on the conservation status of the crop wild relatives of the Americas, genepool by genepool. The authors compiled occurrence records for 38 wild American grapevine taxa, and used fancy GIS to infer the overall distribution and environmental niche of each. They then assessed the degree of representation of each taxon in genebanks and protected areas, and hence any remaining conservation gaps. Here’s the headline finding:

We categorize 25 of 38 of the taxa as urgent priority and 10 as high priority for improving ex situ conservation representation. Three taxa are assessed as urgent and 29 as high priority for enhancing in situ conservation. Further action, with emphasis on conservation gap hotspots, is needed to more comprehensively conserve wild Vitis native to the Americas.

Which is pretty clear.

Or is it?

What if “taxa” are perhaps not always the best units of conservation to use in assessing conservation efforts?

That may in fact be one of the implications of a paper that came out just a few weeks after that of Colin and friends: The dynamics of introgression and parallel adaptation across North American Vitis species.

These authors show that introgression and hybridization are pervasive and evolutionarily important across North American Vitis, based on genomic analysis of 639 accessions representing 48 species. About 14% of the average genome shows evidence of introgression, particularly associated with areas where species come into contact. Some taxa usually regarded as hybrid species are in fact better understood as ever-changing hybrid swarms, rather than distinct evolutionary lineages. Most importantly, the authors find that introgressed genetic variants have repeatedly contributed to adaptation in different species. The paper therefore portrays Vitis diversity as a reticulate network — or tangle — of species, populations and gene flow, rather than a set of discrete species.

This has important implications for conservation: hybrid zones and admixed populations may be really significant reservoirs of adaptive diversity. The framework of the first paper might potentially underestimate the conservation importance of regions where these occur, if they contain substantial genetic variation but aren’t well represented by the taxonomic units used in the gap analysis. For example, it might happen that two neighbouring species are reasonably well represented ex situ, but not from the specific regions where they hybridize and introgression occurs.

This suggests a useful next, synthetic step: take the geographic gaps from the first paper and overlay them with the evidence for introgression and gene flow networks from the second. The resulting map could identify not just under-collected species, but under-collected (or under-conserved in situ) evolutionary processes and genetic mixtures. That could be valuable for designing the next Vitis collecting mission.

Towards a digital workflow for forest restoration

I missed the “From Seeds to Success: Digital Tools for Planning and Managing Forest Restoration” webinar a few weeks ago, jointly organized by the Alliance of Bioversity International and CIAT, the Millennium Seed Bank at the Royal Botanic Gardens, Kew, and the Forest Restoration Research Unit (FORRU), Chiang Mai University. Too busy moving to another continent. But fortunately there are now a handy summary and even a recording online. And there will be a re-run.

To remind everyone what the webinar was about:

The programme featured live demonstrations of five innovative digital tools, developed to support restoration planning, seed sourcing and project management, followed by an interactive discussion, during which participants explored their applications, strengths and future development.

With these five tools, therefore, we have the beginnings of a digital seed-to-restoration pipeline. Which is very exciting to me.

Let’s go through it step by step, and tool by tool.

1. Define the restoration objective and choose species: Diversity for Restoration (D4R)

Picture yourself looking at a site you want to restore. This tool answers the question: What should we plant here?

D4R is the front-end decision-support tool. It combines information on:

  • restoration objectives
  • environmental conditions
  • species distributions
  • functional traits
  • seed zones
  • future climate scenarios
  • to recommend suitable species and seed sources.

    2. Find the seed: SeedPOD

    Then we need to know: Where can we actually obtain suitable seed?

    Once D4R has identified desirable species and potentially appropriate provenances, SeedPOD steps up as the seed-sourcing layer, providing information from different organizations worldwide on the storage, availability, and germination of their seed collections. It is scheduled for launch in September 2026.

    3. Decide how to handle and store the seed: Wyse–Dickie Seed Storage-Behaviour Predictor

    But you won’t necessarily be able to get such information on all the seeds you might need. So how should we handle those seeds?

    This tool predicts whether seeds are likely to be:

  • orthodox — tolerant of drying and suitable for conventional storage
  • recalcitrant — sensitive to drying and therefore requiring different handling
  • using species characteristics, taxonomy and environmental variables. That means you can work out how to process and store them.

    4. Germinate and produce seedlings: Germination Experiment Assistant (GEA)

    Having the seeds is great, but how do we turn them into viable seedlings?

    Where there is little or no published germination information, you will need to generate it yourself. GEA uses experimental data and predictive modelling to produce species-specific germination protocols, including practical procedures for nurseries.

    5. Track the material through the restoration operation: MyFarmTrees

    Finally, once the work has started, you’ll need to monitor how it’s going: What happened to each seed lot, seedling and planting?

    MyFarmTrees is the operational backbone of the five tools. It tracks restoration activities from seed collection to nursery production to field establishment using QR-coded seed lots and mobile technology. This gives you traceability. It also creates the possibility of connecting restoration activities to monitoring and ultimately payment-for-ecosystem-services systems and the like.

    Is anything missing?

    Having these five tools is great, but they don’t quite constitute the entire restoration workflow. They cover the biological-material pipeline extremely well: selection → sourcing → handling → propagation → deployment → monitoring. But the webinar participants did identify restoration approach selection, as well as planning collecting programmes, as gaps in the current digital toolkit. Along with, inevitably, seamless integration, or at least interoperability, of all the different tools.

    So there should maybe be an initial stage to the above sequence:

    0. Assess the site and restoration strategy

    Because before unleashing D4R, you do need to know a bunch of stuff, for example:

  • What is the state of the site?
  • Is active planting actually necessary?
  • What natural regeneration is occurring?
  • What ecosystem are you trying to recover?
  • What functions are missing?
  • What species are already present?
  • What are the local threats?
  • What restoration approach is appropriate?
  • And there should also maybe be a stage parallel to what I labelled 2. Find the seed above. Call it…:

    2b. Collect the seed

    Because if SeedPOD says suitable seed is available, all well and good. But if it isn’t, it would be nice to have a “seed collection planner” to generate a collecting programme for you, factoring in the distribution, ecology, mating system, and phenology of the target species, the accessibility of potential collecting sites, the budget available, you get the idea.

    So maybe eventually the complete architecture of the system will be something along these lines:

  • Site assessment: What needs restoring? Gap
  • Restoration design: What restoration approach is appropriate? Gap
  • Species & provenance selection: What should we plant? D4R
  • Seed sourcing: Where can we obtain seeds? SeedPOD/Gap (collecting)
  • Seed handling/storage: How should we handle the seeds? Wyse–Dickie
  • Germination/propagation: How do we produce seedlings? GEA
  • Traceability & deployment: Where did every seed/seedling go, and what happened to them? MyFarmTrees
  • Monitoring: Did restoration succeed? Gap/MyFarmTrees
  • Interested in seeing where all this goes, as I am? Start by registering for the re-run of the webinar on 30 September.

    The revenge of the sweetpotato

    Speaking of digital imagery and its uses in genebanks, get a load of the recently published catalogues of the Cuban sweet potato collection at the Research Institute of Tropical Roots and Tuber Crops (INIVIT), and of Peruvian cacaos. Beautiful. And do yourself a favour and don’t skip the foreword, preface and introduction to the Cuban volume. I found them very eloquent — and moving. Maybe a touch overwrought, admittedly, but it does take some gumption to say, of sweetpotato, in a genebank catalogue of all places, that: “Today, history grants it its revenge.” I just hope it’s true.

    A picture is worth a thousand descriptors

    For decades, germplasm characterization has relied on people looking at plants, seeds and fruits and recording what they see; first on paper forms, more recently admittedly on tablets and the like. I’ve done that myself, and let me tell you, recording the colour of taro stems on bits of damp paper in the middle of a forest clearing in Vanuatu is no fun.

    Lately, thankfully, the camera has been taking over.

    A recent overview of new tools for plant genebanks highlights digital photography as a way of capturing standardized information on the colour, size and morphology of seeds and other plant parts, alongside more sophisticated technologies such as hyperspectral imaging and mobile field sensors. The attraction is obvious: instead of recording a handful of descriptors by eye, images can capture a much richer set of characteristics that can subsequently be measured and analysed at your leisure.

    The potential is particularly striking for fruit crops. A new study of heritage apples used controlled multi-view imaging to characterize about 350 accessions over three years and two locations. Fancy maths achieved 95% accuracy in distinguishing 38 cultivars, rising to 99% when images of three fruits were combined.

    And the technology does not necessarily require sophisticated equipment. In a recent demonstration with beans with complex colour patterns, Miguel Angel Acosta Chinchilla used ordinary photographs, image pre-processing and clustering algorithms to extract dominant palettes and colour distributions.

    We’re moving from a limited number of human-defined descriptors to infinitely explorable machine-readable phenotypes. An image can preserve information that nobody thought to score at the time, and algorithms can return to it later to measure traits that were not originally part of the characterization protocol. For genebanks, that could be transformative. A photograph taken today may become a source of data for questions that don’t actually arise until tomorrow.

    I wish I had a digital camera with me in that taro patch twenty-odd years ago. Goodness knows what people could be finding in those photos now.