{"data":[{"id":"10.5061/dryad.66t1g1kcz","type":"dois","attributes":{"doi":"10.5061/dryad.66t1g1kcz","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["Technical University of Munich"],"name":"Neumann, Astrid E.","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0002-1736-3946"}]},{"nameType":"Personal","affiliation":["Technical University of Munich"],"name":"Casanelles-Abella, Joan","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Technical University of Munich"],"name":"Conitz, Felix","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Museum für Naturkunde"],"name":"Karlebowski, Susan","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Université Libre de Bruxelles"],"name":"Roberts, Stuart","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Ludwig-Maximilians-Universität München"],"name":"Schmack, Julia M.","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Museum für Naturkunde"],"name":"Sturm, Ulrike","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Technical University of Munich","Cornell University"],"name":"Sexton, Aaron N.","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Technical University of Munich"],"name":"Egerer, Monika","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0002-3304-0725"}]}],"titles":[{"title":"Data and code from: Food resources are more important than nesting resources for explaining wild bee diversity in urban gardens"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Urbanization","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Functional groups","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Community structure","subjectScheme":"PLOS Subject Area Thesaurus"},{"subject":"Apiformes"},{"subject":"Anthophila"},{"subject":"pollinators"},{"subject":"urban filtering"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Cities","subjectScheme":"PLOS Subject Area Thesaurus"}],"contributors":[{"name":"Technical University of Munich","contributorType":"Sponsor","affiliation":[],"nameIdentifiers":[]}],"dates":[{"date":"2025-08-06T10:11:25Z","dateType":"Created"},{"date":"2026-08-18T11:17:51Z","dateType":"Submitted"},{"date":"2026-08-21T00:00:00Z","dateType":"Issued"},{"date":"2026-08-21T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[{"relationType":"IsCitedBy","relatedIdentifier":"10.1016/j.baae.2024.06.004","relatedIdentifierType":"DOI"}],"relatedItems":[],"sizes":["80199 bytes"],"formats":[],"version":"6","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Wild bees are declining worldwide, with urbanisation as one major driver.\n Urban environments can filter for wild bees with specific functional\n traits and resource requirements. Nevertheless, cities can support many\n wild bee species, especially in urban gardens. Yet, we do not know how\n specific habitat features influence bee species persistence and community\n composition in gardens, which prevents implementing appropriate\n conservation measures. We investigated how local and landscape garden\n features affect the taxonomic and functional diversity of wild bees and\n their community composition, and which garden features associate with wild\n bee functional guilds in urban community gardens. Over two years, we\n surveyed community gardens in Munich and Berlin (Germany), assessed bees\n using flower observations and pan traps, and measured local and landscape\n garden features (food and nesting resources, habitat availability incl.\n urbanisation). We found 168 wild bee species, ~1/3 of Germany's bee\n species. Floral richness was the main driver of bee taxonomic and\n functional diversity. Urbanisation, floral resources and open ground\n within gardens significantly affected bee community composition. Finally,\n we found associations between garden features and bee functional guilds.\n For example, the abundance of parasitic bees increased with floral\n richness, that of ground nesters with open ground and that of\n late-emerging species with urbanisation. Urban gardens are important wild\n bee habitats in cities and spaces for conservation action. Diversifying\n available resources may support high bee diversity, and considering\n species functional traits and species requirements may further improve\n urban gardens and other urban ecosystems as bee habitats."},{"descriptionType":"TechnicalInfo","description":"# Data and code from: Food resources are more important than nesting\n resources for explaining wild bee diversity in urban gardens title:\n \"README\" date: \"2026-08-21\" doi:\n [https://doi.org/10.5061/dryad.66t1g1kcz](https://doi.org/10.5061/dryad.66t1g1kcz) ## General information The following repository includes all necessary information on data and code to replicate the analysis for the manuscript: \"Food resources are more important than nesting resources for explaining wild bee diversity in urban gardens\". The repository Dryad provides the R code structure, data and code. The data for the variable of wild bee mean intertegular distance (ITD) will only be published here after the publication of the European Bee Traits Database by S. Roberts et al. in the future. The variables “deadwood_x_y”, “garden_area” and “impervious_1000” have been published before on Zenodo under a CC BY 4.0 licence in association with the publication of Neumann et al. (2024). These variables therefore need to be downloaded separately from the rest of the data and then correctly merged with the datasets provided on survey level (\"Wildbees_features_survey_subset.csv\") and garden/year level (\"Wildbees_features_garden_year_subset.csv\"). A step-by-step instruction on how to merge these variables with the datasets is provided below the dataset description in this README file. **Corresponding author information** ``` Name: Astrid E. Neumann Orcid: 0000-0002-1736-3946 Affiliation: Urban Productive Ecosystems, Technical University of Munich, Freising, Germany email: astrid.neumann@tum.de ``` **Alternative contact information** ``` Name: Monika Egerer Orcid: 0000-0002-3304-0725 Affiliation: Urban Productive Ecosystems, Technical University of Munich, Freising, Germany email: monika.egerer@tum.de ``` **Associated publication** Neumann, A.E.; Casanelles-Abella, J.; Conitz, F.; Karlebowski, S.; Roberts, S.; Schmack, J.M.; Sturm, U.; Sexton, A.N.; Egerer, M. Food resources are more important than nesting resources for explaining wild bee diversity in urban gardens ## Description of the data and file structure ### Project structure The compressed file \"**WildBeeTraitsProject.zip**\" contains three main directories, that is, \"*Input*\", \"*Scripts*\" and \"*Output*\". Load the data and produce the output: ``` WildBeeTraitsProject ├── Input │ ├── Wildbees_features_garden_year_subset.csv │ ├── Wildbees_features_survey_subset.csv │ ├── Wildbees_traits_taxonomy_excl_ITD.csv ├── Scripts │ ├── 01_GLMM │ ├── 02_RLQ │ ├── 03_NMDS_garden_level │ ├── 03_NMDS_survey_level │ ├── 04_Figures_NMDS_garden_level │ ├── 04_Figures_NMDS_survey_level │ ├── 04_Figures_GLMM │ ├── 04_Figures_traits_overview │ ├── 04_Figures_species_overview │ ├── 05_INEXT ├── Output │ ├── figure_1 │ ├── figure_2 │ ├── figure_3 │ ├── figure_4 (a-f) │ ├── figure_5 │ ├── figure_s3 │ ├── figure_s4 (a, b) │ ├── figure_s5 (a-f) ``` * Directory \"*Scripts*\" contains the R scripts. They are described in the section \"Scripts\" below in this README * Directory \"*Input*\" contains three datasets on (1) wild bee traits, (2) wild bee diversity and garden features on a survey level and (3) a garden/year level * Directory \"*Output*\" is empty and will be filled when running the scripts ### Provided data in this repository #### Wild bee diversity and garden features on a garden/year level *Filename*: Wildbees_features_garden_year_subset.csv *Format*: Semicolon-delimited CSV file *Location*: **This dataset is found under the following path: WildBeeTraitsProject/Input/** *Description*: The following dataset contains the bee species and abundances, wild bee taxonomic and functional diversity metrics and local and landscape garden features per garden and year. The dataset is a CSV file and is structured as described below: * garden_year_ID: ID for garden per year level (\"gardencode_year\") * garden_code: Garden abbreviation for analysis * year: Year of data sampling * city: City of data sampling * Andrena_alfkenella - Xylocopa_violacea: Wild bee species abundance based on four rounds per year of flower observations plus netting and pan trap sampling * open_ground: The estimated percentage of bare soil within the 400 m² sampling plots within each garden (study sites). Two values per garden (one each year). 2021: round 4, 2022: round 2 (based on number of observers and their experience) * beehotels_x_y: The number of artificial insect nesting aids (bee hotels) within the 400 m² sampling plots (x) and within a 10 m buffer surrounding the plots (y). Two values per garden (one each year). 2021: round 4, 2022: round 2 (based on number of observers and their experience) * stones: The number of stone structures (sum of the \"rock_structures_x_y\", counting rock structures, and \"drywall_x_y\", counting dry stone walls) within the 400 m² sampling plots (x) and within a 10 m buffer surrounding the plots (y). Two values per garden (one each year). 2021: round 4, 2022: round 2 (based on number of observers and their experience) * Asteraceae_prop_ga_yr: Proportion of Asteraceae in the garden (using average cover of all four rounds of all 8 1x1 m plots) * Lamiaceae_prop_ga_yr: Proportion of Lamiaceae in the garden (using average cover of all four rounds of all 8 1x1 m plots) * Fabaceae_prop_ga_yr: Proportion of Asteraceae in the garden (using average cover of all four rounds of all 8 1x1 m plots) * flow_rich_ga_yr: Number/Richness of flowering plant species per garden and year (using sum of all four rounds of all 8 1x1 m plots; calculated using vegan) * tree_rich_ga_yr: Number/richness of tree species calculated using vegan per garden and year (sum of all four rounds per year) * bee_rich_ga_yr: Number/richness of wild bee species calculated using vegan per garden and year (sum of all four rounds per year) * shannon_ga_yr: Shannon diversity index of wild bees calculated using vegan per garden and year (using \"bee_rich_ga_yr\" and \"bee_abund_ga_yr\") * bee_abund_ga_yr: Number/abundance of wild bee individuals calculated using vegan per garden and year (sum of all four rounds per year) * bee_even_ga_yr: Evenness of wild bees, calculated using vegan per garden and year * fdis: Wild bee functional dispersion calculated using mFD * feve: Wild bee functional evenness calculated using mFD * fric: Wild bee functional richness calculated using mFD The following variables are NOT included in this subset dataset, because they have been published before under a CC BY 4.0 Licence on Zenodo ([https://doi.org/10.5281/zenodo.10961505](https://doi.org/10.5281/zenodo.10961505)) in association with the publication of Neumann et al. (2024): ``` deadwood_x_y: The number of deadwood pieces within the 400 m² sampling plots (x) and within a 10 m buffer surrounding the plots (y). Two values per garden (one each year). 2021: round 4, 2022: round 2 (based on number of observers and their experience) garden_area: Garden size in m² (one value per year) impervious_1000: Percentage of impervious surfaces around the garden within a 1000 m radius ``` *Download and merge missing variables*: To follow the provided analyses, these three variables have to be added to \"Wildbees_features_garden_year_subset.csv\". Follow the following steps to create the full dataset \"Wildbees_features_garden_year.csv\", which can be used for the analyses: 1. Download the dataset “Garden features, diversity and abundance of pollinators in urban community gardens” with the file name “DATA_pollinators_garden_features_11-04-2024” from Zenodo ([https://doi.org/10.5281/zenodo.10961505](https://doi.org/10.5281/zenodo.10961505)) 2. Extract the variables “deadwood”, “garden_area”, “impervious_1000”, and for creating the identifier for merging the data “garden_code” and “year”. 3. Summarize the dataset to a garden/year level. The new variables should consist of the following data: * “deadwood”: For 2021, only use data from round 4 and for 2022 only use data from round 2 * “garden_area”: One value should exist per garden and year. * “impervious_1000”: One value should exist per garden and year. 4. Create an identifier named “garden_year_ID” from “garden_code” and “year” in the format “garden_code”_”year”, e.g. AKT_2021 5. Rename the variable “deadwood” to “deadwood_x_y”. 6. Download the dataset provided here: WildBeeTraitsProject/Input/Wildbees_features_garden_year_subset.csv 7. Merge the variables “garden_area”, “impervious_1000” and “deadwood_x_y” with the dataset for garden/year level from the input folder “Wildbees_features_garden_year_subset” using “garden_year_ID” as identifier. 8. Save the new dataset WildBeeTraitsProject/Input/ and name it “Wildbees_features_garden_year.csv”. #### Wild bee diversity and garden features on a survey level *Filename*: Wildbees_features_survey_subset.csv *Format*: Semicolon-delimited CSV file *Location*: **This dataset is found under the following path: WildBeeTraitsProject/Input/** *Description*: The following dataset contains the bee species and abundances, wild bee taxonomic and functional diversity metrics and local and landscape garden features per survey and garden. The dataset in a CSV file and is structured as described below: * garden_code: Garden abbreviation for analysis * city: City of data sampling * round: Number of garden visit/sampling round per year * year: Year of data sampling * survey_ID: ID for each survey per garden, round and year (\"gardencode_round_year\") * Andrena_alfkenella - Xylocopa_violacea: Wild bee species abundance per round (flower observations plus netting and pan trap sampling) * bee_rich_surv: Number/richness of wild bee species calculated using vegan per survey * bee_shannon_surv: Shannon diversity index of wild bees calculated using vegan per survey (using \"bee_rich_surv\" and \"bee_abund_surv\") * bee_abund_surv: Number/abundance of wild bee individuals calculated using vegan per survey * month: Month of garden visit/sampling * garden: Garden name * open_ground: The estimated percentage of bare soil within the 400 m² sampling plots within each garden (study sites). Two values per garden (one each year). 2021: round 4, 2022: round 2 (based on number of observers and their experience) * beehotels_x_y: The number of artificial insect nesting aids (bee hotels) within the 400 m² sampling plots (x) and within a 10 m buffer surrounding the plots (y). Two values per garden (one each year). 2021: round 4, 2022: round 2 (based on number of observers and their experience) * stones: The number of stone structures (sum of \"rock_structures_x_y\", counting rock structures, and \"drywall_x_y\", counting dry stone walls) within the 400 m² sampling plots (x) and within a 10 m buffer surrounding the plots (y). Two values per garden (one each year). 2021: round 4, 2022: round 2 (based on number of observers and their experience) * garden_year_ID: ID for garden per year level (\"gardencode_year\") * fdis: Wild bee functional dispersion calculated using mFD per garden and year * feve: Wild bee functional evenness calculated using mFD per garden and year * fric: Wild bee functional richness calculated using mFD per garden and year * tree_flow_rich_surv: Number/richness of flowering tree species calculated using vegan per survey * flow_rich_survey: Number/Richness of flowering plant species per survey (using sum of all 8 1x1 m plots; calculated using vegan) * Asteraceae_prop_survey: Proportion of Asteraceae in the garden per survey (using average cover of all 8 1x1 m plots) * Lamiaceae_prop_survey: Proportion of Lamiaceae in the garden per survey (using average cover of all 8 1x1 m plots) * Fabaceae_prop_survey: Proportion of Asteraceae in the garden per survey (using average cover of all 8 1x1 m plots) * bee_even_surv: Evenness of wild bees, calculated using vegan per survey and garden The following variables are NOT included in this subset dataset, because they have been published before under a CC BY 4.0 Licence on zenodo ([https://doi.org/10.5281/zenodo.10961505](https://doi.org/10.5281/zenodo.10961505)): ``` deadwood_x_y: The number of deadwood pieces within the 400 m² sampling plots (x) and within a 10 m buffer surrounding the plots (y). Two values per garden (one each year). 2021: round 4, 2022: round 2 (based on number of observers and their experience) garden_area: Garden size in m² (one value per year) impervious_1000: Percentage of impervious surfaces around the garden within a 1000 m radius (one value per year) ``` *Download and merge missing variables*: To follow the provided analyses, these three variables have to be added to \"Wildbees_features_survey_subset.csv\". Follow the following steps to create the full dataset \"Wildbees_features_survey.csv\", which can be used for the analyses: 1. Download the dataset “Garden features, diversity and abundance of pollinators in urban community gardens” with the file name “DATA_pollinators_garden_features_11-04-2024” from Zenodo ([https://doi.org/10.5281/zenodo.10961505](https://doi.org/10.5281/zenodo.10961505)) 2. Extract the variables “deadwood”, “garden_area”, “impervious_1000”, “round” and “ID” as identifiers for merging the data. 3. Rename the following variables: “ID” -\u0026gt; “survey_ID” 4. Download the dataset provided here: WildBeeTraitsProject/Input/Wildbees_features_survey_subset.csv 5. Merge the variables “garden_area” and “impervious_1000” as they are from “DATA_pollinators_garden_features_11-04-2024” with the dataset for survey level in the input folder “Wildbees_features_survey_subset” using “survey_ID” as identifyer. 6. Create the new variable “deadwood_x_y” in “Wildbees_features_survey_subset”, which is later used in the analysis. 7. Select from the variable “deadwood” for each garden the values from round 4 in 2021. Paste these values into the new variable “deadwood_x_y” for all rounds in 2021. All gardens in the new variable “deadwood_x_y” should have the same values for each round per garden in 2021. Repeat the same for round 2 in 2022. The aim is to have the same value for deadwood per garden and year, using round 4 in 2021 and round 2 in 2022. 8. Delete the former variable “deadwood”. 9. Save the new dataset in WildBeeTraitsProject/Input/ and name it “Wildbees_features_survey.csv”. #### Wild bee trait data *Filename*: Wildbees_traits_taxonomy_excl_ITD.csv *Format*: Semicolon-delimited CSV file *Location*: **This dataset is found under the following path: WildBeeTraitsProject/Input/** *Description*: The following dataset contains all wild bee species that we sampled using flower visitor observations and pan traps and corresponding trait data. Trait data derived from: Westrich (2018), Scheuchl \u0026amp; Willner (2016), Amiet et al. (2001), Amiet et al. (2005), Amiet et al. (2017), Amiet \u0026amp; Krebs (2019), Weissmann \u0026amp; Schaefer (2022), European Bee Traits Database (Roberts, unpublished). The dataset is a CSV file and is structured as described below: * bee_species: bee species names with the structure \"genus_species\" * bee_genus: bee genera * bee_family: bee families * nesting: nesting type with five categories: vegetation and cavities, parasite, ground, cavities, deadwood * sociality: type of social system with four categories: communal, parasitic, social, solitary * specialisation: type of food specialisation with three categories: oligolectic, polylectic, parasitic * activity_period_months: activity period per year with the categories 2-9 (months) * month_1st_flying: month of first emergence: February, March, April, May, June, July * proboscis_length: length of proboscis with two categories: long, short * red_list: Red List Status with IUCN categories: 2, 3, G, no_data, not_threatened, V The following variable is not available in this repository until European Bee Traits Database is published by S. Roberts et al. ``` ITD_females: Mean intertegular distance (ITD) per species from European Bee Traits Database ``` ## Scripts The R scripts provided were written in R Version 4.5.1. The scripts include the following: **01_GLMM**: * Description: The R script produces the LM(M) analyses on a garden/year-level and a survey-level * Input: Wildbees_features_garden_year.csv; Wildbees_features_survey.csv **02_RLQ**: * Description: The R script computes the RLQ and fourth corner analyses and produces Figure 5 * Input: Wildbees_features_garden_year.csv; Wildbees_features_survey.csv; Wildbees_traits_taxonomy_excl_ITD.csv * Output: figure_5 **03_NMDS_garden_level**: * Description: The R script computes the NMDS and vector analysis on garden/year-level and creates three files that are needed as input for the script 04_Figures_NMDS_garden_level * Input: Wildbees_features_garden_year.csv; Wildbees_traits_taxonomy_excl_ITD.csv * Output: NMDS_nms_spec_traits_gy10.csv; NMDS_en_coord_gy10.csv; NMDS_data.scores.gy10.csv **03_NMDS_survey_level**: * Description: The R script computes the NMDS and vector analysis on survey-level and creates three files that are needed as input for the script 04_Figures_NMDS_survey_level * Input: Wildbees_features_survey.csv; Wildbees_traits_taxonomy_excl_ITD.csv * Output: NMDS_en_coord10.csv; NMDS_nms_spec_traits10.csv; NMDS_data.scores10.csv **04_Figures_NMDS_garden_level**: * Description: the R script produces Figure S5 (NMDS and vector analysis on garden/year-level); run script 03_NMDS_garden_level first to produce input files * Input: NMDS_nms_spec_traits_gy10.csv; NMDS_en_coord_gy10.csv; NMDS_data.scores.gy10.csv * Output: figure_S5 (a-f) **04_Figures_NMDS_survey_level**: * Description: The R script produces Figure 4 (NMDS and vector analysis on survey-level); run script 03_NMDS_survey_level first to produce input files * Input: NMDS_en_coord10.csv; NMDS_nms_spec_traits10.csv; NMDS_data.scores10.csv * Output: figure_4 (a-f) **04_Figures_GLMM**: * Description: The R script produces Figure 3 and calculates the descriptive statistics of bee diversity indices and garden features * Input: Wildbees_features_garden_year.csv; Wildbees_features_survey.csv * Output: figure_3 **04_Figures_traits_overview**: * Description: The R script produces Figure 2 and the descriptive statistics for intertegular distance (ITD) * Input: Wildbees_traits_taxonomy_excl_ITD.csv; Wildbees_features_survey.csv * Output: figure_2 **04_Figures_species_overview**: * Description: The R script produces Figures 1 (Rank-abundance curve) and S4 (Bee occurrences per garden) * Input: Wildbees_features_survey.csv * Output: figure_1; figure_s4 **05_INEXT**: * Description: the R script produces a sample-size-based rarefaction-extrapolation curve (Figure S3), and computes the sample coverage (SC) * Input: Wildbees_features_garden_year.csv * Output: figure_s3 ## References: Neumann, A. E., Conitz, F., Karlebowski, S., Sturm, U., Schmack, J. M. \u0026amp; Egerer, M. (2024) Flower richness is key to pollinator abundance: The role of garden features in cities. Basic and Applied Ecology, 79, 102–113. [https://doi.org/10.1016/j.baae.2024.06.004](https://doi.org/10.1016/j.baae.2024.06.004) Neumann, A. E., Conitz, F., Karlebowski, S., Sturm, U., Schmack, J., \u0026amp; Egerer, M. (2024). Garden features, diversity and abundance of pollinators in urban community gardens [Data set]. In Basic and Applied Ecology (Bd. 79, S. 102–113). Zenodo. [https://doi.org/10.5281/zenodo.10961505](https://doi.org/10.5281/zenodo.10961505)"}],"geoLocations":[],"fundingReferences":[{"funderIdentifierType":"ROR","funderName":"Heidehof Stiftung","funderIdentifier":"https://ror.org/02xq7zd76","awardNumber":"57358.01.1/4.20"},{"funderIdentifierType":"ROR","funderName":"Heidehof Stiftung","funderIdentifier":"https://ror.org/02xq7zd76","awardNumber":"57358022"},{"funderName":"Deutsche Postcode Lotterie","awardNumber":"FA-8156"},{"funderIdentifierType":"ROR","funderName":"Swiss National Science Foundation","funderIdentifier":"https://ror.org/00yjd3n13","awardNumber":"217754"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.66t1g1kcz","contentUrl":null,"metadataVersion":0,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-21T18:30:30Z","registered":"2026-08-21T18:30:31Z","published":null,"updated":"2026-08-21T18:30:31Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.79cnp5jcc","type":"dois","attributes":{"doi":"10.5061/dryad.79cnp5jcc","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["University of Oxford"],"name":"López-Idiáquez, David","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0001-9568-4852"}]},{"nameType":"Personal","affiliation":["University of Oxford"],"name":"Cole, Ella","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Oxford"],"name":"Satarkar, Devi","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Oxford"],"name":"Crofts, Samuel","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Oxford"],"name":"McMahon, Keith","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Oxford"],"name":"Sheldon, Ben","nameIdentifiers":[]}],"titles":[{"title":"Early-life environment drives long-term decrease in adult body mass in a wild bird population"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Evolutionary ecology","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Climate change","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Quantitative traits","subjectScheme":"PLOS Subject Area Thesaurus"}],"contributors":[],"dates":[{"date":"2026-07-26T10:02:56Z","dateType":"Created"},{"date":"2026-08-05T17:29:44Z","dateType":"Submitted"},{"date":"2026-08-21T00:00:00Z","dateType":"Issued"},{"date":"2026-08-21T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[{"relationType":"IsCitedBy","relatedIdentifier":"10.64898/2026.02.11.705378","relatedIdentifierType":"DOI"}],"relatedItems":[],"sizes":["20473262 bytes"],"formats":[],"version":"3","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Body mass is a key organismal characteristic that impacts many\n physiological and ecological processes and is often a strong determinant\n of fitness. Many recent studies have documented temporal phenotypic\n changes in this trait in wild animal populations, but identifying the\n mechanisms underpinning these changes can be difficult. While most\n research conducted to date has focused on temperature changes as a driver\n of these trends, the relevance of other environmental variables remains to\n be analysed, limiting our understanding of the factors driving these\n changes. Here, we use 47 years of data to decompose mechanisms behind\n temporal changes in adult and nestling body mass in a great tit Parus\n major population in Wytham Woods (UK). Further, we link those changes to\n temperature, the main driver of temporal trends in mass according to the\n literature, and to two other environmental variables previously recognised\n as drivers of body mass: intra- and inter-specific competition and\n temporal mismatch with a key prey during breeding, winter moth Operophtera\n brumata caterpillars. At the population level we report a marked decrease\n in body mass in adults between 1978 and 2024 (-0.042 Haldanes; -0.020\n grams year-1), and show that this results from phenotypic plasticity,\n driven by a negative between-cohort trend likely reflecting carry-over\n effects of the early environment. Within cohorts, however, trends were\n consistently positive, reflecting an age-dependent mass increase. The\n temporal change in adults was paralleled by a change in nestling body mass\n (-0.036 Haldanes). Nestling mass was negatively associated with estimated\n intensity of intraspecific competition, as well as inter-specific\n competition from blue tits Cyanistes caeruleus, as quantified by local\n population density. These effects carried over to adulthood, as shown by a\n negative association between adult mass and the population density of\n great tits and blue tits experienced at early life. Seasonal and\n developmental temperature and mismatch with the caterpillar food supply,\n despite being associated with adult and nestling mass, did not explain the\n observed declines in mass. Our results illustrate the potential for\n effects mediated early in development to carry-over into long-term\n phenotypic change at later life history stages, and emphasise the\n importance of considering different environmental variables as drivers of\n phenotypic change in natural populations."},{"descriptionType":"TechnicalInfo","description":"# Early-life environment drives long-term decrease in adult body mass in a\n wild bird population Dataset DOI:\n [10.5061/dryad.79cnp5jcc](https://doi.org/10.5061/dryad.79cnp5jcc) ##\n Description of the data and file structure Data from the long term\n monitoring population of Wytham Woods, near Oxford. ### Files and\n variables #### File: code.Rmd **Description:** Code required to replicate\n the analyses of the paper. This code allows to replicate all the models in\n the manuscript, including analysing the temporal trends in adult and\n nestling mass, analysing the links between adult and nestling mass and the\n environmental conditions (temperature, breeding density and mismatch with\n the peak of food availability) and run the quantitative genetic models to\n compute the heritability of adult mass and the correlations at the genetic\n and year of birth levels between adult and nestling mass (For further\n details, see manuscript). #### File: adult_mass_ran_slo.csv\n **Description:** Data required to replicate the analysis exploring the\n temporal trends in adult mass within and between cohorts. ##### Variables\n * ring: Individual identity * Pnum: Brood Identity * year: Year of\n measurement * sex: Sex (males (M) vs females (F)) * weight: Body mass (in\n grams) * box: Nestbox identity * Section: Section identity (9 levels\n representing the names of the sections) * year_cont: Year of measurement\n as a continuous variable * b_year: year of birth * est_byear: estimated\n year of birth for those individuals dispersing to the population (they are\n assumed to arrive in the population at age=2 years old) * year_fac: year\n as a categorical variable * est_byear_fac: Estimated year of birth as a\n categorical variable #### File: data_adults.csv **Description:** Data\n required to replicate the analysis exploring the temporal trends in adult\n mass and the links between adult mass and the environmental variables, and\n the quantitative genetic analyses. ##### Variables * ring: Individual\n identity * Pnum: Brood identity * year: Year of measurement * sex: Sex\n (males (M) vs females (F)) * weight: Body mass (in grams) * box: Nestbox\n identity * age_cod: Age as juveniles (juv) vs adults (ad) * Section:\n Section identity (9 levels representing the names of the sections) *\n year_cont: Year as a continuous variable * animal: Individual identity to\n link the pedigree * tmean_breeding: Mean temperature at breeding season\n (in ºC) * tmean_winter: Mean temperature in winter (in ºC) *\n average_temperature_capture: Mean temperature during capture (in ºC) *\n n_days: Period duration * buffer_Xm_Y: Number of blue tits (Y=b) or great\n tits (Y=g) pairs breeding in a buffer X meters around each focal nest box\n where adults breed in a certain year. There are 17 buffers of different\n sizes (from X=100 meters to X=4000). * mismatch: Mismatch, in days, with\n the caterpillar half fall date. Mismatch was not available for all years\n and rows with no mismatch information have NA (see *Linking body mass and\n mismatch* section for further information). #### File: puned_ped.csv\n **Description:** Social pedigree of Wytham Woods, pruned to retain\n informative individuals and describing the relatedness of the individuals\n in the population. NAs correspond to unknown parents. ##### Variables *\n id: Id of the individual * dam: Id of the mother * sire: Id of the father\n #### File: data_nestlings.csv **Description:** Data required to replicate\n the analysis exploring the temporal trends in nestling mass and its links\n to the environmental variables (breeding density, temperature, and\n mismatch) and the quantitative genetic analyses ##### Variables * pnum:\n Brood  * bto_ring: Individual ID * year: Year of measurement * nest:\n Nestbox ID * year_cat: Year of measurement as a categorical variable *\n Section: Section ID * tmean_breeding: Mean temperature during the breeding\n season (in ºC) * tmean_winter: Mean temperature in winter (in ºC) *\n average_temperature_capture: Mean temperature during development (in ºC) *\n Num_Days_laying: Number of developmental days * buffer_Xm_Y: Number of\n blue tits (Y=b) or great tits (Y=g) pairs breeding in a buffer X meters\n around each focal nest box where adults breed in a certain year. There are\n 17 buffers of different sizes (from X=100 meters to X=4000). *\n mismatch: Mismatch, in days, with the caterpillar half fall date. Mismatch\n was not available for all years and rows with no mismatch information have\n NA (see *Linking body mass and mismatch* section for further information).\n * recr: Defines whether a nestling was recruited (recr=1) to the breeding\n population or not (recr=NA) * mass_m: Nestling mass (in grams) ##\n Code/software The analyses were conducted using R and the following\n packages: - lmertest (v. 3.1-3) - lme4 (v. 1.1-37) - ggeffects (v. 2.3.1)\n - ggplot2 (v. 4.0.3) - dplyr (v. 1.1.4) - brms (v. 2.23.0) - MCMCglmm (v.\n 2.36) - broom.mixed (v. 0.2.9.6) - splines (v. 4.5.2)"}],"geoLocations":[],"fundingReferences":[{"funderIdentifierType":"ROR","funderName":"Natural Environment Research Council","funderIdentifier":"https://ror.org/02b5d8509"},{"funderIdentifierType":"ROR","funderName":"Biotechnology and Biological Sciences Research Council","funderIdentifier":"https://ror.org/00cwqg982"},{"funderIdentifierType":"ROR","funderName":"UK Research and Innovation","funderIdentifier":"https://ror.org/001aqnf71"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.79cnp5jcc","contentUrl":null,"metadataVersion":0,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-21T18:02:17Z","registered":"2026-08-21T18:02:18Z","published":null,"updated":"2026-08-21T18:02:18Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.69p8cz9bw","type":"dois","attributes":{"doi":"10.5061/dryad.69p8cz9bw","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["Wageningen University \u0026 Research"],"name":"Cribellier, Antoine","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0001-5113-7531"}]},{"nameType":"Personal","affiliation":["Wageningen University \u0026 Research","Institut de Recherche en Sciences de la Santé"],"name":"Poda, Serge","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Institut de Recherche en Sciences de la Santé"],"name":"Dabiré, Roch K.","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Institut de Recherche en Sciences de la Santé"],"name":"Diabaté, Abdoulaye","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Maladies Infectieuses et Vecteurs: Écologie, Génétique, Évolution et Contrôle"],"name":"Roux, Olivier","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Wageningen University \u0026 Research"],"name":"Muijres, Florian T.","nameIdentifiers":[]}],"titles":[{"title":"Data and code from: The complex swarming dynamics of malaria mosquitoes emerge from simple minimally-interactive behavioral rules"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"subject":"mosquito swarm"},{"subject":"boids model"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Agent-based modeling","subjectScheme":"PLOS Subject Area Thesaurus"},{"subject":"behavioural rules"},{"subject":"Anopheles coluzzii"}],"contributors":[],"dates":[{"date":"2024-09-26T13:37:07Z","dateType":"Created"},{"date":"2026-08-18T11:12:08Z","dateType":"Submitted"},{"date":"2026-08-21T00:00:00Z","dateType":"Issued"},{"date":"2026-08-21T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[{"relationType":"IsCitedBy","relatedIdentifier":"10.1101/2024.08.31.610631","relatedIdentifierType":"DOI"},{"relationType":"IsDerivedFrom","relatedIdentifier":"10.5281/zenodo.13843637","relatedIdentifierType":"DOI"}],"relatedItems":[],"sizes":["240158764 bytes"],"formats":[],"version":"6","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Many insect species exhibit swarming behavior, often to reproduce. In such\n mating swarms, tens to thousands of male insects congregate near a visual\n ground marker, and females join for in-flight mating. Despite its\n importance for reproduction and survival, the biology and physics of\n insect swarming remains poorly understood. In particular, it is unclear\n whether swarming insects follow similar highly interactive behavioral\n rules as flocking birds and schooling fish. In flocking and schooling,\n neighbors exhibit strong mutual attraction, alignment, and collision\n avoidance. Here, we combined high-speed videography-based experiments with\n agent-based modeling to study the mating swarms of male malaria mosquitoes\n at sunset. By video-tracking the three-dimensional flight kinematics of\n swarming mosquitoes, we revealed that these animals exhibited highly\n stereotypic flight behaviors. Swarming mosquitoes tend to fly straight\n across the swarm center, and perform sharp saccadic turns at the swarm\n boundaries. Based on these observations, we derived parsimonious\n behavioral rules for swarming: for flight path alignment and swarm\n centering, mosquitoes rely on visual cues of the ground marker and sunset;\n and interactions between conspecifics only occur at very short distances,\n to avoid collisions. Using agent-based modelling and systematic\n simulations, we showed that this small set of behavioral rules is both\n necessary and sufficient to reproduce the complex emergent flight\n kinematics and coordinated patterns observed in real swarms. This suggests\n that, unlike highly interactive bird flocking and fish schooling, insect\n swarming may be primarily driven by responses to environmental cues\n instead of mutual interactions."},{"descriptionType":"TechnicalInfo","description":"# Data and code from: The complex swarming dynamics of malaria mosquitoes\n emerge from simple minimally-interactive behavioral rules Dataset DOI:\n [10.5061/dryad.69p8cz9bw](https://doi.org/10.5061/dryad.XXXXXXXXX) ##\n Associated publication Cribellier A., Poda B.S., Dabiré R.K., Diabaté A.,\n Roux O., Muijres F.T. *The complex swarming dynamics of malaria mosquitoes\n emerge from simple minimally-interactive behavioral rules.* Submitted to\n *PLOS Computational Biology*. Corresponding author: Antoine Cribellier,\n Experimental Zoology Group, Wageningen University, Wageningen, The\n Netherlands\n ([antoine.cribellier@wur.nl](mailto:antoine.cribellier@wur.nl)). ##\n Description of the data and file structure ### Aim and scope of the\n dataset Male malaria mosquitoes (*Anopheles coluzzii*) aggregate at dusk\n in station-keeping swarms above a visually contrasting ground object (the\n *swarm marker*), where females come to mate. This deposit contains the two\n datasets that underlie the study: (i) the **measured** three-dimensional\n flight tracks of *An. coluzzii* males swarming above a 40 x 40 cm black\n marker in a laboratory arena under simulated-sunset conditions (Data S1),\n and (ii) the **simulated** three-dimensional flight tracks produced by the\n agent-based model (ABM) that reproduces the observed swarming dynamics\n from a minimal set of individual behavioral rules (Data S2). The central\n hypothesis tested with these data is that swarm-level structure and\n dynamics do not require rich mosquito-to-mosquito interactions, but emerge\n from two simple, mostly non-social rules: (1) individuals perform saccadic\n turns triggered by the apparent size of the marker in their visual field\n (the marker viewing angle, α), and (2) individuals avoid collisions with\n nearby conspecifics at short range. Data S1 provides the empirical flight\n kinematics and the empirical turn-angle distributions used to parameterise\n the model; Data S2 provides the model output for the two key model\n variants (with and without the collision-avoidance rule) that are compared\n against the measurements in the manuscript. ### Overview of the deposited\n files | File | Type | Content | | --------------- |\n ---------------------------------------------- |\n ---------------------------------------------------------------------------------------------------------------------------------------- | | `Data_S1.zip` | ZIP archive containing one MATLAB `.mat` file | Measured 3D flight tracks of swarming *An. coluzzii* males (6 experiments) | | `Data_S2.zip` | ZIP archive containing two MATLAB `.mat` files | Simulated 3D flight tracks from the agent-based model (2 model variants x 20 replicates) | | `Table_S1.xlsx` | Excel spreadsheet | Complete list of agent-based model parameters, their symbols, names in the MATLAB code, descriptions, default values, units, and sources | | `Video_S1.mp4` | MPEG-4 video | Swarming flight of an initiator mosquito (measured) | | `Video_S2.mp4` | MPEG-4 video | Flight kinematics of a single swarming mosquito (measured) | | `Video_S3.mp4` | MPEG-4 video | Swarming flight of multiple mosquitoes (measured) | | `Video_S4.mp4` | MPEG-4 video | Swarming flight of ten simulated mosquitoes (agent-based model, with collision avoidance) | ### Conventions shared by all `.mat` files **File format.** All `.mat` files are MATLAB v7.3 files, i.e., HDF5 containers. They open natively in MATLAB (`load('filename.mat')`) and can also be read in Python with `h5py`, or in R with `rhdf5`. They cannot be read with `scipy.io.loadmat`, which supports only MATLAB v5-v7.2 files. Note that HDF5 stores MATLAB arrays in transposed order, so an `N x 1` MATLAB column vector appears as shape `(N, 1)` in `h5py` but character arrays appear as arrays of `uint16` character codes that must be converted back to text. **Top-level structure.** Each `.mat` file contains a single MATLAB structure named `all_data` with two branches: * `all_data.cst` — constants that apply to the whole dataset (frame rate; for Data S2, all model parameters). * `all_data.exp` — one sub-structure per experiment (Data S1) or per simulation replicate (Data S2), each holding `metadata` and `tracks`. **Coordinate system.** All positions are in metres, in the world reference frame defined in Figure 1 of the manuscript: the origin is at the centre of the swarm marker on the ground, *z* points vertically upwards, *y* points horizontally towards the simulated sun / sunset horizon, and *x* is horizontal and parallel to the sunset horizon, completing a right-handed frame. Velocities are in m s⁻¹ in the same frame. Angles are in degrees unless stated otherwise. **Time base.** Tracks were sampled at 50 Hz (`all_data.cst.fps = 50`). `frames` is the video frame index and `time = frames / fps` is the elapsed time in seconds from the start of the recording session (Data S1) or from the start of the simulation (Data S2). **Missing values.** Missing samples are encoded as `NaN`. In Data S1, `NaN` marks frames within an individual's tracked interval in which the mosquito was not detected in enough camera views to be reconstructed in 3D (typically short gaps between stitched track segments). Data S2 contains no missing values. --- ### Data S1 (`Data_S1.zip`): measured flight tracks of swarming *Anopheles coluzzii* males The archive contains a single file: **`all_data-exp1-coluzzii_males-selection_swarming.mat`** — 3D flight tracks of male *An. coluzzii* mosquitoes swarming above a 40 x 40 cm black marker in a climate-controlled laboratory arena (2.0 x 0.7 x 1.8 m) under simulated-sunset lighting. Trajectories were reconstructed from synchronised multi-camera video as described in the Materials and Methods of the manuscript. As the file name indicates, the deposit contains a *selection* of the recorded swarming episodes: the swarm onset, plus fixed time windows sampling the build-up, plateau and decay phases of each swarm (see `dynamic` below). **Structure** ``` all_data ├── cst │ └── fps 50 (Hz), frame rate of all tracks └── exp ├── exp20221007_183000 one field per experiment (6 in total) │ ├── metadata │ │ ├── species 'coluzzii' │ │ ├── sex 'male' │ │ ├── age age of the mosquitoes at the time of the experiment (days) │ │ ├── nb number of mosquitoes released in the arena │ │ ├── date date and time of the experiment (MATLAB datetime) │ │ └── marker x, y, z: position of the marker centre in the raw │ │ tracking frame of the arena (m); tracks are already │ │ expressed relative to this point │ └── tracks │ └── object │ ├── obj1 one field per tracked individual │ ├── obj2 │ └── ... ├── exp20221010_183000 └── ... ``` Experiment fields are named `expYYYYMMDD_HHMMSS` after the date and start time of the recording session (all sessions started at 18:30 local time). **Per-experiment metadata** | Experiment | Mosquito age (days) | Mosquitoes released | Tracked individuals | | -------------------- | ------------------- | ------------------- | ------------------- | | `exp20221007_183000` | 5 | 30 | 162 | | `exp20221010_183000` | 6 | 30 | 137 | | `exp20221012_183000` | 7 | 30 | 182 | | `exp20221028_183000` | 6 | 50 | 137 | | `exp20221030_183000` | 6 | 50 | 209 | | `exp20221103_183000` | 7 | 50 | 111 | | **Total** | \n \n | \n \n | **938** | Together the six experiments contain 938 individual\n trajectories, 1,816 raw track segments and 709,920 valid 3D position\n samples (≈ 3.9 h of tracked flight at 50 Hz). **Variables inside each**\n **`obj`** Each `obj` is one tracked individual mosquito\n over one continuous tracked interval. `objN` (together with the experiment\n name) is the unique identifier of a trajectory in this dataset. The\n time-series fields below are all `N x 1` column vectors of the same length\n `N`, aligned sample by sample. | Variable | Type / size | Units |\n Description | | ---------- | -------------- | ----- |\n --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `frames` | `N x 1` double | — | Video frame index, counted from the start of the recording session. Consecutive (step 1) over the tracked interval. | | `time` | `N x 1` double | s | Elapsed time since the start of the recording session, `time = frames / 50`. | | `x` | `N x 1` double | m | Position along the horizontal axis parallel to the sunset horizon. `NaN` where the mosquito was not reconstructed. | | `y` | `N x 1` double | m | Position along the horizontal axis pointing towards the simulated sun. | | `z` | `N x 1` double | m | Height above the plane of the swarm marker. | | `x_vel` | `N x 1` double | m s⁻¹ | Velocity component along *x*, obtained by numerical differentiation of the (smoothed) position. | | `y_vel` | `N x 1` double | m s⁻¹ | Velocity component along *y*. | | `z_vel` | `N x 1` double | m s⁻¹ | Velocity component along *z*. | | `track_id` | `N x 1` double | — | Identifier of the raw track segment that produced each sample. Individuals whose trajectory was interrupted and re-acquired are represented by several segments stitched into one `obj`; `track_id` records which segment each sample came from. `NaN` in the gaps between segments. | **Scalars inside each** **`obj.scalars`** (constant for that individual) | Variable | Type / size | Description | | ----------------- | --------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `mosqID` | `1 x 1` double | Index of the individual within its swarm phase (see `dynamic_id`). Not globally unique — use the experiment name plus `obj` as the primary key. | | `sex` | char | Sex of the mosquito; `'male'` throughout. | | `dynamic` | char | Phase of the swarm during which the individual was tracked: `'increasing'` (swarm forming), `'top'` (swarm at its plateau), `'decreasing'` (swarm dispersing). | | `dynamic_id` | `1 x 1` double | Numeric code of the swarm phase, with a finer split than `dynamic`: `1` = swarm initiator (the first male to start swarming; `dynamic` = `'increasing'`), `2` = build-up phase (`dynamic` = `'increasing'`), `3` = plateau, a window of up to 150 s starting 900 s after the start of the session (`dynamic` = `'top'`), `4` = decay, the last \\~150 s of the session (`dynamic` = `'decreasing'`). | | `was_initiator` | `1 x 1` logical | `true` for the single initiator of each experiment (6 individuals in total, one per experiment), `false` otherwise. | | `num_recording` | `1 x 1` double | Index of the video recording, within the session, that contains this trajectory (each session was captured as several successive recordings). | | `start_recording` | `1 x 1` logical | `true` if the individual was already being tracked at the first frame of that recording, i.e. the trajectory is truncated at its start. | | `trackID` | `M x 1` double | List of the `M` raw track segments stitched into this individual; the values are those found in the `track_id` time series. | | `trackID_frames` | `M x 1` double | First video frame of each of those `M` segments. | | `trackID_counts` | `M x 1` double | Number of samples in each of those `M` segments. | **Derived quantities.** The kinematic and visual variables analysed in the manuscript — accelerations, jerk, flight speed, angular speed, cylindrical coordinates (*r*, θ), the marker viewing angles α_x and α_y, saccade detection and saccade properties, swarm-level metrics and heat maps — are **not** stored in this file. They are recomputed from the raw positions above by the analysis code (see *Code/software*); `main.m` and `gen_all_data_tracks.m` list them explicitly. --- ### Data S2 (`Data_S2.zip`): simulated flight tracks from the agent-based model The archive contains two files, one per model variant. Both were produced with the same behavioral rules and the same parameter values (Table S1); they differ only in whether the short-range collision-avoidance rule was active: | File | Collision avoidance | | ----------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------- | | `all_data-abm-along_y_rand-null-threshold-any-alpha_x_percent-threshold_alpha_x.mat` | **off** (`cst.with_collision_avoidance = 0`) | | `all_data-abm-along_y_rand-null-threshold-any-alpha_x_percent-threshold_alpha_x-avoid_collisions.mat` | **on** (`cst.with_collision_avoidance = 1`, `cst.min_distance = 0.006` m) | The file names encode the model configuration as `-----`; each of these components is a field of `all_data.cst` and is defined in `Table_S1.xlsx`. Each file contains 20 independent replicate simulations of 10 agents flying for 120 s at 50 Hz (6,000 frames), i.e. 200 simulated trajectories and 1,200,000 position samples per file. **Structure** ``` all_data ├── cst all model parameters (see Table_S1.xlsx) └── exp ├── rep1 one field per replicate simulation (20 in total) │ ├── metadata │ │ ├── species 'boid' │ │ └── sex 'male' │ └── tracks │ └── object │ ├── boid1 one field per simulated agent (10 per replicate) │ ├── boid2 │ └── ... ├── rep2 └── ... ``` **`all_data.cst`: model parameters** `Table_S1.xlsx` is the authoritative, annotated description of these parameters: it gives, for every parameter, the symbol used in the manuscript, the name used in the MATLAB code (which is the field name used here), a description, the default value, the units, the source or justification, the submodel in which it is used, and whether it is stochastic. **Users should consult `Table_S1.xlsx` for the definition of every field of `all_data.cst`.** For orientation, the fields present are: | Field | Value in this deposit | Meaning | | ------------------------------------------------ | ----------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | `fps`, `dt` | 50 Hz, 0.02 s | Simulation frame rate and time step | | `init_vel` | 0.5 m s⁻¹ | Constant flight speed during straight flight | | `boid_type` | `'swarming_saccades'` | Agent behavioural model used | | `sim_name` | see file names above | Full configuration string of the simulation | | `saccade_type` | `'along_y_rand'` | Rule defining the direction of a saccadic turn | | `acceleration_type` | `'null'` | No additional acceleration field applied | | `start_saccade_type` | `'threshold'` | Saccades are triggered by threshold crossing | | `start_saccade_field_names` | `{'alpha_x_percent'}` | Variable used as the saccade trigger: the relative deviation of the marker viewing angle from its maximum, R\\_α = (α\\_max − α)/α\\_max × 100 % | | `start_saccade_combination_type` | `'any'` | How multiple trigger criteria are combined | | `with_threshold_alpha_x` | 1 | Absolute bounds on α are also enforced | | `alpha_x_threshold_min`, `alpha_x_threshold_max` | 24°, 55° | Lower and upper bounds on the marker viewing angle that trigger a saccade | | `alpha_x_percent_threshold` | 4.5 % | Threshold on R\\_α that triggers a saccade | | `use_prev_max_alpha_x` | 0 | α\\_max is taken as the viewing angle at the swarm centre at the agent's current height (not the agent's own previous maximum) | | `vel_after_saccade` | `'start'` | Speed is reset to `init_vel` after each saccade | | `with_collision_avoidance` | 0 or 1 | Collision-avoidance rule off / on (see table above) | | `min_distance` | 0.006 m | Inter-agent distance below which avoidance is triggered (present only in the collision-avoidance file) | | `marker` | `x_center` = `y_center` = `z_center` = 0; `x_width` = `y_width` = 0.4 m | Geometry and position of the swarm marker; `x`, `y`, `z`, `r`, `theta`, `order` give its corner coordinates for plotting | | `bounds` | x ∈ \\[−1.45, 0.55] m, y ∈ \\[−0.35, 0.35] m, z ∈ \\[0, 1.8] m | Virtual arena, matching the dimensions of the experimental arena; `bounce_bounds = 1` means agents reflect off the walls | | `limits` | — | Plotting limits per variable, used by the figure code | | `dist` | — | Empirical distributions from which stochastic model quantities are drawn (see below) | | `show_vel` | 0 | Display flag used by the simulation's live plotting; no effect on the data | `all_data.cst.dist` holds the empirical distributions measured in Data S1 and used to draw stochastic quantities in the model. `dist.linear` contains the fitted/interpolated distributions actually sampled by the model — `angle_v_pks_theta_after_linear` (azimuth of flight direction after a saccadic turn, degrees) and `angle_v_pks_phi_after_linear` (climb angle after a saccadic turn, degrees), both shown in Figure S2, plus `r_acc`. `dist.raw` contains the corresponding raw measured distributions of a broader set of kinematic variables (flight speed `vel`, acceleration `acc`, angular speed `angle_v`, the velocity and acceleration components `x_vel`…`z_acc`, `phi_vel`, `phi_acc`, saccade speed `angle_v_pks_vel`, and the cumulative trigger variables `cumsum_alpha_x_percent`, `cumsum_time_pks`, `cumsum_distance_pks`), retained for reference. **Variables inside each** **`boid`** Each `boid` is one simulated agent over the full 120 s simulation. All time-series fields are `6000 x 1` column vectors with no missing values. | Variable | Type / size | Units | Description | | ------------------------- | ----------------- | ------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `frames` | `6000 x 1` double | — | Simulation frame index, 1 to 6000. | | `time` | `6000 x 1` double | s | Elapsed simulation time, `time = frames / 50`, 0 to 119.98 s. | | `x`, `y`, `z` | `6000 x 1` double | m | Agent position in the marker-centred world frame (same convention as Data S1). | | `x_vel`, `y_vel`, `z_vel` | `6000 x 1` double | m s⁻¹ | Agent velocity components. The speed is held at `init_vel` = 0.5 m s⁻¹ between saccades. | | `alpha_x` | `6000 x 1` double | degrees | Marker viewing angle α: the angle subtended at the agent's position by the two edges of the marker along the *x* axis. This is the visual cue that drives the saccade rule. | | `max_alpha_x` | `6000 x 1` double | degrees | Reference maximum viewing angle α\\_max: the value α would take directly above the marker centre (*x* = *y* = 0) at the agent's current height. The saccade trigger is the relative deviation R\\_α = (α\\_max − α)/α\\_max × 100 %. | | `state` | `1 x 1` double | — | Behavioural state of the agent; a single state (`1`) is implemented in this model version, so the field is constant. | | `saccades.frames` | `K x 1` double | — | Frame indices at which the agent performed a saccadic turn (`K` varies per agent; median 185, range 1-212, in the collision-avoidance simulations). The first entry is the initial frame, at which the flight direction is set. | **Scalars inside each** **`boid.scalars`** | Variable | Type / size | Units | Description | | --------------- | --------------- | ----- | --------------------------------------------------------------------------------------------------------------------------------------- | | `init_r` | `1 x 1` double | m | Radial distance from the *z* axis at which the agent was initialised (drawn uniformly). | | `init_theta` | `1 x 1` double | rad | Azimuth at which the agent was initialised (drawn uniformly in \\[−π, π]). Agents start flying towards the arena centre. | | `init_z` | `1 x 1` double | m | Height at which the agent was initialised (drawn uniformly within the arena bounds). | | `is_living` | `1 x 1` logical | — | `true` while the agent is inside the arena bounds and active. All agents remain active for the full simulation in both deposited files. | | `num_replicate` | `1 x 1` double | — | Index of the replicate simulation (1-20), matching the `rep` field name. | --- ### `Table_S1.xlsx` Single sheet (\"Table S1\") listing all 24 parameters of the agent-based model, one per row, with the columns: *Parameter name*, *Symbol*, *Name in Matlab code*, *Description*, *Default value*, *Units*, *Range / values explored*, *Source / Justification*, *Used in submodel(s)* and *Stochastic?*. The model description follows a simplified form of the ODD protocol (Grimm et al., 2020, *JASSS* 23(2):7). This table is the reference documentation for every field of `all_data.cst` in the Data S2 files and for every parameter of the simulation code. --- ### Videos All videos are MPEG-4 (H.264) renderings generated from the datasets above with `gen_video_tracks.m` and `gen_video_tracks_vel_acc.m`. **`Video_S1.mp4`: swarming flight of an initiator mosquito.** Three-dimensional and two-dimensional views of the flight of the initiator — the first *An. coluzzii* male to start swarming — recorded above a 40 x 40 cm black swarm marker. Source data: the individual with `was_initiator = true` in Data S1. **`Video_S2.mp4`: flight kinematics of a single swarming mosquito.** Three-dimensional position, flight speed, acceleration and angular speed of one swarming *An. coluzzii* male recorded above the same marker. This individual was swarming together with conspecifics. Source data: Data S1. **`Video_S3.mp4`: swarming flight of multiple mosquitoes.** Three-dimensional and two-dimensional views of the simultaneous flight of multiple *An. coluzzii* males recorded above the same marker. Source data: Data S1. **`Video_S4.mp4`: swarming flight of ten simulated mosquitoes.** Three-dimensional and two-dimensional views of ten agents simulated with the agent-based model including collision avoidance, above a virtual 40 x 40 cm marker. Source data: one replicate of `all_data-abm-along_y_rand-null-threshold-any-alpha_x_percent-threshold_alpha_x-avoid_collisions.mat` in Data S2. --- ## Code/software Code is hosted on Zenodo, in the compressed file Code_S1.zip. **Reading the data.** The `.mat` files are MATLAB v7.3 (HDF5). In MATLAB (R2022b or later recommended): `load('all_data-exp1-coluzzii_males-selection_swarming.mat')` returns the structure `all_data`. In Python, use `h5py` (not `scipy.io.loadmat`), remembering that HDF5 transposes MATLAB arrays and stores MATLAB character arrays as `uint16` character codes: ```python import h5py, numpy as np f = h5py.File('all_data-exp1-coluzzii_males-selection_swarming.mat', 'r') track = f['all_data/exp/exp20221007_183000/tracks/object/obj1'] x, y, z = (np.array(track[k]).ravel() for k in ('x', 'y', 'z')) t = np.array(track['time']).ravel() ``` **Analysis and model code (S1 Code).** All original MATLAB code written to run the agent-based model, analyse both datasets and generate the figures of the manuscript is provided as **S1 Code** with the article. It was developed and run in MATLAB R2022b. Its main entry points are: * `_Code/main.m` — full analysis pipeline for a dataset: loads an `all_data` file, computes the derived kinematic and visual variables (`gen_all_data_tracks.m`), the swarm-level metrics (`gen_all_data_swarms.m`) and the spatial heat maps (`gen_all_data_heatmaps.m`), and produces the corresponding figures (`gen_all_fig_swarms.m`, `gen_all_fig_heatmaps.m`). Set `exp_num = 1` to analyse Data S1, or `exp_num = 0` together with the appropriate `sim_name` to analyse one of the Data S2 files. * `_Code/Agent based model/main_all_abm.m` — runs the agent-based model for the listed configurations (20 replicates x 10 agents x 120 s each) and writes the `all_data-abm-*.mat` files deposited as Data S2. `main_abm.m` runs a single simulation, and `boids/swarming_saccades/update_boids.m` implements the behavioural rules themselves (saccade triggering, turn-angle sampling, collision avoidance). * `_Code/get_alpha_xy.m` and `_Code/Agent based model/get_alpha_x.m` — compute the marker viewing angles α_x and α_y from a 3D position, assuming coordinates centred on the marker. * `_Code/gen_video_tracks.m`, `_Code/gen_video_tracks_vel_acc.m` — generate Videos S1-S4. To reproduce the analysis: unzip the data, place the `.mat` file in the `_Data` folder, set `data_path` at the top of `main.m` to that location, and run `main.m`; figures are written to the `_Figures` folder. Note that `main.m` also calls a small number of general-purpose helper functions from the authors' shared MATLAB library (e.g. `derivative`, `save_figures`, `gen_all_data_split_tracks`, `gen_all_data_tracks_subset`, `angle_btw_vectors`), which are not part of S1 Code; these can be obtained from the corresponding author."}],"geoLocations":[],"fundingReferences":[{"funderIdentifierType":"ROR","funderName":"International Human Frontier Science Program Organization","funderIdentifier":"https://ror.org/02ebx7v45","awardNumber":"RGP0044/2021"},{"funderIdentifierType":"ROR","funderName":"Institut de Recherche pour le Développement","funderIdentifier":"https://ror.org/05q3vnk25"},{"funderIdentifierType":"ROR","funderName":"Wageningen University \u0026 Research","funderIdentifier":"https://ror.org/04qw24q55"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.69p8cz9bw","contentUrl":null,"metadataVersion":0,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-21T17:34:30Z","registered":"2026-08-21T17:34:31Z","published":null,"updated":"2026-08-21T17:34:31Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.0k6djhbh8","type":"dois","attributes":{"doi":"10.5061/dryad.0k6djhbh8","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["University of North Carolina at Chapel Hill"],"name":"Sockman, Keith","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0002-3846-5386"}]},{"nameType":"Personal","affiliation":["University of North Carolina at Chapel Hill"],"name":"Frasson, Nicolas","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of North Carolina at Chapel Hill"],"name":"Reinhardt, Emma","nameIdentifiers":[]}],"titles":[{"title":"Data from: Effect of egg size on size of subsequent eggs: experimental evidence for performance-based feedback within the clutch of a wild bird"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"subject":"Egg laying"},{"subject":"life-history"},{"subject":"Lincoln's sparrow"},{"subject":"i\u0026gt;Melospiza lincolnii"},{"subject":"reproductive effort"},{"subject":"unpredictable environment"},{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Animal and dairy science","subjectScheme":"fos"},{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Natural sciences","subjectScheme":"fos"}],"contributors":[],"dates":[{"date":"2026-08-18T19:57:40Z","dateType":"Created"},{"date":"2026-08-18T19:57:40Z","dateType":"Submitted"},{"date":"2026-08-21T00:00:00Z","dateType":"Issued"},{"date":"2026-08-21T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[{"relationType":"IsCitedBy","relatedIdentifier":"10.1086/685881","relatedIdentifierType":"DOI"}],"relatedItems":[],"sizes":["26000 bytes"],"formats":[],"version":"3","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"An egg's size should reflect the net benefit of its production, which\n itself should vary with the environment. Although the environment is often\n complex, performance-based feedback from the size of an egg could signal\n the cumulative effects of the environment and enable females to optimize\n size of subsequent eggs, as a recent model and experiment support. That\n experiment did not rule out the signaling role of size-correlates,\n however. Therefore, using free-ranging Lincoln's sparrows (Melospiza\n lincolnii), we conducted an experiment in which we substituted the\n first-laid egg—on the day it was laid—with an artificial egg from one of\n two groups that were identical except in size. We discovered that,\n compared to a small substitute, a large substitute yielded a larger\n increase in size from a previous egg to a subsequent egg (egg 4 in the\n clutch). This was consistent with the model and with the hypothesis that\n the size specifically of an initial egg can affect the size of\n subsequently laid eggs. Under unpredictably variable environmental\n conditions, such performance-based feedback may contribute to the\n optimization of not only egg size but a diversity of organismal processes."},{"descriptionType":"TechnicalInfo","description":"# Data from: Effect of egg size on size of subsequent eggs: experimental\n evidence for performance-based feedback within the clutch of a wild bird\n Dataset DOI:\n [10.5061/dryad.0k6djhbh8](https://doi.org/10.5061/dryad.0k6djhbh8) ##\n Description of the data and file structure ### Files and variables ####\n File: Data_for_Submission.xlsx * **Sheet 1: **describes\n the variables in the other sheets * **Sheet 2**: raw data for experimental\n nests only * **Sheet** 3: raw data for all natural (non-experimental)\n nests analyzed in the study * **Sheet 4**: raw data for natural nests in\n the years 2023 and 2024 only ##### Variables | Variable | Description | |\n :------ | :------------------------------------------------------- | |\n year | Year in which data were collected | | nestid | Unique identifier of\n individual nest | | lgtreat | Experimental treatment: 0=small, 1=large | |\n egg | Laying order of individual egg within an individual nest | | length\n | Length of egg in mm | | width | Width of egg in mm |"}],"geoLocations":[],"fundingReferences":[],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.0k6djhbh8","contentUrl":null,"metadataVersion":0,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-21T15:27:01Z","registered":"2026-08-21T15:27:02Z","published":null,"updated":"2026-08-21T15:27:02Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.4xgxd25t1","type":"dois","attributes":{"doi":"10.5061/dryad.4xgxd25t1","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["University of California, Los Angeles"],"name":"Daversa, David","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0002-8984-8897"}]},{"nameType":"Personal","affiliation":["Yosemite National Park"],"name":"Grasso, Rob","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of California, Los Angeles"],"name":"Posta, Molly","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of California, Los Angeles"],"name":"Lloyd-Smith, James","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of California, Los Angeles"],"name":"Shaffer, H. Bradley","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0002-5795-9242"}]}],"titles":[{"title":"Data and code from: \u003cem\u003eBatrachochytrium dendrobatidis\u003c/em\u003e infections proliferate during amphibian terrestrial dormancy"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Ecology","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Infectious diseases","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Amphibians","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Biology and life sciences","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Epidemiology","subjectScheme":"PLOS Subject Area Thesaurus"}],"contributors":[],"dates":[{"date":"2026-08-12T20:20:03Z","dateType":"Created"},{"date":"2026-08-12T20:20:31Z","dateType":"Submitted"},{"date":"2026-08-21T00:00:00Z","dateType":"Issued"},{"date":"2026-08-21T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[{"relationType":"IsCitedBy","relatedIdentifier":"10.1101/2025.07.07.663599","relatedIdentifierType":"DOI"}],"relatedItems":[],"sizes":["421533 bytes"],"formats":[],"version":"5","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"The spread and impact of wildlife pathogens is often seasonal, and\n identifying the season(s) of highest impact can be critical to\n biodiversity and public health management. We report new evidence that\n winter host dormancy, a period generally neglected in terms of pathogen\n seasonal dynamics, strongly promotes the proliferation of the globally\n threatening amphibian fungal pathogen, Batrachochytrium dendrobatidis\n (Bd), in an endangered high-elevation anuran host species. Surveillance of\n Bd in Yosemite toads (Anaxyrus canorus) during their first year of life\n showed that Bd prevalence and intensity increased fourfold during their\n first terrestrial winter dormancy. High prevalence and intensity of Bd\n infections following winter dormancy were observed in terrestrial\n first-year juveniles across multiple sites and three consecutive years,\n and their Bd prevalence was much higher than that observed in aquatically\n breeding adults during the same timeframe. The infections persisted into\n the active season. We integrated these results into a framework\n summarizing possible mechanisms of overwinter Bd maintenance. We recommend\n a reconsideration of seasonal Bd proliferation (i.e., increases in\n prevalence and/or intensity) and highlight a mechanism of interannual Bd\n maintenance in dormant terrestrial juveniles. The importance of host\n dormancy to pathogen persistence and seasonal infection spread may be far\n greater than previously thought."},{"descriptionType":"Methods","description":"Definitions Maintenance - the continued\n presence of the pathogen at the host population scale.\u003cbr\u003e\n Prevalence - The proportion of samples that exhibit infection loads\n exceeding our limits of detection (100 ITS1 copies).\u003cbr\u003e\n Proliferation – increases in the observed prevalence and/or intensity of\n infection.\u003cbr\u003e Transmission – the acquisition of new infections by\n susceptible hosts, either from the environment or from other infected\n hosts. Study site Our surveys\n occurred from 2021-2023 and spanned six meadow complexes in the Tioga Pass\n region of Yosemite National Park: Lower Gaylor Meadow, Dana Meadow,\n Delaney Meadow, Gaylor Forested Meadow, Mono Pass Meadow, and Spillway\n Meadow (Figure 1b). All the sites are open, high-elevation grasslands\n surrounded by conifer forest and include shallow ephemeral waterbodies\n that comprise aquatic breeding habitat for toads and terrestrial\n non-breeding habitat (willows, mammal burrows, rock piles) (Karlstrom,\n 1962; Thomson et al., 2016). Sample\n collection All field sampling procedures were conducted\n under approved protocols from the Institutional Animal Use and Care\n Committee (protocol ID: CA_YOSE_Herps_2021.A3), the US Fish and Wildlife\n Service (permit # TE86906B-1), and the US National Park Service (permit#\n YOSE-2022-SCI-0062). We carried out two sets of surveys\n to quantify temporal and life-stage variation in \u003cem\u003eBd\u003c/em\u003e\n prevalence and intensity. Survey 1 consisted of repeated cross-sectional\n surveys of recently metamorphosed toads (metamorphs) throughout their\n first year of life to characterize \u003cem\u003eBd\u003c/em\u003e dynamics before\n and after dormancy (Figure 1c). Metamorphs that emerged at the same time\n and location form distinct cohort groups that enter and emerge from\n dormancy together, which enabled reliable resurveying of the same cohorts\n (but not necessarily the same individuals). Surveys spanned the full\n active season of metamorphs in 2021 and the beginning of the active season\n in 2022 and 2023 when individuals were emerging from their first winter\n dormancy in the late spring (May/June) (Figure 1c). We hand-captured up to\n fifty metamorphs per meadow on three different occasions during their\n first active season in 2021 before brumation: immediately after\n metamorphosis (sample 1, July), one month post-metamorphosis (sample 2,\n August), and two months post-metamorphosis (sample 3, September) (Figure\n 1c). The final sampling event in mid-September coincided with cooling\n temperatures and daylength shortening that precedes brumation (Calatayud\n et al., 2021; Sherman, 1980). We collected individual skin samples for\n \u003cem\u003eBd\u003c/em\u003e detection using a sterile swab (MWE 113 medical\n wire, UK) that was stroked over the toad’s belly 30 times (thighs and feet\n webbing were too small to sample). We located the same\n cohorts as they were emerging from dormancy in 2022 as first-year\n juveniles (Figure 1a, c). To minimize the possibility of emerging toads\n contracting \u003cem\u003eBd\u003c/em\u003e infections after dormancy but before\n sampling, we scanned each meadow for toads every 2-3 days starting on May\n 17\u003csup\u003eth\u003c/sup\u003e, 2022, before the first emergence of any bare\n ground from the season’s snowpack (Figure S5c). We did not observe any\n toads in our initial surveys, suggesting that our surveillance began when\n individuals were still dormant. Aerial imagery supports this conclusion\n with evidence of extensive snowpack on our initial survey weeks (Figure\n S5). We first observed metamorphs in Dana Meadow on 25 May 2022 as the\n snowpack was just beginning to recede (Figure S5d). As the cohorts emerged\n (late May-early June, depending on the site), we captured and swabbed\n individuals at three time points:  immediately after emergence (sample 4,\n N = up to 50/meadow), two weeks after emergence (sample 5, N = 30/meadow),\n and four weeks after emergence (sample 6, N = 30/meadow).  Cohorts could\n not be relocated in Delaney Meadow in 2022, and so only five meadows were\n included for samples 4-6. Survey 2 included all life\n stages, with the goal of comparing early-season infection levels across\n toad life stages starting when they emerged from brumation (25 May 2021 –\n 16 Sept, 2021). We focused early season sampling at Lower Gaylor Meadow in\n 2021 because it is a central population with consistently high abundances\n of breeding toads (Maier et al., 2022). Again, we began scanning the\n meadow when the snowpack was still fully intact (Figure S5a-b), and our\n initial surveys did not detect any toad activity. We first observed toads\n at Lower Gaylor Meadow on 25 May, 2021 as the snowpack was receding\n (Figure S5b), at which point we initiated weekly surveys. We walked the\n entire meadow and captured and swabbed male and female adult toads that\n were not in amplexus as well as first-year (2020 cohort) and second-year\n juveniles. Swabbing of juveniles and adults involved rubbing the belly 15\n times, each thigh 5 times, and each hind foot webbing surface 5 times,\n following Vredenburg \u003cem\u003eet al\u003c/em\u003e. (2010). We used a cutoff\n of 23mm snout-to-vent length to distinguish first-year (\u0026lt; 23 mm)\n from second-year (\u0026gt; 23 mm) juveniles. This cutoff was based on\n published literature (Karlstrom, 1962) and our data from the cohort\n surveys (the largest first-year juvenile captured in the first two weeks\n post-dormancy in 2022 was 22mm SVL). We also collected and swabbed thirty\n tadpoles (Gosner stage 30-40) from Lower Gaylor Meadow in June. Swabbing\n of tadpoles included 30 strokes of the labial tooth rows and mandibles,\n which are the only keratinized structures of tadpoles, and therefore the\n only potential \u003cem\u003eBd\u003c/em\u003e infection sites (Vredenburg et al.,\n 2010). Starting in the third sampling week, we expanded our surveillance\n to the five other focal meadows to assess temporal \u003cem\u003eBd\u003c/em\u003e\n dynamics more broadly. Sample Processing\n We screened swabs for \u003cem\u003eBd\u003c/em\u003e using\n well-established DNA extraction and qPCR procedures (Boyle et al., 2004).\n Briefly, we used Prepman-based DNA extraction procedures that involved\n homogenization with a bead-beater and centrifugation (Boyle et al., 2004).\n We prepared 1:10 DNA extract dilutions and quantified the number of copies\n of the ITS1 region of the \u003cem\u003eBd\u003c/em\u003e genome in 5 µL dilution\n aliquots using real-time Taqman qPCR assays (Boyle et al., 2004). The\n assays included negative controls and five concentration standards serving\n as positive controls (concentrations: 10\u003csup\u003e2\u003c/sup\u003e,\n 10\u003csup\u003e3\u003c/sup\u003e, 10\u003csup\u003e4\u003c/sup\u003e,\n 10\u003csup\u003e5\u003c/sup\u003e, 10\u003csup\u003e6\u003c/sup\u003e ITS1 copies).\n Standards were prepared by the Briggs Lab at the University of California,\n Santa Barbara. Swabs were processed in singlicate due to the documented\n efficacy of \u003cem\u003eBd\u003c/em\u003e detection from single samples (Joseph\n \u0026amp; Knapp, 2018; Kriger et al., 2006; Vredenburg et al., 2010), and\n any ambiguous outputs were rerun. We defined \u003cem\u003eBd\u003c/em\u003e\n intensity (also known as \u003cem\u003eBd\u003c/em\u003e load) as the number of\n ITS1 copies. A sample was considered \u003cem\u003eBd\u003c/em\u003e-positive when\n the ITS1 copy number was 100 or greater, matching our lowest-level, most\n sensitive positive controls. We also reran all analyses considering all\n samples with non-zero \u003cem\u003eBd\u003c/em\u003e loads as\n \u003cem\u003eBd\u003c/em\u003e-positive; this changed the classification of just\n 31 samples, and the results were essentially identical (see the\n Supplementary Material for the results). Data\n analysis We evaluated the extent to which\n \u003cem\u003eBd\u003c/em\u003e prevalence and intensity varied by season and host\n life stage using generalized linear models (GLMs) and mixed models\n (GLMMs). Prevalence refers to ‘observed prevalence’ based on our detection\n assays (see definition in Section 2.1), meaning that changes in prevalence\n could arise from the acquisition of new infections (transmission) or from\n newly detectable infections. For models of \u003cem\u003eBd\u003c/em\u003e\n prevalence, we used a binary infection status (0 = uninfected, 1 =\n infected) of individuals as the response variable and a binomial error\n structure. For models of \u003cem\u003eBd\u003c/em\u003e intensity, we focused\n exclusively on \u003cem\u003eBd\u003c/em\u003e-positive samples, used\n log-transformed \u003cem\u003eBd\u003c/em\u003e loads as the response variable,\n and a Gaussian error structure.  To evaluate\n within-season \u003cem\u003eBd\u003c/em\u003e dynamics in the 2021 metamorph\n cohorts, we ran GLMMs with meadow ID as a random effect to account for\n site-level variation and the sampling event (1, 2, 3) as a fixed effect to\n measure changes over time. We used the same modelling structure to compare\n \u003cem\u003eBd\u003c/em\u003e prevalence and intensity in cohorts before vs.\n after dormancy. Specifically, we compared infection data from the last\n cohort sample in 2021 (sample 3, Figure 1c) with the infection data from\n the first cohort sample in 2022 (sample 4, Figure 1c), using GLMMs with\n the sampling event as a fixed effect and meadow ID as a random\n effect.  We ran GLMs to compare early-season\n \u003cem\u003eBd\u003c/em\u003e prevalence and intensity across life stages at\n Lower Gaylor Meadow. We ran GLMs with life stage as a fixed effect with\n four levels: first-year juvenile (2020 cohort), second-year juvenile,\n adult male, and adult female. For all analyses, we\n tested the influence of sampling event and life stage on\n \u003cem\u003eBd\u003c/em\u003e prevalence and intensity by comparing the fit of\n models containing those factors to models omitting them, using likelihood\n ratio tests (MASS package in R) following either a chi-squared\n distribution in the case of GLMMs and binomial GLMs, or an\n \u003cem\u003eF\u003c/em\u003e distribution for GLMs with Gaussian error\n structures. When a significant effect of a factor was detected, we\n examined differences between factor levels with Tukey’s HSD tests, using\n the emmeans function in R (emmeans package). Ninety-five percent\n confidence intervals for prevalence estimates were derived from an exact\n binomial test based on Clopper \u0026amp; Pearson (1934), using the\n binom.test function in R. Figures were created using the ggplot2 package\n in R, with graphics subsequently added in Adobe Illustrator."},{"descriptionType":"TechnicalInfo","description":"# Data and code from: *Batrachochytrium dendrobatidis* infections\n proliferate during amphibian terrestrial dormancy Dataset DOI:\n [10.5061/dryad.4xgxd25t1](https://doi.org/10.5061/dryad.4xgxd25t1) ##\n Description of the data and file structure The following script below\n complements a paper reporting on Batrachochytrium dendrobatidis (Bd)\n surveillence in Yosemite toad (Anaxyrus canorus) populations in 2021 \u0026amp;\n 2022 in the Tioga Pass area of Yosemite National Park. Data were collected\n by Dave Daversa and Molly Posta. Bd samples were processed with qPCR at\n the Sierra Nevada Aqutic Research Laboratory (SNARL). ### Files and\n variables #### File: YOSE_canorus_2026_FINAL.Rmd **Description:** R script\n for data management and analysis #### File: YOSE_Acanorus_Db_rev.csv\n **Description:**  Original raw data from our field surveillance #####\n Variables * entry_id: a unique identifier for the row/entry * species:\n which species to which the datum pertains (A canorus, P regilla) *\n swab_id: Unique identification label give to swab used to sample for Bd on\n toads * sex_lifestage: the life stage of the individual (larva, metamorph,\n subadult, adult) and, in the case of adults, the sex (Male vs. Female) *\n svl_gosner: the snout-to-vent length of the individual in millimeters *\n wt: weight of individual in grams * habitat: whether animals were aquatic\n or terrestrial when captured * date: date of sample collection *\n sample_no: the sampling event in which the sample was collected * week:\n which week of the study period the sample was collected * year: which year\n the sample was collected * meadow_name: name of breeding meadow where\n sample was collected * meadow_ID: unique identification as assigned by the\n US Geological Survey * macrohabitat: substrates or features of the\n macrohabitat (anecdotes only) #### File: YOSE_Daversa_qPCR_210529_rev.csv\n **Description:** 2021 qPCR data for analyses of infection ##### Variables\n * swab_id - unique identifier given to the skin swab sample collected from\n a toad individual * date_extracted: day, month, and year when the DNA\n extractions of the sample were carried out * plate: PCR plate in which the\n sample was processed through qPCR * date_qPCR: date that the sample was\n run through qPCR processing * quant_cycle: the Ct value, or the period in\n the qPCR cycle when the detection threshold is crossed * start_quant: the\n initial amount of target DNA before amplification * bd_load: quantitative\n estimate of Bd infection strength, expressed in ITS1 copies #### File:\n YOSE_Daversa_qPCR_230220_rev.csv **Description:** 2022 qPCR data for\n analyses of infection ##### Variables * swab_id - unique identifier given\n to the skin swab sample collected from a toad individual * plate: PCR\n plate in which the sample was processed through qPCR * quant_cycle: the Ct\n value, or the period in the qPCR cycle when the detection threshold is\n crossed * start_quant: the initial amount of target DNA before\n amplification * bd_load: quantitative estimate of Bd infection strength,\n expressed in ITS1 copies #### File: YOSE_Daversa_qPCR_240802_rev.csv\n **Description:** 2023 qPCR data for analyses of infection ##### Variables\n * swab_id - unique identifier given to the skin swab sample collected from\n a toad individual * plate #: PCR plate in which the sample was processed\n through qPCR * quant_cycle: the Ct value, or the period in the qPCR cycle\n when the detection threshold is crossed * start_quant: the initial amount\n of target DNA before amplification * bd_load: quantitative estimate of Bd\n infection strength, expressed in ITS1 copies * notes: notes from the lab\n manager about any samples that were re-run #### File:\n DanaMeadows_DailyData_20210501_20220701.csv **Description:** Mean, max,\n and min ambient temperature data collected at a Dana Meadows weather\n station over the study period. ##### Variables * station_id: Unique\n weather station identification from which data were collected * data_time:\n Date and time of day the data were collected * swe_in: snow-water\n equivalent, recorded in inches * min_temp_f: minimum temperature reading,\n in Fahrenheit * max_temp_f: maximum temperature reading, in Fahrenheit *\n avg_temp_f: average temperature reading, in Fahrenheit * min_temp_c:\n minimum temperature reading, in Celcius * max_temp_c: maximum temperature\n reading, in Celcius * avg_temp_c: average temperature reading, in Celcius\n ## Code/software R programming software is required. We used R version\n 4.3.1 (2023-06-16) -- \"Beagle Scouts\". The following R libraries\n are also required: library(ggplot2): For making plots library(ggsci): for\n customizing colors of plots library(grid): for combining plots into a\n single figure library(performance): for the check_model funcation that\n assesses model performance library(lme4): for mixed modelling\n library(MASS): For statistical analysis library(dplyr): for data\n organization library(MuMIn): for r.squaredGLMM library(emmeans): for post\n hoc tests library(lmerTest): for ranova function that tests significance\n of random effect terms library(binom): for calculating confidence\n intervals of binomial data The full annotated workflow is documented in\n the .Rmd file (YOSE canorus_2026_FINAL.Rmd) ## Access information Other\n publicly accessible locations of the data: * N/A Data was derived from the\n following sources: * N/A"}],"geoLocations":[],"fundingReferences":[{"funderName":"UCLA La Kretz Center for California Conservation Science","awardTitle":"La Kretz Postdoctoral Fellowship"},{"funderName":"Yosemite Conservancy"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.4xgxd25t1","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-21T08:48:25Z","registered":"2026-08-21T08:48:26Z","published":null,"updated":"2026-08-21T14:41:15Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.tht76hfdr","type":"dois","attributes":{"doi":"10.5061/dryad.tht76hfdr","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["Michigan State University"],"name":"Collins, Erin","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0002-6097-8479"}]},{"nameType":"Personal","affiliation":["Wild Salmon Center"],"name":"Thompson, Tasha","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0003-3482-2101"}]},{"nameType":"Personal","affiliation":["California Department of Water Resources"],"name":"Goertler, Pascale","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["California Department of Water Resources"],"name":"Baerwald, Melinda","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Michigan State University"],"name":"Meek, Mariah","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0002-3219-4888"}]}],"titles":[{"title":"Data from: Diversity under pressure: Long-term genetic decline in California’s Central Valley Chinook salmon ESA listed runs"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Earth and related environmental sciences","subjectScheme":"fos"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Genetics","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Salmon","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Effective population size","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Conservation genetics","subjectScheme":"PLOS Subject Area Thesaurus"}],"contributors":[],"dates":[{"date":"2026-04-08T12:15:12Z","dateType":"Created"},{"date":"2026-08-04T07:15:11Z","dateType":"Submitted"},{"date":"2026-08-05T00:00:00Z","dateType":"Issued"},{"date":"2026-08-05T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[{"relationType":"IsSupplementedBy","relatedIdentifier":"10.5061/dryad.280gb5mxx","relatedIdentifierType":"DOI"},{"relationType":"IsCitedBy","relatedIdentifier":"10.1002/ece3.74176","relatedIdentifierType":"DOI"}],"relatedItems":[],"sizes":["126910251 bytes"],"formats":[],"version":"2","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Intraspecific diversity enables populations to persist under stochastic\n and extreme environmental conditions. One system that is representative of\n notable intraspecific diversity is Chinook salmon (Oncorhynchus\n tshawytscha) in the Central Valley of California, USA. It is the only\n place in the species range where four distinct migration timings (Winter,\n Spring, Fall, Late-fall) co-occur. These populations are declining, with\n Winter run listed as Endangered and Spring run as Threatened under the\n Endangered Species Act (ESA). To quantify temporal changes in genetic\n diversity of the different runs, we genotyped outmigrating juveniles using\n RAD-sequencing across 20+ years of annual sampling. Tajima's D\n revealed significant shifts in neutral genetic variation over time in\n Spring and Winter runs. Effective population size (Ne) declined in all\n listed populations, and the historical Ne estimates indicate severe\n plummets in genetic diversity 25-50 generations ago. Our results\n demonstrate how anthropogenic forces have eroded the genetic diversity of\n ESA listed Chinook salmon populations over the past century. Moreover,\n this study demonstrates how an organism with genetically-based life\n history variation (migration timing), traditionally advantageous under\n natural environmental variability, struggles to persist when faced with\n anthropogenically altered habitat and changing\n climate.  "},{"descriptionType":"TechnicalInfo","description":"# Data from: Diversity under pressure: Long-term genetic decline in\n California’s Central Valley Chinook salmon ESA listed runs Dataset DOI:\n [10.5061/dryad.tht76hfdr](https://doi.org/10.5061/dryad.tht76hfdr) ##\n Description of the data and file structure All Chinook salmon in the\n Central Valley of California, USA (CV) are spawned in tributaries to the\n Sacramento and San Joaquin Rivers and migrate out to the ocean as\n juveniles through the CV Delta, passing by Chipps Island. We utilized\n archived samples collected from Chipps Island during the juvenile\n outmigration over a twenty-year period (1996-2018) to evaluate\n intraspecific-level diversity in listed CV Chinook salmon (see Fig. 1 in\n the associated article). The subpopulations of interest are the Mill\n Creek, Deer Creek, and Butte Creek populations within the Spring run (Fig.\n 1). The outmigrating juvenile Chinook salmon collected for this study were\n sampled from the territories of Miwok, Patwin, Me-Wuk (Bay Miwok), and the\n Confederated Villages of Lisjan and the individuals collected were spawned\n across the Central Valley, which covers the territory of over 100 tribes\n (Native Land Digital). To evaluate genetic diversity and effective\n population sizes in CV Spring and Winter run Chinook salmon over the last\n 20 years, we utilized data from a previously published study (Thompson et\n al. 2024). This dataset genetically sequenced approximately 622 Chinook\n salmon juveniles sampled while outmigrating past Chipps Island in the\n lower Sacramento River Delta (Fig. 1) using a RAD-sequencing protocol (Ali\n et al., 2016), then assigned each sample to one of the major demographic\n groups in the CV (Winter, Spring, Fall, and Late-Fall runs), as well as to\n subpopulations within the Spring-run (Mill/Deer Creek, and Butte Creek;\n Fig. 1). In total, 325 samples confidently assigned to the Spring-run (159\n Mill/Deer and 166 Butte Creek) and 220 samples confidently assigned to the\n Winter run, and these samples were included in the current study. ###\n Files and variables #### File:\n cv_juv_1kaln_cutoff.snp_panel.IBSonly.sites.ibs **Description:**  The\n first column is the chromosome, the second column is the position, the\n third column is the major allele, the fourth column is the minor allele,\n and the rest of the columns are the genotype likelihood data for each\n individual. The allele calling samples a single allele instead of calling\n genotypes (it picks a single read at the position and takes the allele\n from that read). Nothing is homozygous or heterozygous, it's just a\n single allele. Effectively, it down samples everything to 1x at each\n position, which greatly reduces technical artifacts from differences in\n coverage (something that doesn't matter much in relatively high\n coverage data because genotype calls are all pretty good once you get a\n little over about 5x, but can matter a lot for low coverage data that have\n a lot of uncertainty in genotype calls). 1 corresponds to the major\n allele, 0 to the minor, and -1 to missing. ## Code/software Sequencing\n data alignment and initial processing is described in Thompson et al.\n (2024). Briefly, raw fastq files were mapped to the Chinook salmon\n reference genome (Otsh_v2.0; GCA_018296145.1; Christensen et al., 2018)\n with bwa-mem (Vasimuddin et al. 2019), and samtools (Danecek et al. 2021)\n was used to quality filter the mapped reads. Only reads that mapped\n uniquely, had mapping qualities \u0026gt; 30, base qualities \u0026gt; 30,\n properly-paired reads, and \u0026gt; 1,000 final aligned reads were retained.\n Thompson et al. (2024) identified a panel of single nucleotide\n polymorphisms (SNPs) present in at least 50% of samples from that study\n with minor allele frequencies \u0026gt; 0.05, and that panel of SNPs was used\n for analyses in this current study. We filtered the samples by missingness\n by population according to Supplemental Table 2 in Thompson et al. (2024).\n To get the most information from the 20 years of samples with varying\n quality, Thompson et al. (2024) called a single allele at each SNP locus\n instead of calling a genotype."}],"geoLocations":[],"fundingReferences":[{"funderIdentifierType":"ROR","funderName":"California Department of Water Resources","funderIdentifier":"https://ror.org/04w9m8m13"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.tht76hfdr","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":1,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-05T22:31:41Z","registered":"2026-08-05T22:31:42Z","published":null,"updated":"2026-08-21T14:36:08Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.zs7h44jpg","type":"dois","attributes":{"doi":"10.5061/dryad.zs7h44jpg","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["The Ohio State University"],"name":"Boot, Matthew R","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0009-0002-5939-9610"}]},{"nameType":"Personal","affiliation":["The Ohio State University"],"name":"Sozanski, Kyle S","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Tiffin University"],"name":"Davis, Mazie","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["The Ohio State University"],"name":"Hamilton, Ian M","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0002-1198-3431"}]},{"nameType":"Personal","affiliation":["The Ohio State University"],"name":"Adams, Rachelle MM","nameIdentifiers":[]}],"titles":[{"title":"Data and code from: Guest-ant social parasites avoid conflict with social hosts using venom signaling"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Animal behavior","subjectScheme":"PLOS Subject Area Thesaurus"},{"subject":"communication"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Symbiosis","subjectScheme":"PLOS Subject Area Thesaurus"},{"subject":"Social parasitism"},{"subject":"social antagonism"},{"subject":"Interdependence"},{"subject":"host manipulation"},{"subject":"parasite signal"},{"subject":"venom signaling"},{"subject":"Formicidae"},{"subject":"fungus-farming ants"},{"subject":"social insects"},{"subject":"chemical communication"}],"contributors":[{"name":"Smithsonian Tropical Research Institute","contributorType":"Sponsor","affiliation":[],"nameIdentifiers":[]}],"dates":[{"date":"2026-05-20T17:56:03Z","dateType":"Created"},{"date":"2026-06-12T14:34:17Z","dateType":"Submitted"},{"date":"2026-07-02T00:00:00Z","dateType":"Issued"},{"date":"2026-07-02T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[{"relationType":"IsCitedBy","relatedIdentifier":"10.1371/journal.pone.0345143","relatedIdentifierType":"DOI"}],"relatedItems":[],"sizes":["218343 bytes"],"formats":[],"version":"3","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Megalomyrmex symmetochus “guest ant” parasites cohabit, long-term, within\n the nest of a single colony of Sericomyrmex amabilis, a fungus-farming\n ant. Although M. symmetochus exploit their stingless hosts for resources,\n they directly rely on their host colony for survival and reproduction, and\n thus, have been demonstrated to protect their host colony from threats to\n the shared nest using venom weaponry. We use behavioral observation of\n staged host-parasite conflict, lethal-dose assays, and direct venom\n measurements to show that M. symmetochus uses conspicuous and costly\n displays of venom during interactions with hosts. Megalomyrmex symmetochus\n is observed to dispense alkaloid venom directly through stinging, but more\n frequently, indirectly through airborne venom dispersal of volatile\n pyrrolizidine alkaloids. We further demonstrate that indirect venom use\n can alter interaction outcomes if hosts switch to non-aggressive\n behavioral tactics. This data repository contains both the data and code\n necessary to reproduce all analyses from the publication: Boot et al.,\n 2026, Guest-ant social parasites avoid conflict with social hosts using\n venom signaling. The dataset is structured into two primary folders\n (Behavior, Venom) from which separate analyses are run. All code and data\n necessary to run the code is provided in the respective folder, as well as\n separate README files which enumerate the data files/structures contained\n within each of the two folders. For methods, please see the main\n publication."},{"descriptionType":"TechnicalInfo","description":"# Data and code from: Guest-ant social parasites avoid conflict with\n social hosts using venom signaling AUTHOR(S): Matthew Boot \u0026amp; Kyle\n Sozanski DATE: 5/15/2026 [Access this dataset on Dryad]\n ([https://doi.org/10.5061/dryad.zs7h44jpg](https://doi.org/10.5061/dryad.zs7h44jpg) ) ## EXPERIMENTAL CONTEXT This data repository contains data and codebase from the publication: Boot *et al.*, 2026. Guest-ant social parasites avoid conflict with social hosts using venom signaling. Several datasets and analyses are provided, broken down between 1) behavioral data/analyses and 2) venom data/analyses. ### Background \u0026amp; motivation: As context-dependent mutualists, the two focal species generally cohabit peacefully within shared nests for many years. However, conflict can also occur in these mixed interspecific nests and has been observed at certain points in the host colonies’ lifecycle. Parasite venom is thought to play a role in host colony infiltration and integration as it is toxic and can have communicative properties. *Sericomyrmex amabilis* hosts lack venom and are stingless but have strong mandibles which they use when in conflict. *Megalomyrmex symmetochus* social parasites have alkaloid-based venom stings, and these parasites have been observed using conspicuous, indirect venom display behavior—which could play a role in maintaining nest cohesion, even outside of overt conflict. The provided data and code were used to investigate parasite venom use, and if it is used in signaling during host-parasite behavioral conflict. In addition, we also characterized the venom expenditure necessary for parasites to neutralize hosts. ### Behavioral interactions: To analyze host-parasite conflict, we observed staged conflict between naive *S. amabilis* hosts and *M. symmetochus* social parasites (5 hosts to a single social parasite) at the group level. When host ants have not lived with a parasite, their initial behaviors are aggressive. By staging these altercations, we are able to induce interspecific conflict. Conflict-related host/parasite behaviors were scored and recorded during video playback of arena trials, along with certain interaction metadata (e.g., video ID, ant colony ID, etc.) for each behavioral observation. Using a generalized linear mixed modeling framework to account for multiple observations at the colony and/or video level, we used these observations to assess: a) a comparison of direct and indirect venom use by parasites, b) the sequence of associated conflict behaviors between hosts and parasites, c) the relationship between parasite venom use and host behavior, d) escalation of aggression between hosts and parasites, e) parasite behavior following host biting, f) parasite venom use relative to conflict escalation, g) host submission following venom use, and h) interaction termination. Through our analyses, host ants were found to be the primary aggressors, while parasites tended to avoid conflict and conflict escalation. Parasites were observed to dispense alkaloid venom directly through stinging, but more frequently, indirectly via airborne venom dispersal behavior earlier in interactions. We further demonstrated that indirect venom use can alter interaction outcomes if hosts switch to non-aggressive behavioral tactics. Specifically, parasites themselves tended to terminate interactions in which venom was used. All data for the behavioral aspect of our investigation is contained in the \"Behavior\" folder. The data is completely contained in a single, large data set, and the analyses are presented in a single code document. Analyses within the behavior code-base are presented in sections dedicated to each major finding/figure presented in the PLOS One publication. Behavioral observations in this dataset are encoded, with associated interaction metadata, line by line. Each row represents a unique record based on a single observed behavior, and includes interaction metadata associated with that observation. Behaviors were categorized for the host as aggressive or submissive. Parasite venom behaviors included venom behavior. Irrelevant behaviors were classified in the \"other\" category. Cells without recorded data in a given line are not applicable to the current line record, and thus, do not represent missing data. Further information is provided in a dedicated Behavior README document inside the \"Behavior\" folder. ### Venom economy: To characterize the costs of parasite venom use within the context of interactions, we also assessed limitations to parasite venom usage in terms of 1) parasite anatomical venom storage, and 2) host neutralization capacity. We then used these concepts to define a \"venom economy\" between hosts and parasites in conflict. To establish venom storage capacity (in “number of stings”), we calculated the relative ratio in size between the individual droplets discharged by cold-anesthetized parasites directly from their sting, and the total venom liquid stored in the venom sac. To assess the number of hosts a single parasite can neutralize using venom directly (stinging), we performed a combination of mortality assays using both whole venom sting and synthetic venom proxy (stereoisomeric mix of 3-butyl-5-hexylpyrrolizidine). We used the synthetic venom proxy to calibrate a lethal dose-response curve for hosts, which we used as a baseline model against which to compare mortality from whole-venom single stings. Through our analyses, we showed the costs of parasite venom-use (via stinging) in terms of overall parasite venom reserves (per parasite). Note, however, that we assume that parasite venom supply is not immediately replenishable, that the rate of production is slow relative to the timescale over which host-parasite interactions occur, and that host mortality is not immediate. Additionally note that ‘sting’, in this analysis, refers to the droplet size that is produced by parasites under anesthetized conditions; however, parasite individuals may have finer control over their venom delivery when using venom under normal conditions. Analyses in the venom codebase are provided in a single code document, but the venom measurement data are split between several datasheets. Data structures within each datasheet vary, but are explained fully in a separate, dedicated venom README document, located inside the \"Venom\" folder. ## Data Navigation Within the main folder, Boot_et_al_2026_Dryad.zip, the data and code are split between 2 secondary folders (Behavior, Venom), representing the two primary analysis types presented in the associated publication. Cells without recorded data in a given line are not applicable to the current line record, and thus, do not represent missing data. ### FOLDER STRUCTURE: 1\\) Behavior 2\\) Venom *Each folder separately contains both the data and code necessary to reproduce all analyses used in the paper. #### The Behavior folder contains: 1 data file (MRB2024_BehaviorData_Submission2.xlsx), 1 R-markdown code document (MRB_SignalingMS_Analysis_Subm2_Behavior.Rmd), and 1 README file (1_BEHAVIOR_README.ods) explaining the datafile. #### The Venom folder contains: 4 data files (KSS_Msymmetochus_venomsize.csv, Venom_sting_summary.xlsx, KSS+MRB_Msymmetochus_Samabilis_StingMort.xlsx, MRB_LD50data.xlsx) 1 R-markdown code document (MRB_SignalingMS_Analysis_Subm2_Venom.Rmd), and 1 README file (1_VENOM_README.ods) explaining its datafiles. ## README DOCS \u0026amp; DATA ENUMERATION: The README docs within the Behavior and Venom folders are spreadsheets that explain what data is included in each data file, enumerate datatypes (columns), and define the values used in each column. Please refer to the respective READMEs for specific details regarding the datasets in each folder. ## CODE: The codebase uses R-markdown notation to split the code into multiple sections. The code is intended to be viewed in outline form using an IDE such as RStudio (or other similar software) that can parse R-markdown type script since each code document has multiple sections, and the Behavioral analysis in particular is quite long. In general, each code document has several main sections (in caps), and multiple subsections. For example: SECTIONS: LOAD PACKAGES BEHAVIORAL ANALYSIS PRIMARY ANALYSES - SET-UP \u0026amp; DESCRIPTIVE ANALYSIS 1\\) Behavioral summary 2\\) Behavioral overview 3\\) ... Each numbered section contains a set of analyses (incl. sub-analyses) which pertains to a named method/results passage in Boot et. al, 2026. Manuscript figure outputs are demarcated in the detailed outline (R-markdown) above the code chunk which will produce plot outputs used in those figures. Each code document contains further instructions that pertain to the specifics of that document. In the case that further instructions are provided in the document, line numbers are provided in the document notes header."}],"geoLocations":[],"fundingReferences":[{"funderIdentifierType":"ROR","funderName":"Division of Graduate Education","funderIdentifier":"https://ror.org/00whkrf32","awardTitle":"Graduate Research Fellowship Program (GRFP)","awardNumber":"2240614"},{"funderIdentifierType":"ROR","funderName":"Division of Environmental Biology","funderIdentifier":"https://ror.org/03g87he71","awardTitle":"\n        CAREER: Integrative Systematics: Taxonomy and Evolution of Megalomyrmex\n        Ants and Their Venom\n      ","awardNumber":"2146104"},{"funderIdentifierType":"ROR","funderName":"Division of Integrative Organismal Systems","funderIdentifier":"https://ror.org/01rvays47","awardTitle":"\n        Illumination of behavior leading to host exploitation by a\n        context-dependent mutualist\n      ","awardNumber":"2127521"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.zs7h44jpg","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-07-02T17:40:49Z","registered":"2026-07-02T17:40:50Z","published":null,"updated":"2026-08-21T14:35:45Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.bnzs7h4sz","type":"dois","attributes":{"doi":"10.5061/dryad.bnzs7h4sz","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["Institute of Chronoecology","Technical University of Munich"],"name":"Monecke, Stefanie","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0003-1167-752X"}]},{"nameType":"Personal","affiliation":[","],"name":"Gerrits, Theo","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Igneous software, Netherlands"],"name":"Brand, Rene","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Wageningen Environmental Research"],"name":"van Kats, Ruud","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Gaiazoo"],"name":"ter Meulen, Tjerk","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Institut des Neurosciences Cellulaires et Intégratives"],"name":"Pévet, Paul","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Technical University of Munich"],"name":"Vadkertikova, Monika","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Technical University of Munich"],"name":"Altunkaya, Alp","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Wageningen Environmental Research"],"name":"Müskens, Gerard","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Zoogdiervereniging","Radboud"],"name":"La Haye, Maurice","nameIdentifiers":[]}],"titles":[{"title":"Data from: Monitoring rodents' life history by body temperature and activity"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Environmental biotechnology","subjectScheme":"fos"},{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Natural sciences","subjectScheme":"fos"},{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"subject":"activity"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Body temperature","subjectScheme":"PLOS Subject Area Thesaurus"},{"subject":"Cricetus cricetus"},{"subject":"litter-size"},{"subject":"reproduction"}],"contributors":[],"dates":[{"date":"2026-07-30T09:00:12Z","dateType":"Created"},{"date":"2026-07-30T09:00:14Z","dateType":"Submitted"},{"date":"2026-08-21T00:00:00Z","dateType":"Issued"},{"date":"2026-08-21T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[{"relationType":"IsCitedBy","relatedIdentifier":"10.1002/wlb3.01707","relatedIdentifierType":"DOI"}],"relatedItems":[],"sizes":["17130095 bytes"],"formats":[],"version":"4","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"European hamster populations are rapidly declining all over their\n distribution area. One of the suggested reasons is a strong and ongoing\n decrease in reproductive output of about 77% since the nineteen twenties.\n The cause of this decline is, however, unknown. To find out why the\n reproduction rate has decreased, precise information on the phenology of\n reproduction and litter sizes is necessary. In this study, we show that\n recordings of body temperature (Tb) and activity of European hamster\n females provide accurate information on the reproductive phenology as well\n as individual health and death. In contrast to non-gravid females,\n thermograms and actograms of gestating and lactating females clearly show\n arrhythmic patterns. Parturition is indicated by a profound hypothermia.\n The duration of the decline part correlates with the litter size. Thus,\n the timing and number of litters as well as the litter size can be\n precisely determined in Tb recordings. Moreover, after parturition,\n females show rhythmic patterns of activity and Tb only after separation\n from the pups, no matter at which postnatal day that happens, allowing\n conclusions whether at least one pup could be raised to independence. In a\n pilot study, we recorded Tb in wild European hamster females in a Dutch\n population which is regularly restocked by captive-bred animals. These\n data support the hypotheses that European hamsters currently reproduce too\n less and too late in order to keep the population stable. A helpful\n side-effect of activity and temperature recordings is that they allow to\n determine the exact date and time of death of transmitter-tagged\n individuals and thus conclusions about the reason. With the exception of\n the sensor implantation, monitoring Tb is a non-invasive method to study\n the life history of small mammals. Monitoring reproduction of small\n mammals might be a helpful tool to understand the massive global decline\n of species."},{"descriptionType":"TechnicalInfo","description":"# Data from: Monitoring rodents' life history by body temperature and\n activity Dataset DOI:\n [10.5061/dryad.bnzs7h4sz](https://doi.org/10.5061/dryad.bnzs7h4sz) ##\n Description of the data and file structure These are the raw data on body\n temperature and activity (the latter only for Dutch animals) for the\n article Monitoring rodents’ life history by body temperature and activity.\n The ID of the animal is given in the file names as well as the recording\n period (YYYYMMDD). The IDs of the animals correspond to the IDs given in\n the article. Except for the Strasbourg animals, whose IDs were abbreviated\n in the article. The data ID 09.151.5 corresponds to 151 in the text and\n 09.152.5 to 152 The BioWise files of the Dutch animals were recorded by\n self-made dataloggers (@madebyTheo) and include the following columns:\n date and time (6min intervals), body temperature, activity, sound, light,\n and battery on or off. Please note the following: * Devices were\n programmed and started but not immediately implanted so that the relevant\n recording start only when body temperature reaches physiological values *\n Devices were not stopped immediately at death of the animal or\n explantation of the loggers. Thus, the relevant recording ends when body\n temperature shows unphysiological values * These were self-made sensors.\n The original idea was to track heart rate by sound, which was of course\n not possible at this recording interval. The data turned out to be\n useless. The next generation of sensors (v0.4) recorded heart rate\n instead, but this did not work out. * The idea of the light sensors was to\n help retrieving the sensors in free-ranging animals after the death of the\n animals. The data could be transmitted and received, and devices be\n located roughly by telemetry. If the sensor for light recorded light the\n remains of the animal, the logger needed to be searched above ground;\n otherwise, below ground. * the battery could be turned on or off to save\n battery. You see in the first lines of each file that the device was\n programmed (battery on), then the battery was switched off for a while\n until close to the intended implantation. Thus, there is a delay in time.\n Please read the method section of the article for details. Data set 1a\n comprises data of captive animals in a Dutch breeding colony: It includes\n animals 206,289,292,293,294,295,298,299. Data set 1b comprises data of\n captive animals in a French breeding colony: It includes animals 09.151.5\n and 09.152.5 (corresponding to animal 151 and 152 in the article). The\n data for the animals from Strasbourg 09.151.5 and 09.152.5 contain only a\n column for date/time and body temperature. They were recorded with\n iButtons in 30min intervals. Data set 2 comprises data of free-ranging\n animals from a hamster-friendly managed field in the province Limburg in\n the Netherlands. It includes animals 259,260,261,274,276,278,286,316,324\n Please note that the transmitter of animal 278 had a failure since the\n recorded body temperature dropped continuously over time, which\n doesn't allow to include it in each type of analysis. However, the\n changes in the Tb pattern are clearly visible. All other information is\n given in the associated article ## Code/software Given in the associated\n article"}],"geoLocations":[],"fundingReferences":[],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.bnzs7h4sz","contentUrl":null,"metadataVersion":0,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-21T11:58:05Z","registered":"2026-08-21T11:58:06Z","published":null,"updated":"2026-08-21T11:58:06Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.tx95x6bc2","type":"dois","attributes":{"doi":"10.5061/dryad.tx95x6bc2","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["Goethe University Frankfurt"],"name":"Rongstock, Lydia","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0002-9395-537X"}]},{"nameType":"Personal","affiliation":["Goethe University Frankfurt"],"name":"Scheepens, J.F.","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0003-1650-2008"}]},{"nameType":"Personal","affiliation":["University of Bremen"],"name":"Classen, Alice","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Goethe University Frankfurt"],"name":"Grünewald, Bernd","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Goethe University Frankfurt"],"name":"Krämer, Jan","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Goethe University Frankfurt"],"name":"Schuler, Aileen","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Stanford University"],"name":"Pate, Braden","nameIdentifiers":[]}],"titles":[{"title":"Summer resource deficits strongly constrain mid-season bumblebee colony development despite high early-season performance"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Bumblebees","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Conservation biology","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Urban environments","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Animal behavior","subjectScheme":"PLOS Subject Area Thesaurus"}],"contributors":[],"dates":[{"date":"2026-02-18T15:33:36Z","dateType":"Created"},{"date":"2026-08-15T11:08:52Z","dateType":"Submitted"},{"date":"2026-08-20T00:00:00Z","dateType":"Issued"},{"date":"2026-08-20T00:00:00Z","dateType":"Available"},{"date":"2026-08-21T00:00:00Z","dateType":"Updated"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[],"relatedItems":[],"sizes":["1026958 bytes"],"formats":[],"version":"5","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Land-use and climate change pose severe challenges for insect pollinators.\n Increasing urbanisation, drought, and high temperatures directly affect\n pollinator physiology; indirect effects occur when floral resource\n availability and quality decline. While absolute resource losses across\n environments have received considerable scientific attention, their\n seasonal availability and the direct consequences for social pollinators\n remain understudied. Here, we examined how seasonal resource availability\n translated into bumblebee colony growth and individual forager behaviour\n of Bombus terrestris. During the early, mid-, and late phases of a single\n vegetation period, we monitored fresh sets of colonies (n = 60 per season)\n in three urban and three semi-natural habitats and related fitness\n parameters to local resource availability and weather conditions. We\n additionally tested whether forager sucrose responsiveness reflected\n seasonal resource availability. Early-season colonies developed well in\n both habitat types, while mid- and late-season colony growth rates\n stagnated, mirroring the strong drop in resources after the early season.\n Habitat differences were small compared to seasonal effects. Elevated mid-\n and late-season temperatures further reduced reproductive success. Local\n late-season resource peaks benefited nearby colonies. Sucrose response\n thresholds reflected temporal and spatial nectar resource availability:\n they were high in early-season hives, low in mid-season, and partially\n increased again in late-season hives at sites with higher nectar\n availability. These results demonstrate how seasonal resource gaps,\n compounded by elevated temperatures, can severely limit bumblebee colony\n growth and reproductive success, yet local resource peaks can mitigate\n these impacts and should therefore be prioritised in conservation efforts."},{"descriptionType":"TechnicalInfo","description":"# Summer resource deficits strongly constrain mid-season bumblebee colony\n development despite high early-season performance Dataset DOI:\n [10.5061/dryad.tx95x6bc2](https://doi.org/10.5061/dryad.tx95x6bc2) ##\n Description of the data and file structure Data files (txt) and R Codes of\n the experimental field study 2022, including bumblebee colony performance,\n bumblebee sucrose responsiveness, and environmental conditions. ### Files\n and variables #### **File: Colony_Env.txt** **Description**: Bumblebee\n colony performance data and mean floral resources. Resources are either\n given per square metre (density) or as the covered proportion (coverage).\n ##### **Variables**: Grouping variables * Hive_ID * Site_Iteration *\n Iteration * Site * Category * Hive_type Colony performance * Larvae_queen\n * Larvae_midi (drones and workers) * Eggs_Pollen * Pupae_midi (drones and\n workers) * Pupae_queen * Brood_total (sum of all larvae, pupae, and eggs)\n * Nectar_pot, Pollen_pot * Brood_parasite (sum of parasitic specimens) *\n Bumblebee_spec (sum of other species found inside nests) * Food_thieves\n (sum of food harvesting individuals including wasps, honeybees) *\n Larvae_Pollen * Brood_Pollen * Pupae_total * WML (wax moth larvae) *\n Count_Hive_drone * Count_Hive_queen * Count_Hive_undefined (sex unclear,\n missing abdomen) * Count_Hive_worker * Mean_Thorax_Hive_drone *\n Mean_Thorax_Hive_queen * Mean_Thorax_Hive_undefined (sex unclear) *\n Mean_Thorax_Hive_worker Biomass metrics * Week_mB (week of maximum\n biomass) * maxBiomass (g, maximum biomass) * Week_End (week of last\n biomass value or End_Mass) * End_Mass (g) Flight Activity * mean_Entry *\n mean_Exit * mean_Entry_Pollen * mean_Activity_total * sum_Activity_total\n Resource Availability * Food_total * FU_per_sqm * pFU_area * Z_FU_per_sqm\n * FU_per_sqm_Nektar * pFU_area_Nektar * FU_per_sqm_Pollen *\n pFU_area_Pollen * log_FU_per_sqm * log_pFU_area, * log_FU_per_sqm_Nektar *\n log_FU_per_sqm_Pollen * log_pFU_area_Nektar * log_pFU_area_Pollen *\n FU_per_sqm_centered * FU_per_sqm_Nektar_centered *\n FU_per_sqm_Pollen_centered * pTree_area * pTree_area_Nektar *\n pTree_area_Pollen * log_pTree_area * log_pTree_area_Nektar *\n log_pTree_area_Pollen * pTree_area_centered * pTree_area_Nektar_centered *\n pTree_area_Pollen_centered * log_Sum_Flower_cover_Nek #### **File:\n NetBiomass.txt** **Description**: Weekly measurments of biomass (g) and\n temperature (°C) ##### **Variables**: Grouping variables * Hive_ID,\n Iteration * Site, Category * Hive_type * Week (from 0-6 across a single\n season) * Week_Number (from 0-19 across entire study) Biomass metrics *\n Biomasse (net biomass change) * Week_mB (week of maximum biomass) *\n maxBiomass (maximum biomass) * Week_End (week of last biomass value or\n End_Mass) * End_Mass * Vitality Temperature * Temp_week (temperature mean\n per week) * Temp_week_centered (centered variable) #### **File:\n SR_Specimens.txt** **Description**: Specimen Summary of Sucrose\n Responsiveness Tests (SR) ##### **Variables**: Grouping variables *\n Iteration * Category * Site * Hive_ID * Hive_Week (combined variable of\n Hive_ID and Week_SR) * Week_SR (SR test-week) * Week_Number (from 0-19\n across entire study) SR-Test * Specimen_count * Deceased_na (deceased\n without further information) * Deceased_bfeed (deceased before feeding) *\n Deceased_btest (deceased before testing) * Deceased_dtest (deceased during\n testing) * Deceased_count (deceased sum) * Mortality_rate * Escaped *\n Tested_count (sum of all tested specimens, excluding deceased and escaped)\n #### **File: Temp_daily.txt** **Description**: Daily mean temperatures\n (°C) ##### **Variables**: Grouping variables * Exp_Day (experiment day\n across a single season) * Field_Day (continuous day count across entire\n study) * Iteration * Site * Category Temperature * Temp_new_day\n (temperature daily mean) * Temp_new_day_centered (centered variable) ####\n **File: SR_Test.txt** **Description**: Sucrose Responsiveness Tests (SR)\n ##### **Variables**: Grouping variables * Bee_ID * Hive_ID *\n Site_Iteration_ID * Site * Category * Iteration (combined variable of\n Hive_ID and Week_SR) * Week_SR (SR test-week) * Week_Number (from 0-19\n across entire study) Test * Fed_btesting (μL, fed volume after harnessing)\n * W1 (first water stimulus) * 0.1 (first sucrose concentration) * W2 * 0.3\n * W3 * 1.0 * W4 * 3.0 * W5 * 10 * W6 * 30 * Water_Score * Sucrose_Score *\n Vitality #### **File: Vegetation.txt** **Description**: Floral resource\n data recorded per plot and season (includes only records with all\n measurments). Either given per square metre (density) or as the covered\n proportion (coverage). ##### **Variables**: Grouping variables * Aufnahme\n (Record ID) * Plot * Iteration * Site_ID * Site * Date * Category *\n BIO_TYP_XX (biotope type ID) * BIOTYPE (biotope type) * Field_Day *\n Colour_Cat * Exp_Day * Week_Number Resources * FU_per_sqm (floral unit\n density) * pFU_area (floral unit coverage) * pTree_area (tree area\n coverage) * pTree_area_Nektar (tree area coverage nectar-weighted) *\n pTree_area_Pollen(tree area coverage pollen-weighted) * pVeg_area (ratio\n of vegetated area per site) * Z_FU_per_sqm (z-transformed) *\n FU_per_sqm_Nektar * pFU_area_Nektar * FU_per_sqm_Pollen * pFU_area_Pollen\n * log_FU_per_sqm (log-transformed) * log_pFU_area * log_pTree_area *\n log_FU_per_sqm_Nektar * log_FU_per_sqm_Pollen * log_pFU_area_Nektar *\n log_pFU_area_Pollen * log_pTree_area_Nektar * log_pTree_area_Pollen *\n log_Sum_Flower_cover_Nektar (sum of floral unit and tree area coverage) *\n log_Sum_Flower_cover_Pollen * Sum_Flower_cover_Nektar *\n Sum_Flower_cover_Pollen * FU_per_sqm_centered * FU_per_sqm_Nektar_centered\n * FU_per_sqm_Pollen_centered * pTree_area_centered *\n pTree_area_Nektar_centered * pTree_area_Pollen_centered *\n Sum_Flower_cover_Nektar_centered * Sum_Flower_cover_Pollen_centered *\n Temperature (mean per site and season) #### **File: Vegetation_area.txt**\n **Description:** Floral resource data recorded per plot and season\n (includes all records, e.g., \u0026lt; 3 measurments), size and ratio of\n biotope types ##### **Variables**: Grouping variables * Iteration * Site *\n BIOTYPE (biotope type) * Aufnahme (Record ID) * Plot * Date * BIO_TYP_XX\n (biotope type ID) * Category * Field_Day * Colour_Cat * Exp_Day *\n Temperature * Week_Number Resources * FU_per_sqm (floral unit density) *\n pFU_area (floral unit coverage) * pTree_area (tree area coverage) *\n pTree_area_Nektar (nectar-weighted) * pTree_area_Pollen (pollen-weighted)\n * pVeg_area (ratio of vegetated area per site) * Z_FU_per_sqm\n (z-transformed) * FU_per_sqm_Nektar * pFU_area_Nektar * FU_per_sqm_Pollen\n * pFU_area_Pollen * log_FU_per_sqm * log_pFU_area * log_pTree_area *\n log_FU_per_sqm_Nektar * log_FU_per_sqm_Pollen * log_pFU_area_Nektar *\n log_pFU_area_Pollen * log_pTree_area_Nektar * log_pTree_area_Pollen *\n log_Sum_Flower_cover_Nektar * log_Sum_Flower_cover_Pollen *\n Sum_Flower_cover_Nektar * Sum_Flower_cover_Pollen * FU_per_sqm_centered *\n FU_per_sqm_Nektar_centered * FU_per_sqm_Pollen_centered *\n pTree_area_centered * pTree_area_Nektar_centered *\n pTree_area_Pollen_centered * Sum_Flower_cover_Nektar_centered *\n Sum_Flower_cover_Pollen_centered * area_ha (biotope type area size per\n site) * area_p (ratio of area covered of a given biotope type per site)\n #### **File: Rongstock_et_al_2026_R_Code.Rmd** **Description:** R Codes\n for statistical analysis ## Code/software R (version 4.4.1, R Core Team\n 2024), R-Studio (2024.4.2.764) packages: tidyverse (1.3.1), boot (1.3-30),\n lmerTest (3.1-3), glmmTMB (1.1.9), randomForest (4.7-1.2)"}],"geoLocations":[],"fundingReferences":[],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.tx95x6bc2","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-20T09:49:46Z","registered":"2026-08-20T09:49:47Z","published":null,"updated":"2026-08-21T11:10:42Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.x69p8czzx","type":"dois","attributes":{"doi":"10.5061/dryad.x69p8czzx","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["University of Kansas Medical Center"],"name":"Sadeghi, Maryam","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0003-0057-2562"}]},{"nameType":"Personal","affiliation":["University of Tehran"],"name":"Kordi, Mohammadreza","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Royan Institute"],"name":"Daemi, Mehdi","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Tehran"],"name":"Tabasi, Seyed Maziyar","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Tehran"],"name":"Ebrahimnezhad Bashiri, Mohammad Sina","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Royan Institute"],"name":"Nabavi, Seyed Massood","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Royan Institute"],"name":"Khaligh-Razavi, Seyed-Mahdi","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Kansas Medical Center"],"name":"Thompson, Jeffrey","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Kansas Medical Center"],"name":"Sosnoff, Jacob","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Kansas Medical Center"],"name":"Devos, Hannes","nameIdentifiers":[]}],"titles":[{"title":"Data from: Comparing exercise with virtual reality gaming on gait and cognition in relapsing-remitting multiple sclerosis: a randomized controlled trial"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Medical and health sciences","subjectScheme":"fos"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Multiple sclerosis","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Virtual reality","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Exercise","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Gait rehabilitation","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Cognitive impairment","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Machine learning","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Biomarkers","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Precision medicine","subjectScheme":"PLOS Subject Area Thesaurus"}],"contributors":[{"name":"         University-affiliated Neurology and Rehabilitation Research Center,         clinical trial and digital health research facilities       ","contributorType":"Sponsor","affiliation":[],"nameIdentifiers":[]}],"dates":[{"date":"2026-02-10T23:43:28Z","dateType":"Created"},{"date":"2026-02-10T23:43:34Z","dateType":"Submitted"},{"date":"2026-03-16T00:00:00Z","dateType":"Issued"},{"date":"2026-03-16T00:00:00Z","dateType":"Available"},{"date":"2026-08-21T00:00:00Z","dateType":"Updated"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[],"relatedItems":[],"sizes":["37518 bytes"],"formats":[],"version":"11","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Background: Exercise and virtual reality gaming may mitigate gait and\n cognitive deficits in relapsing-remitting multiple sclerosis (RRMS). The\n main aim was to compare the efficacy of both interventions on gait and\n cognition and gait in RRMS. Secondary aims were to explore the efficacy of\n both interventions on serum biomarkers and to explore the predictors of\n treatment response. Methods: Forty-eight participants with RRMS were\n randomized to exercise (n=19), VR (n=19), or wait-list control (n=10) for\n eight weeks. Primary outcomes were the 10-meter walk test (10MWT) and the\n Symbol Digit Modalities Test (SDMT). Secondary outcomes included serum\n levels of neurofilament light chain (NfL), brain-derived neurotrophic\n factor (BDNF), and insulin-like growth factor-1 (IGF-1). Extreme Gradient\n Boosting (XGBoost), Random Forest, and logistic regression models were\n trained to predict treatment response. Results: The exercise group\n improved 10MWT performance by 2.41 seconds and increased IGF-1 levels by\n 100.25 ng/ml, significantly more than the VR and control groups (both\n p\u0026lt;0.001). The VR group improved on the SDMT by 1.95 points (p=0.001\n vs. control; p=0.05 vs. exercise). Both interventions reduced NfL\n concentrations compared to control (exercise: –2.07 pg/ml; VR: –0.60\n pg/ml), with exercise showing a greater reduction than VR (p=0.02).\n XGBoost demonstrated highest predictive accuracy (10MWT: 87%; SDMT: 86%).\n SHapley Additive exPlanations (SHAP) analysis identified baseline IGF-1\n and BDNF as top predictors of 10MWT, and baseline CognICA, BDNF, and age\n as predictors of SDMT performance. Conclusion: Exercise preferentially\n improves gait and IGF-1, whereas VR gaming yields modest cognitive gains.\n Serum biomarkers enhance machine learning prediction of treatment\n response, supporting a precision rehabilitation approach in RRMS.\n Keywords: Multiple Sclerosis, Virtual Reality, Exercise, Gait, Cognition,\n Biomarkers, Machine Learning"},{"descriptionType":"TechnicalInfo","description":"# Data from: Comparing exercise with virtual reality gaming on gait and\n cognition in relapsing-remitting multiple sclerosis: a randomized\n controlled trial **Dataset DOI:**\n [https://doi.org/10.5061/dryad.x69p8czzx](https://doi.org/10.5061/dryad.x69p8czzx) ## Description of the Data and File Structure ### 1. Study Description This dataset was generated from an outcome assessor-blinded randomized controlled trial investigating the effects of exercise training and immersive virtual reality (VR) gaming, compared with a wait-list control condition, on gait performance, cognitive function, functional outcomes, and circulating biomarkers in individuals with relapsing-remitting multiple sclerosis (RRMS). Participants were randomly assigned to one of three groups: exercise intervention, VR gaming intervention, or wait-list control. Outcomes were assessed at baseline (pretest) and immediately following the 8-week intervention period (post-intervention). The data were analyzed using conventional statistical methods. Selected baseline variables were also used in exploratory supervised machine-learning analyses to investigate potential predictors of individual treatment response. ### 2. Study Design **Study type:** Randomized controlled trial\\ **Blinding:** Outcome assessor blinded\\ **Intervention duration:** 8 weeks\\ **Assessment points:** Baseline (pretest) and post-intervention **Groups:** * Exercise group: supervised aerobic and resistance training * Virtual reality group: immersive motor-cognitive VR gaming * Control group: wait-list control maintaining usual activity **Total participants:** 48 ### 3. Participants **Diagnosis:** Relapsing-remitting multiple sclerosis according to the 2010 McDonald criteria\\ **Age range:** 18–55 years\\ **Disability level:** Expanded Disability Status Scale (EDSS) ≤ 5.0\\ **Disease-modifying therapy:** Stable for ≥ 6 months The released data are de-identified. Participant identifiers in the dataset are anonymized and do not contain direct personal identifiers. ### 4. File Contents The dataset is provided in tabular CSV format and includes: * Participant demographic characteristics * Clinical disability measures * Cognitive test scores * Gait and functional mobility outcomes * Balance and flexibility measures * Serum biomarker concentrations * Baseline and post-intervention values * Calculated change scores Unless otherwise specified, change scores are calculated as: **Post-intervention value − baseline value** ### 5. Definition of Treatment Responders Responder definitions were specified for the exploratory supervised machine-learning analyses. **Gait responder (10MWT):** A participant was classified as a gait responder if 10-meter walk test time improved by at least 20%: **(Baseline − Post-intervention) / Baseline × 100 ≥ 20%** Because lower 10MWT times indicate faster gait, this formulation expresses improvement as a positive percentage. **Cognitive responder (SDMT):** A participant was classified as a cognitive responder if the Symbol Digit Modalities Test (SDMT) score increased by at least 4 points from baseline to post-intervention. These thresholds were selected based on clinically meaningful change criteria described in the associated manuscript. The machine-learning analyses should be interpreted as exploratory because of the modest sample size. ### 6. Data Processing Notes * Data are provided in non-imputed form. * No normalization or scaling has been applied to the deposited participant-level raw variables. * Change-score variables are calculated as post-intervention minus baseline unless otherwise noted. * The deposited data may be used for independent reanalysis, secondary analysis, and methodological replication subject to the CC0 license. ### 7. Ethical Approval and Informed Consent The study was conducted in accordance with the Declaration of Helsinki and approved by the Royan Institute Ethics Committee. **Ethics approval ID:** IR.ACECR.ROYAN.REC.1396.98 Written informed consent was obtained from all participants before study participation. ### 8. Data License This dataset is released under the **CC0 1.0 Universal Public Domain Dedication**. The data may therefore be copied, modified, distributed, and reused in accordance with the terms of the CC0 dedication. ## Files and Variables ### Files * `MSRCT_Data_Set__anonymized.csv` * `MSRCT_ALL_Mean_SD__anonymized.csv` * `MSRCT_ALL_Class__anonymized.csv` * `DisciplineSpecificMetadata.json` The CSV files contain de-identified data generated from the randomized controlled trial comparing exercise training, immersive VR gaming, and a wait-list control condition in individuals with RRMS. Data were collected at two principal assessment time points: * **Baseline (pretest)** * **Post-intervention (after 8 weeks)** The files include demographic characteristics, clinical disability measures, blood-based biomarkers, cognitive outcomes, and physical performance measures. ### Identifiers and Grouping * **ID** – Unique anonymized participant identifier * **group name** – Intervention assignment (`exercise`, `VR`, or `control`) * **gender** – Biological sex as coded in the deposited dataset (`m`, `f`) ### Demographic and Clinical Characteristics * **age** – Age in years * **EDSS (pretest)** – Expanded Disability Status Scale score at baseline * **EDSS (post)** – EDSS score after the intervention period * **EDSS diff** – Change in EDSS score (post − pre) ### Serum Biomarkers * **IGF-1 (ng/ml) pretest** – Baseline insulin-like growth factor-1 concentration * **IGF-1 (ng/ml) post** – Post-intervention IGF-1 concentration * **IGF-1 diff** – Change in IGF-1 concentration (post − pre) * **BDNF (pg/ml) pretest** – Baseline brain-derived neurotrophic factor concentration * **BDNF (pg/ml) post** – Post-intervention BDNF concentration * **BDNF diff** – Change in BDNF concentration (post − pre) * **NFL (pg/ml) pretest** – Baseline serum neurofilament light chain (NfL) concentration * **NFL (pg/ml) post** – Post-intervention serum NfL concentration * **NFL diff** – Change in serum NfL concentration (post − pre) **Note:** `NFL` is retained above where it reflects the deposited column name; the biomarker is referred to scientifically as neurofilament light chain (NfL). ### Cognitive Outcomes * **ICA score pretest** – Baseline Integrated Cognitive Assessment (CognICA) composite score * **ICA score post** – Post-intervention ICA score * **ICA diff** – Change in ICA score (post − pre) * **SDMT pretest** – Baseline Symbol Digit Modalities Test score (number correct) * **SDMT post** – Post-intervention SDMT score * **SDMT diff** – Change in SDMT score (post − pre) ### Functional and Mobility Outcomes * **Timed get up \u0026amp; go test (s) pretest** – Baseline Timed Up and Go (TUG) test time in seconds * **Timed get up \u0026amp; go test (s) post** – Post-intervention TUG time * **Timed get up \u0026amp; go test diff** – Change in TUG time (post − pre) * **Three minutes step test pretest** – Baseline number of steps completed * **Three minutes step test post** – Post-intervention number of steps completed * **Three minutes step test diff** – Change in step-test performance (post − pre) * **10-metre timed walk test (s) pretest** – Baseline 10-meter walk test (10MWT) time in seconds * **10-metre timed walk test (s) post** – Post-intervention 10MWT time * **10-metre timed walk test diff** – Change in 10MWT time (post − pre) * **Standing balance test (s) pretest** – Baseline standing-balance duration in seconds * **Standing balance test (s) post** – Post-intervention standing-balance duration * **Standing balance test diff** – Change in standing-balance duration (post − pre) * **The sit \u0026amp; reach-A test (cm) pretest** – Baseline sit-and-reach flexibility score in centimeters * **The sit \u0026amp; reach-A test (cm) post** – Post-intervention sit-and-reach score * **The sit \u0026amp; reach-A test diff** – Change in flexibility score (post − pre) ### `MSRCT_ALL_Class__anonymized.csv` This derived file was used in the exploratory machine-learning component of the study. It contains variables used in classification analyses of treatment response. Variables include: * **class** – Classification variable used in the derived machine-learning dataset. The exact coding should be interpreted according to the deposited data-generation workflow and associated manuscript. * **ICA score** – Integrated Cognitive Assessment (CognICA) composite score * **SDMT** – Symbol Digit Modalities Test score * **Timed get up \u0026amp; go test (s)** – Timed Up and Go result in seconds * **Three minutes step test** – Number of steps completed during the 3-minute step test * **10-metre timed walk test (s)** – Time required to walk 10 meters, in seconds * **Standing balance test (s)** – Static standing-balance duration, in seconds * **The sit \u0026amp; reach-A test (cm)** – Sit-and-reach flexibility score, in centimeters ### `DisciplineSpecificMetadata.json` This file contains discipline-specific metadata automatically generated during the Dryad submission process. It supports repository indexing and interoperability and is not required for analysis of the participant-level dataset. ### Missing Values No missing values are present in the final deposited dataset. ## Code and Software The deposited CSV files can be opened with standard tabular-data software, including Microsoft Excel, LibreOffice Calc, Google Sheets, R, or Python. No proprietary or custom software is required to access the raw data. ### Statistical Analysis Conventional statistical analyses reported in the associated manuscript were conducted using **SAS version 9.4**. ### Exploratory Machine-Learning Analysis Machine-learning analyses were conducted using **Python 3.11**, including: * **scikit-learn 1.3.0** – Logistic Regression, Random Forest, cross-validation, and model evaluation * **XGBoost 1.7.6** – Extreme Gradient Boosting classification * **SHAP** – Model interpretability and feature-attribution analyses * **NumPy** and **Pandas** – Numerical and tabular data processing The exploratory machine-learning workflow included: 1. Preparing de-identified baseline predictor variables 2. Defining responder and non-responder outcomes according to the prespecified thresholds described above 3. Training Random Forest, XGBoost, and logistic-regression classifiers 4. Performing stratified model validation and hyperparameter tuning 5. Evaluating classification performance 6. Exploring feature contributions using SHAP values Detailed machine-learning hyperparameters and methodological information are reported in the associated manuscript and Supplementary Methods. ## Access Information The dataset was generated by the study authors and was not derived from external or third-party datasets. The de-identified participant-level data and accompanying documentation are publicly available through the **Dryad Digital Repository**: [https://doi.org/10.5061/dryad.x69p8czzx](https://doi.org/10.5061/dryad.x69p8czzx) The dataset is provided to support transparency, independent reanalysis, and reproducibility of the associated research. All deposited data are released under the CC0 Public Domain Dedication. ## Human subjects data All data included in this dataset were collected from human participants following approval by the relevant institutional ethics committee. Written informed consent was obtained from all participants, including explicit consent for the use and publication of de-identified research data in the public domain. The dataset has been fully anonymized prior to deposition. All direct identifiers (such as names, contact information, national identification numbers, and exact dates of birth) were removed. Participants are represented only by randomly assigned study IDs. The dataset contains only non-identifiable demographic variables (e.g., age in years, sex), clinical measures, biochemical markers, and functional test outcomes. No personally identifiable information (PII) is included in this dataset, and the risk of re-identification is minimal."}],"geoLocations":[],"fundingReferences":[{"funderName":"University of Kansas Medical Center Endowment","awardTitle":"Robert and Aileen Button MS Research Scholarship Fund (2024)"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.x69p8czzx","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":44,"downloadCount":14,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-03-16T17:41:50Z","registered":"2026-03-16T17:41:51Z","published":null,"updated":"2026-08-21T10:32:32Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.h9w0vt4zb","type":"dois","attributes":{"doi":"10.5061/dryad.h9w0vt4zb","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["Simon Fraser University"],"name":"Johnson, Sarah","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0003-0506-5201"}]},{"nameType":"Personal","affiliation":["Simon Fraser University"],"name":"M'Gonigle, Leithen","nameIdentifiers":[]}],"titles":[{"title":"Data and code from: Bumble bees on edge: Species- and caste-specific variation after wildfire supports complementary resource use hypothesis"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Earth and related environmental sciences","subjectScheme":"fos"},{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Natural sciences","subjectScheme":"fos"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Bumblebees","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Ecology","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Wildfires","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Climate change","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Forest ecology","subjectScheme":"PLOS Subject Area Thesaurus"}],"contributors":[],"dates":[{"date":"2026-05-14T22:55:18Z","dateType":"Created"},{"date":"2026-05-22T20:25:13Z","dateType":"Submitted"},{"date":"2026-08-21T00:00:00Z","dateType":"Issued"},{"date":"2026-08-21T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[],"relatedItems":[],"sizes":["20231173 bytes"],"formats":[],"version":"5","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Fire impacts landscape structure and heterogeneity in many terrestrial\n ecosystems and climate change is accelerating and expanding its influence.\n One consequence of forest fire is the removal of canopy cover, which can\n lead to increases in abundance and diversity of flowering plants and\n changes to pollination dynamics. Bumble bees are an important group of\n temperate forest pollinators, where similar species vary in their\n conservation status and ecological preferences. The genus exhibits a\n complex life cycle, as different life stages require unique resources\n through a season that are likely to be differentially impacted by\n wildfire. When comparing burned and unburned forested habitat, positive\n effects of fire on bumble bees have been almost ubiquitous, but studies\n rarely address or control for distance to edge and often focus on a single\n species. Here we examine how the post-wildfire environment might influence\n species-specific bumble bee abundance over spatial distances that\n individual bees can feasibly move, and ask whether the potential\n complementarity of spatially distinct resources at a forest fire edge may\n lead to differential effects on queens, workers, and males. We identified\n an edge effect: bumble bee abundance was highest at the edge and declined\n with increasing distance into each habitat, potentially stabilizing at\n approximately 750-1000m from the edge. The strength of that decline varied\n by habitat, over the course of the season, and among species and castes,\n and this observed effect is likely not fully explained by fire-induced\n changes in floral abundance. Our findings indicate that distance to a\n forested edge may be an important factor for bumble bees benefiting from\n wildfire, especially for species with less generalized nesting\n preferences, as forest-specific nesting habitat is a likely explanation\n for our observed patterns. This has negative implications for the\n downstream regeneration potential of the distant interiors of larger-scale\n fires that we expect to see more frequently under climate change."},{"descriptionType":"Methods","description":"We conducted our study on the unceded territories of the Ulkatcho\n First Nation in Tweedsmuir Provincial Park in British Columbia, Canada\n from 25 June to 22 August 2021. We established sites in a grid pattern\n spanning a portion of the edge of a large (7,366.7 ha) wildfire that\n burned during the summer of 2017, situated within primarily lodgepole pine\n forest with a mix of closed and open canopy. We established a total of 90\n sites, each 10m in diameter, in unburned adjacent forest (42 sites), as\n close as possible to the discernible edge of the burn (6 sites), and\n within the burned forest (42 sites). To sample bumble bees, we set up one\n blue vane trap on the ground in the centre of each site to continuously\n collect from setup to takedown. Each site was sampled up to 3 consecutive\n times (sampling periods or visits); some site-collections were destroyed\n by bears, and any destroyed samples are excluded from all datasets. We\n sorted each bumble bee to caste (queen, worker, male) based on sex\n characteristics and size differences, and identified each bumble bee to\n species. During sampling visits we also collected\n quantitative and qualitative habitat information. We recorded floral\n community characteristics at the beginning and end of every sample\n collection period (four times)---censusing open flowers within a 10m\n diameter circular radius around the trap. We recorded several additional\n stable habitat characteristics at the site level (i.e., one measurement\n per site)---elevation using a handheld GPS device, canopy cover from four\n evenly spaced (equidistant from the trap in each cardinal direction)\n upward-facing photographs using a fish-eye lens, and descriptive estimates\n of three categories of ground cover---bare ground, exposed rock, and woody\n debris---as a visual percent of total site area, averaged between two\n observers."},{"descriptionType":"TechnicalInfo","description":"# Data and code from: Bumble bees on edge: Species- and caste-specific\n variation after wildfire supports complementary resource use hypothesis\n Dataset DOI:\n [10.5061/dryad.h9w0vt4zb](https://doi.org/10.5061/dryad.h9w0vt4zb) ##\n Description of the data and file structure ### Files and variables ####\n File: Bumble-bees-on-edge.zip Zip file contains all necessary scripts\n (code) and datasets to run models and/or produce figures, in the\n appropriate structure required. ### Folder: data **Description**: contains\n all datasets used in the paper as raw csv files \u0026amp; an RData file *\n **File: dd-bb.csv** - pooled bumble bee dataset * **File: dd-veg.csv** -\n vegetation dataset * **File: dd-species.csv** - bumble bee dataset, with\n abundances separated per species (3 common species) * **File:\n dd-caste.csv** - bumble bee dataset, with abundances separated per caste *\n **File: dd-species-rare.csv** - bumble bee dataset, with abundances\n separated per species (3 rare species) * **File: dd-datasets.RData** -\n RData file containing above datasets ready for import into R to be loaded\n in the scripts that require them **Variables shared among datasets (all\n variables ending in \"*.s*\" are the original variable scaled, by\n subtracting the column mean and dividing by column standard deviation)** *\n **site**: unique site index for which the sample was collected (one for\n each of the 90 sites) * **visit**: sample collection index (1-3 for bumble\n bees, 0-3 for vegetation), nested within site * **sample.quality**:\n indicator for level of bear disturbance of sample (destroyed samples are\n excluded) * yes = complete sample, no evidence of disturbance * partial =\n some insects sampled, but evidence of trap disturbance was observed *\n **habitat2; habitat.binary:** habitat type category of site (all sites,\n including \"edge\", were categorized by location of site centre\n relative to the fire perimeter) * Burn; 0 = burned forest * Forest; 1 =\n unburned forest * **edge.distance:** minimum straight line distance of\n centre of site from fire perimeter (m) * **julian.day:** continuous count\n of day of year for the respective sample/observation (median of start and\n end day of trapping for bumble bees) * **trap.hours**: duration of\n sampling (trapping) period in hours (start to end) ##### **Variables\n unique to dd-bb.csv** * **abundance**: total number of bumble bees\n collected in the sample * **elevation**: site elevation (m) *\n **canopy**.mean: site canopy openness (%), average from 4 photos\\ *\n **rock**: visual estimate of proportion of site ground cover consisting of\n exposed rock (%) * **ground**: visual estimate of proportion of site\n ground cover consisting of bare soil (%) * **wood**: visual estimate of\n proportion of site ground cover consisting of dead woody debris (%) #####\n Variables unique to dd-veg.csv * **veg.abundance:** total open flowers\n counted in the site ##### Variables unique to dd-species.csv \u0026amp;\n dd-species-rare.csv * **species:** species of bumble bee collected *\n flavifrons = *Bombus flavifrons* * bifarius = *Bombus bifarius*; * mixtus\n = *Bombus mixtus*; * terricola = *Bombus terricola*; * rufocinctus\n = *Bombus rufocinctus*; * melanopygus = *Bombus melanopygus* *\n **sp.abundance:** total number of bumble bees of the respective *species*\n collected in the sample ##### Variables unique to dd-caste.csv *\n **caste**: caste of bumble bee collected (queen, worker, male), all\n species poold * **caste.abundance**: total number of bumble bees of the\n respective caste collected in the sample ## Code/software Annotations are\n provided throughout each script to guide the user through important steps\n and clarify complex procedures. There is no mandatory order for running\n scripts, as all required file dependencies have been provided within the\n folder structure.* *To run, make sure to set the working\n drive for each script to the location of the downloaded main zip folder on\n your machine. All models were run in Stan using the R package rstan\n version 2.32.6 in R version 4.4.1. Figures are all constructed using base\n R and any required additional packages are loaded within the script for\n which they are used. ### Folder: models **Description**: contains all\n required model scripts and respective RData files containing full model\n runs from paper ready for import, loaded in the scripts that require them\n **File: b-model.R** **Description:** Stan model script for pooled bumble\n bee analysis, specifying the required data inputs and parameter structure\n for Bayesian analysis of the fixed effects of (scaled) *edge.distance*,\n *trap.hours*, *julian.day*, *edge.distance* x *habitat*, *edge.distance* x\n *julian.day*, and the random effect of *site* on log transformed bumble\n bee abundance, called in the *analyses.R* script **File:\n bb-model-run.RData** **Description:** the saved model run presented in our\n paper, which can be loaded directly in *figures.R* to reproduce all\n figures **File: veg.model.R** **Description:** the same model as\n *bb-model.R* but for floral analysis, excluding the fixed effect of\n *trap.hours (veg-model-run.RData* saved run) **File: species-model.R**\n **Description**: the same model as bb-model.R for species or caste\n analysis, depending on which case is selected in the analyses.R script,\n where each parameter is indexed by species or caste\n (species-model-run.RData, caste-model-run.RData saved runs) ### Folder:\n figs **File: plotting.R** **Description**: script with required plotting\n functions loaded in scripts that require them ### Other files: **File:\n analyses.R** **Description**: script containing all code for analyses (if\n entire script is run, it will overwrite the full model run(s) from the\n paper pre-saved in the models folder) **File: figures.R** **Description**:\n script containing all code to produce figures, and when run will save\n figures as PDF files into the figs folder"}],"geoLocations":[],"fundingReferences":[{"funderIdentifierType":"ROR","funderName":"Natural Sciences and Engineering Research Council of Canada","funderIdentifier":"https://ror.org/01h531d29"},{"funderIdentifierType":"ROR","funderName":"Simon Fraser University","funderIdentifier":"https://ror.org/0213rcc28"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.h9w0vt4zb","contentUrl":null,"metadataVersion":0,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-21T10:10:29Z","registered":"2026-08-21T10:10:30Z","published":null,"updated":"2026-08-21T10:10:30Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.dncjsxmf0","type":"dois","attributes":{"doi":"10.5061/dryad.dncjsxmf0","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["China Jiliang University"],"name":"Li, Hong-Liang","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0001-7094-2472"}]},{"nameType":"Personal","affiliation":["China Jiliang University"],"name":"Zhang, Hong-Qi","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Ningxia Academy of Agriculture and Forestry Sciences"],"name":"Zhang, Zhi-Ke","nameIdentifiers":[]}],"titles":[{"title":"Data and code from: Binding characteristics and physicochemical basis of C-minus odorant-binding protein 8 in western flower thrips, \u003cem\u003eFrankliniella occidentalis\u003c/em\u003e with various host vegetable volatiles"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Agricultural sciences","subjectScheme":"fos"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Odorant binding proteins","subjectScheme":"PLOS Subject Area Thesaurus"},{"subject":"Western flower thrips"},{"subject":"Host plant volatiles"}],"contributors":[],"dates":[{"date":"2026-04-16T05:34:51Z","dateType":"Created"},{"date":"2026-08-17T10:13:26Z","dateType":"Submitted"},{"date":"2026-08-21T00:00:00Z","dateType":"Issued"},{"date":"2026-08-21T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[],"relatedItems":[],"sizes":["58002 bytes"],"formats":[],"version":"5","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Western flower thrips (WFT), Frankliniella occidentalis (Pergande)\n (Thysanoptera: Thripidae), is a globally invasive polyphagous pest that\n relies on olfaction to sense host plants. Odorant-binding proteins (OBPs)\n in insect sensillar lymph mediate interactions with plant volatiles.\n However, the physicochemical mechanisms by which OBPs in F. occidentalis\n binding to host-plant volatiles, the key amino acid residues involved in\n these interactions, and the structural characteristics of the preferred\n ligands remain unclear. Here, we cloned a novel OBP, FoccOBP8, which\n belongs to the C-minus OBP subfamily based on sequence alignment and\n phylogenetic analyses. Competitive fluorescence binding assays using\n purified recombinant FoccOBP8 revealed broad binding interactions toward\n host-plant volatiles. Notably, 4-ethylacetophenone and ethyl isonicotinate\n showed the strongest binding affinity with dissociation constants (KD) of\n 2.00 and 2.46 µmol/L, respectively. Thermodynamic analyses indicated the\n binding processes were dynamic quenching processes and mainly driven by\n hydrophobic interactions, spontaneously. Molecular docking and the\n corresponding residue-interaction energy heatmap analyses predicted that\n the conserved Gln4 residue contributed substantially to the ligand-protein\n interactions. Molecule clustering and structural analyses predicted that\n the ligands containing aromatic moieties π-π conjugated with carbonyl\n group or carbon-carbon double bond preferential bound to FoccOBP8. This\n study elucidates the broad binding profiles of FoccOBP8 with ligands and\n their physicochemical basis, providing insights into the olfactory-driven\n polyphagia of WFT and theoretical foundations for developing odor-based\n attractants or repellents targeting this invasive pest. Keywords: Western\n flower thrips (WFT); Odorant-binding proteins (OBPs); Host-plant\n volatiles; Fluorescence spectroscopy; Thermodynamics; Molecular docking;\n Molecular Clustering."},{"descriptionType":"Methods","description":"The study characterized the C-minus odorant-binding protein\n FoccOBP8 from the western flower thrips, \u003cem\u003eFrankliniella\n occidentalis\u003c/em\u003e, using sequence analysis, recombinant-protein\n fluorescence binding assays, thermodynamic fluorescence analysis,\n molecular docking, residue-interaction energy analysis, and molecular\n clustering. The fluorescence competitive-binding experiment tested 17\n candidate host-plant volatiles. The two strongest-affinity ligands were\n 4-ethylacetophenone (4-EAP; KD = 2.00 μmol/L) and ethyl isonicotinate (EI;\n KD = 2.46 μmol/L). The FoccOBP8-1-NPN binding experiment used 1-NPN\n (N-phenyl-1-naphthylamine; CAS 90-30-2) as the fluorescent probe. Protein\n concentration was 1 μmol/L, the 1-NPN stock solution was 1 mmol/L,\n fluorescence spectra were collected from 290 to 500 nm, and the excitation\n wavelength used for FoccOBP8 detection was 282 nm.Thermodynamic and\n Stern-Volmer analyses for these two ligands were performed at 290 K, 300\n K, and 310 K. Molecular docking and residue-interaction energy analyses\n were performed for seven selected high-affinity ligands. ECFP4/Tanimoto\n molecular clustering and maximum common substructure (MCS) analyses were\n also performed for these seven ligands."},{"descriptionType":"TechnicalInfo","description":"# Data and code from: Binding characteristics and physicochemical basis of\n C-minus odorant-binding protein 8 in western flower thrips, *Frankliniella\n occidentalis* with various host vegetable volatiles Dataset DOI:\n 10.5061/dryad.dncjsxmf0 ## Description of the data and file structure ###\n Principal investigator contact information * **Name:** Hong-Liang Li *\n **Institution:** College of Life Sciences, China Jiliang University,\n Hangzhou 310018, China * **Email:**\n [hlli@cjlu.edu.cn](mailto:hlli@cjlu.edu.cn) ### Alternate contact\n information * **Name:** Hong-Qi Zhang * **Institution:** College of Life\n Sciences, China Jiliang University, Hangzhou 310018, China * **Email:**\n [2985425150@qq.com](mailto:2985425150@qq.com) ### Dataset overview This\n dataset contains the original measurements, derived values, sequence data,\n docking-energy data, and analysis code generated for the accepted\n manuscript named in the title of this README. ### File inventory and\n relationships All files are located in the dataset root archive\n IMB-R1-Original_data.zip\\ \\ The five CSV files are UTF-8 comma-separated\n tables with one header row and one observation or matrix-entry record per\n row. 1. `FoccOBP8_OBP_sequences_unaligned.fasta`: 48 unaligned amino-acid\n sequences. 2. `FoccOBP8_OBP_sequences_aligned_MEGA11.fas`: the same 48\n sequences after MEGA11 multiple-sequence alignment. 3.\n `FoccOBP8_binding_to_1-NPN.csv`: FoccOBP8 titration and Scatchard-analysis\n data. 4. `FoccOBP8-1-NPN_binding_to_17_ligands.csv`: the complete\n 17-ligand competitive-binding experiment in long format. 5.\n `Double-logarithm_equation_plot.csv`: double-logarithmic analysis data for\n two ligands at three temperatures. 6. `S-V_plot.csv`: Stern-Volmer\n analysis data for two ligands at three temperatures. 7.\n `Heatmap_of_key_residues.csv`: residue-ligand interaction-energy data for\n seven ligands. 8. `Molecular_analysis_and_MCS.py`: Python code for\n molecular clustering and MCS analysis. 9. `README.md`: this data\n description and reuse guide. The fluorescence files are related as\n follows: `FoccOBP8_binding_to_1-NPN.csv` contains the initial 1-NPN\n titration; `FoccOBP8-1-NPN_binding_to_17_ligands.csv` contains competitive\n displacement by 17 ligands; and `Double-logarithm_equation_plot.csv` and\n `S-V_plot.csv` contain temperature-dependent analyses for ethyl\n isonicotinate and 4-ethylacetophenone. `Heatmap_of_key_residues.csv`\n contains docking-derived interaction energies for seven selected ligands.\n The sequence files provide the input sequence collection and its MEGA11\n alignment. The Python script independently defines its seven ligand\n structures as SMILES strings. ## Files and variables ### 1.\n `FoccOBP8_OBP_sequences_unaligned.fasta` This FASTA file contains 48\n amino-acid sequence records used for sequence characterization and\n homology analysis. Each record begins with a FASTA header containing an\n accession identifier followed by an amino-acid sequence. ### 2.\n `FoccOBP8_OBP_sequences_aligned_MEGA11.fas` This FAS file contains the\n same 48 amino-acid sequence records after multiple-sequence alignment in\n MEGA11. All aligned records have the same length. Gap characters (`-`)\n indicate alignment positions and are not amino acids. ### 3.\n `FoccOBP8_binding_to_1-NPN.csv` This table contains 19 FoccOBP8 titration\n records used for Scatchard analysis of the FoccOBP8-1-NPN complex. |\n Variable | Description | Unit or interpretation | |\n -------------------------- |\n ------------------------------------------------------------- |\n ---------------------------------------------------------- | |\n `NPN_concentration_umol_L` | Concentration of 1-NPN added during titration\n | μmol/L | | `Fluorescence_intensity` | Measured fluorescence intensity\n after addition of 1-NPN | Instrument fluorescence-intensity units | |\n `Bound_NPN_umol_L` | Calculated concentration of protein-bound 1-NPN |\n μmol/L; `NA` means not reported in the source table | |\n `Bound_to_free_NPN_ratio` | Calculated bound/free 1-NPN ratio used for\n Scatchard analysis | Dimensionless; `NA` means not reported in the source\n table | ### 4. `FoccOBP8-1-NPN_binding_to_17_ligands.csv` This table\n contains the complete competitive-binding experiment for 17 candidate\n host-plant volatiles. It is intentionally one experiment-level file rather\n than 17 separate files. Each quantitative fluorescence observation\n occupies one row. The ligands are Benzaldehyde, β-ionone, p-cymene,\n o-cymene, (E)-cinnamaldehyde, p-anisaldehyde, linalool, 1,8-cineole,\n 1,2-diethylbenzene, ethyl isonicotinate, eugenol, 1-octen-3-ol, dodecane,\n salicylaldehyde, 4-ethylacetophenone, 1-vinylpyrene, and p-xylene. |\n Variable | Description | Unit or interpretation | |\n ----------------------------------------------------- |\n -------------------------------------------------- |\n -------------------------------------------------------------------------------------------- | | `Ligand` | Host-plant volatile represented by the record | Chemical name | | `Sample_ID` | Identifier for the FoccOBP8-1-NPN measurement | Categorical identifier; `NA` for the eugenol status-only record | | `Ligand_concentration_umol_L` | Concentration of the competing ligand | μmol/L; `NA` where no quantitative measurement was recorded | | `Fluorescence_intensity` | Measured fluorescence intensity | Instrument fluorescence-intensity units; `NA` where no quantitative measurement was recorded | | `Relative_fluorescence_fraction` | Fluorescence relative to the corresponding control | Fraction; `1` represents 100% of reference fluorescence | | The ligand stock solution concentration was 1 mmol/L. | \n \n | \n \n | ### 5. `Double-logarithm_equation_plot.csv` This table contains 108\n records used for double-logarithmic fitting of FoccOBP8 binding to ethyl\n isonicotinate and 4-ethylacetophenone at 290 K, 300 K, and 310 K. Each row\n represents one ligand, quencher concentration, and temperature\n combination. | Variable | Description | Unit or interpretation | |\n ------------------------------- |\n --------------------------------------------------------------------------------- | ----------------------------------------------------------------------- | | `Ligand` | Ligand analyzed | `ethyl isonicotinate` or `4-ethylacetophenone` | | `Quencher_concentration_umol_L` | Ligand/quencher concentration | μmol/L | | `log10_Quencher_concentration` | Base-10 logarithm of quencher concentration | Dimensionless; `NA` at zero concentration because log10(0) is undefined | | `Temperature_K` | Measurement temperature | K; 290, 300, or 310 | | `Fluorescence_intensity` | Measured fluorescence intensity | Instrument fluorescence-intensity units | | `F0_over_F` | Initial fluorescence without ligand divided by fluorescence after ligand addition | Dimensionless | | `log10_F0_over_F_minus_1` | Base-10 logarithm of F0/F - 1 | Dimensionless; `NA` where undefined or not reported | ### 6. `S-V_plot.csv` This table contains 108 records used to construct Stern-Volmer plots for ethyl isonicotinate and 4-ethylacetophenone at 290 K, 300 K, and 310 K. Each row represents one ligand, quencher concentration, and temperature combination. | Variable | Description | Unit or interpretation | | ------------------------------- | --------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------- | | `Ligand` | Ligand analyzed | `ethyl isonicotinate` or `4-ethylacetophenone` | | `Quencher_concentration_umol_L` | Ligand/quencher concentration | μmol/L | | `Temperature_K` | Measurement temperature | K; 290, 300, or 310 | | `Fluorescence_intensity` | Measured fluorescence intensity | Instrument fluorescence-intensity units | | `F0_over_F` | Initial fluorescence without ligand divided by fluorescence after ligand addition | Dimensionless; `NA` preserves the blank source cell at 310 K and zero quencher concentration | ### 7. `Heatmap_of_key_residues.csv` This table contains 84 long-format residue-ligand interaction-energy records used to generate the molecular-docking heatmap. | Variable | Description | Unit or interpretation | | -------------- | -------------------------------------- | ------------------------------------------------------------------- | | `Residue` | FoccOBP8 residue name and position | For example, `Gln 4` or `Glu50` | | `Ligand` | Docked ligand | Chemical name | | `EPair_kJ_mol` | Residue-ligand pair interaction energy | kJ/mol; numeric zero values are recorded values and are not missing | The seven ligands are 4-ethylacetophenone, ethyl isonicotinate, p-anisaldehyde, (E)-cinnamaldehyde, p-cymene, 1-octen-3-ol, and o-cymene. ### 8. `Molecular_analysis_and_MCS.py` This Python script contains the ligand definitions, ECFP4/Tanimoto-distance hierarchical clustering procedure, molecular-structure visualization, and MCS analysis for the seven selected high-affinity ligands. It does not require a separate input file because ligand structures are defined internally as SMILES strings. ## Code and software requirements ### Sequence analysis software * BLASTp was used to retrieve homologous protein sequences. * MEGA11 was used for amino-acid sequence alignment and phylogenetic analysis. * The neighbor-joining phylogenetic method used 1,000 bootstrap replicates. * iTOL v7 was used to visualize the phylogenetic tree. ### Fluorescence, thermodynamic, and docking analysis Fluorescence spectra were recorded with an RF-5301 PC spectrofluorometer (Shimadzu, Japan). Stern-Volmer plots were generated by linear regression of `F0/F` against `[Q]`. Double-logarithmic plots used `log10(F0/F - 1)` against `log10[Q]` to estimate the apparent binding constant `KA` and binding-site parameter `n`. `F0` is initial fluorescence without ligand and `F` is fluorescence after ligand addition. SWISS-MODEL was used to construct the FoccOBP8 homology model using the *F. occidentalis* OBP template with accession A0A6J1S430. Molegro Virtual Docker 4.2 was used for molecular docking, and PyMOL was used to visualize docking interactions. Residue-level pair-interaction energies (`EPair`) from docking were used for `Heatmap_of_key_residues.csv`. ### Molecular clustering and MCS script Recommended environment: ```bash conda create -n mol_cluster_env python=3.9 -y conda activate mol_cluster_env conda install -c conda-forge rdkit matplotlib scipy pandas numpy -y ``` Run from the dataset directory: ```bash python Molecular_analysis_and_MCS.py ``` The script requires RDKit, NumPy, pandas, SciPy, and matplotlib. Exact package versions used for the original analysis were not recorded. The clustering procedure uses ECFP4 fingerprints (`radius = 2`, `nBits = 1024`), Tanimoto similarity, a distance defined as `1 - Tanimoto similarity`, and hierarchical clustering. It generates: * `ligand_clustering_dendrogram.png`: clustering dendrogram. * `ligand_mcs_results.csv`: cluster identifier, ligand pair, MCS SMILES, and MCS image path. * `mcs_results/`: pairwise MCS visualization PNG files for ligand pairs within multi-member clusters. Cluster identifiers in the script are manually assigned after hierarchical clustering to reproduce the grouping used in the associated figure. These are reproducible outputs, not required input files. The script closes the plot after saving it so it can run in a non-interactive environment."}],"geoLocations":[],"fundingReferences":[{"funderIdentifierType":"ROR","funderName":"National Natural Science Foundation of China","funderIdentifier":"https://ror.org/01h0zpd94"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.dncjsxmf0","contentUrl":null,"metadataVersion":0,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-21T09:51:14Z","registered":"2026-08-21T09:51:15Z","published":null,"updated":"2026-08-21T09:51:15Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.0k6djhbg4","type":"dois","attributes":{"doi":"10.5061/dryad.0k6djhbg4","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["University Of Thessaly"],"name":"Bali, Eleftheria-Maria","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0009-0002-4475-6507"}]},{"nameType":"Personal","affiliation":["University Of Thessaly"],"name":"Verykouki, Eleni","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University Of Thessaly"],"name":"Papadopoulos, Nikos","nameIdentifiers":[]}],"titles":[{"title":"Data from: Thermal acclimation and environment drive field dispersal patterns of adult Mediterranean fruit flies"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"subject":"Ceratitis capitata"},{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Agricultural sciences","subjectScheme":"fos"},{"subject":"release-recapture"},{"subject":"environmental conditions"}],"contributors":[],"dates":[{"date":"2026-03-01T08:34:08Z","dateType":"Created"},{"date":"2026-08-18T08:07:27Z","dateType":"Submitted"},{"date":"2026-08-21T00:00:00Z","dateType":"Issued"},{"date":"2026-08-21T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[],"relatedItems":[],"sizes":["1052648 bytes"],"formats":[],"version":"3","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"The study examines how thermal acclimation and environmental conditions\n interact to influence dispersal of Mediterranean fruit fly adults in the\n field. Four release–recapture trials were conducted using wildish marked\n adults acclimated at 20, 25 or 30 °C. Trials were performed in two\n contrasting habitats: a mixed-fruit orchard containing host plants and an\n olive orchard lacking hosts. Trapping grids centered on release points\n were established using McPhail and Jackson traps positioned at increasing\n distances. At ten days of age, 100 individuals per sex and treatment were\n released at each site. The overall recapture rate was 14.8%, with males\n recaptured more frequently than females. Recapture rates increased with\n ambient temperature and declined rapidly over time, with most individuals\n captured close to release points and within ten days. Sites containing\n host plants yielded higher recaptures, whereas thermal acclimation had no\n significant effect."},{"descriptionType":"TechnicalInfo","description":"# Data from: Thermal acclimation and environment drive field dispersal\n patterns of adult Mediterranean fruit flies Dataset DOI:\n [10.5061/dryad.0k6djhbg4](https://doi.org/10.5061/dryad.0k6djhbg4) ##\n Description of the data and file structure This study examines how thermal\n acclimation and environmental conditions interact to influence dispersal\n of Mediterranean fruit fly adults in the field. Four release–recapture\n trials were conducted using wildish marked adults acclimated at 20, 25 or\n 30 °C. Trials were performed in two contrasting habitats: a mixed-fruit\n orchard containing host plants and an olive orchard lacking hosts.\n Trapping grids centered on release points were established using McPhail\n and Jackson traps positioned at increasing distances. At ten days of age,\n 100 individuals per sex and treatment were released at each site. ###\n Files and variables #### File:\n Data_Thermal_acclimation_and_environment_drive_field_dispersal_patterns_of_adult_Mediterranean_fruit_flies.xlsx ##### Variables | \n \n | \n \n | | :--------------------------------- |\n :------------------------------------------------------------------------------- | | No release | Number of release conducted | | Mean Temperature of release period | The mean Temperature (°C) recorded during the whole periode for each release | | Mean Temperature daily | The mean daily Temperature (°C) recorded in each day of the experiment | | Days after release | The days after release in which the traps were inspected | | Plot | Type of plot where the experiments were conducted \\[1=Non host, 2= Host] | | RP\\_Latitude | Latitude of the release point in each plot | | RP\\_Longitude | Longitude of the release point in each plot | | Trap type | Type of the traps which were used to recapture the flies \\[1=Jackson, 2=McPhail] | | Trap number | Number of each trap deployed in each plot | | T\\_Latitude | Latitude of each trap in each plot | | T\\_Longitude | Longitude of each trap in each plot | | Distance | Distance (m) between each trap and the release point of the plot | | Sex  | Sex of the flies recaptured \\[1=males, 2=females] | | Acclimation | Acclimation temperature of the recaptured flies \\[1=20°C, 2=25°C, 3=30°C] | | Captures | Number of recaptured individuals | ## Code/software To run the files we used R version 4.5.2 (R Core Team, Vienna, Austria). All models were implemented using the glmmTMB package, and estimated marginal means were calculated using the emmeans package."}],"geoLocations":[],"fundingReferences":[{"funderIdentifierType":"ROR","funderName":"European Union","funderIdentifier":"https://ror.org/019w4f821","awardNumber":"818184 (FF-IPM)"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.0k6djhbg4","contentUrl":null,"metadataVersion":0,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-21T09:24:30Z","registered":"2026-08-21T09:24:31Z","published":null,"updated":"2026-08-21T09:24:31Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.z34tmpgwg","type":"dois","attributes":{"doi":"10.5061/dryad.z34tmpgwg","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["University of Illinois Urbana-Champaign"],"name":"Zhou, Yi","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0002-4177-8818"}]},{"nameType":"Personal","affiliation":["University of Illinois Urbana-Champaign"],"name":"Ge, Yifei","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Illinois Urbana-Champaign"],"name":"Harrison, Wesley","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0001-8256-5926"}]},{"nameType":"Personal","affiliation":["University of Illinois Urbana-Champaign"],"name":"Zhao, Huimin","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0002-9069-6739"}]}],"titles":[{"title":"Data for: Asymmetric enzymatic hydrophosphorylation via O₂ activation"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Biocatalysis","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Computational biology","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Computational chemistry","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Chemical sciences","subjectScheme":"fos"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Photochemistry","subjectScheme":"PLOS Subject Area Thesaurus"}],"contributors":[],"dates":[{"date":"2026-07-27T16:55:23Z","dateType":"Created"},{"date":"2026-08-14T19:29:21Z","dateType":"Submitted"},{"date":"2026-08-21T00:00:00Z","dateType":"Issued"},{"date":"2026-08-21T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[],"relatedItems":[],"sizes":["8522567 bytes"],"formats":[],"version":"5","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Enzymatic carbon–phosphorus (C–P) bond formation is extremely rare in\n nature, limiting biocatalytic access to phosphorus-containing compounds\n that are widely used in pharmaceuticals and agrochemicals. Here, we report\n an asymmetric enzymatic hydrophosphorylation via O₂ activation using a\n repurposed flavin-dependent enzyme. Mechanistic studies reveal that\n reactive oxygen species (ROS) are converted into productive P-centered\n radicals, followed by radical addition and enzymatic hydrogen atom\n transfer, achieving high enantioselectivity. The enzyme accommodates\n diverse P–H donors that pose challenges to chemical catalysis, enabling\n the biosynthesis of valuable phosphorus-containing scaffolds. This work\n expands the scope of biocatalysis to programmable C–P bond formation and\n establishes a new paradigm for channeling oxygen reactivity in enzymes."},{"descriptionType":"TechnicalInfo","description":"# Data for: Asymmetric enzymatic hydrophosphorylation via O₂ activation\n Dataset DOI:\n [10.5061/dryad.z34tmpgwg](https://doi.org/10.5061/dryad.z34tmpgwg) ##\n Description of the data and file structure This dataset supports the\n research published in \"Asymmetric Enzymatic Hydrophosphorylation via\n O₂ Activation\" and contains all the computational results as\n described in the article and Supplementary Information. Molecular docking\n was performed using AutoDock Vina, MD simulation was carried out using\n Amber 24, and theozyme model was computed using Gaussian 16. ### Files and\n variables #### File: OYE1-WT-P-R1-soft_docking.pse Description:\n Soft-docking results for the wild-type OYE1–carbon radical R1 complex.\n #### File: MD_NAC-OYE1-WT-R1.pdb **Description:** Selected near-attack\n conformation for hydrogen atom transfer obtained from a restrained MD\n simulation of wild-type OYE1 bound to R1. #### File: MD_NAC_R-OYE1-M5a.pdb\n **Description:** Representative near-attack conformation for hydrogen atom\n transfer leading to the *R*-product, obtained from restrained MD\n simulations of the OYE1-M5a–R1 complex. #### File: MD_NAC_S-OYE1-M5a.pdb\n **Description:** Representative near-attack conformation for hydrogen atom\n transfer leading to the *S*-product, obtained from restrained MD\n simulations of the OYE1-M5a–R1 complex. #### File: MD-OYE1-M6.pdb\n **Description:** Representative conformation obtained from MD simulations\n of the OYE1-M6–R1 complex. #### File:\n Cartesian_coordinates_of_all_optimized_structures.txt\n **Description:** Cartesian coordinates of all stationary points located by\n DFT calculations. ## Code/software The deposited data can be accessed\n using freely available software. The deposited `.pse` and `.pdb` files can\n be opened using the open-source version of PyMOL. The Cartesian\n coordinates in the TXT document can be copied into a compatible XYZ or\n Gaussian input file and visualized with PyMOL or GaussView."}],"geoLocations":[],"fundingReferences":[{"funderIdentifierType":"ROR","funderName":"Center for Advanced Bioenergy and Bioproducts Innovation","funderIdentifier":"https://ror.org/002gqkt51","awardNumber":"DE-SC0018420"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.z34tmpgwg","contentUrl":null,"metadataVersion":0,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-21T07:56:19Z","registered":"2026-08-21T07:56:21Z","published":null,"updated":"2026-08-21T07:56:21Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.ksn02v7jq","type":"dois","attributes":{"doi":"10.5061/dryad.ksn02v7jq","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["University of Pennsylvania","Boston University"],"name":"Glass, Benjamin","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0002-2288-6389"}]},{"nameType":"Personal","affiliation":["University of Pennsylvania"],"name":"Barott, Katie","nameIdentifiers":[]}],"titles":[{"title":"Data and code from: Nematostella developmental plasticity"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"subject":"Cnidarian"},{"subject":"development"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Physiology","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Gene expression","subjectScheme":"PLOS Subject Area Thesaurus"},{"subject":"carryover effects"},{"subject":"plasticity"}],"contributors":[],"dates":[{"date":"2025-11-17T14:07:08Z","dateType":"Created"},{"date":"2026-08-14T13:18:29Z","dateType":"Submitted"},{"date":"2026-08-21T00:00:00Z","dateType":"Issued"},{"date":"2026-08-21T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[{"relationType":"IsDerivedFrom","relatedIdentifier":"10.5281/zenodo.18883339","relatedIdentifierType":"DOI"}],"relatedItems":[],"sizes":["9346552 bytes"],"formats":[],"version":"8","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"This repository contains data and code associated with a study\n investigating temperature-induced developmental plasticity in early life\n stages of the sea anemone Nematostella vectensis. Files are associated\n with a manuscript by Glass and Barott."},{"descriptionType":"TechnicalInfo","description":"# Data and code from: Nematostella developmental plasticity (Glass and\n Barott) This dataset includes files for data associated with a manuscript\n by Glass and Barott. Within this README, \"Nvec\" indicates the\n sea anemone Nematostella vectensis, the species investigated in this\n study. The overall aim of this study was to determine how acute heat shock\n influences developmental outcomes. For the data described below, treatment\n groups include larvae exposed to ambient conditions or those exposed to a\n brief heat shock (1 h at 39°C) at day 3 post-fertilization. Following the\n heat shock treatment, larvae were raised for several weeks, which included\n settlement into the juvenile stage, and various metrics were quantified.\n ## Data and file structure ### File 1: Settlement_data.csv This file is\n for data pertaining to the quantification of settlement rates in Nvec.\n Settlement refers to the transition from the motile planula life stage to\n the benthic juvenile stage, and was determined via morphological\n observations through bright field microscopy. Settlement rates were\n compared for animals exposed to ambient vs. heat shock treatments as well\n as within treatment over time. This file has headers, which are: -\n Treatment = experimental treatment group - Group = unique letter assigned\n to groups of larvae (biological replication within treatments) -\n Day_post_fertilization = the day after fertilization on which animals were\n observed - Animals_settled = number of animals in group settled into\n juvenile stage - Total_number = number of animals in group being observed\n - Proportion_settled = Animals_settled/Total_number - Percent_settled =\n Proportion_settled*100 ### File 2: Size_data.csv This file is for data\n pertaining to the quantification of size in Nvec larvae/juveniles.\n Measurements were collected from images of organisms taken using bright\n field microscopy, and size was compared for animals exposed to ambient vs.\n heat shock treatments as well as within each treatment over time. This\n file has headers, which are: - Treatment = experimental treatment group -\n Group = unique letter assigned to groups of larvae (biological replication\n within treatments) - Day_post_fertilization = the day after fertilization\n on which animals were observed - Length_cm = animal length expressed in cm\n - Length_mm = animal length expressed in mm - Width_cm = animal width\n expressed in cm - Width_mm = animal width expressed in mm - Aspect_ratio =\n length/width - Volume_x100_mm_3 = animal volume expressed in mm^3 x 100 -\n Surface_area_x100_mm_2 = animal surface area expressed in mm^2 x 100 -\n Surface_area_volume_ratio_mm_neg_1 =\n Surface_area_x100_mm_2/Volume_x100_mm_3 ### File 3: Lipid_data.csv This\n file is for data pertaining to the quantification of lipid content in Nvec\n larvae. Measurements were collected using laser scanning confocal\n microscopy in larvae at days 3 and 7 post-fertilization, and results were\n compared for animals exposed to ambient vs. heat shock treatments as well\n as over time. This file has headers, which are: - Treatment = experimental\n treatment group - Animal_ID = unique identifying number assigned to each\n larva in which lipids were quantified - Day_post_fertilization = the day\n after fertilization on which animals were observed - Reduced_fluorescence\n = average fluorescence pixel value (0-256) in reduced range for larva\n being imaged - Reduced_background = average fluorescence pixel value for\n background of image - Reduced_corrected = Reduced_fluorescence -\n Reduced_background - Oxidized_fluorescence = average fluorescence pixel\n value (0-256) in oxidized range for larva being imaged -\n Oxidized_background = average fluorescence pixel value for background of\n image - Oxidized_corrected = Oxidized_fluorescence - Oxidized_background -\n LPO_ratio = Reduced_corrected/Oxidized_corrected ### File 4:\n Respiration_data.csv This file is for data pertaining to the\n quantification of aerobic respiration rates in Nvec larvae. Aerobic\n respiration (i.e., oxygen consumption) rates were measured using a\n microplate-based respirometry rig, and rates were compared for animals\n exposed to ambient vs. heat shock conditions as well as over time. This\n file has headers, which are: - Treatment = experimental treatment group -\n Day_post_fertilization = the day after fertilization on which animals were\n observed - Animals = number of animals in group for which respiration\n rates were measured - Protein_ug = total protein (in ug) for animals in\n group for which respiration rates were measured - Protein_ug_per_animal =\n Protein_ug / Animals - O2_consumed_nmol_min = rate of oxygen consumption\n expressed as nmol per minute - O2_consumed_pmol_min_ug_protein =\n (O2_consumed_nmol_min/Protein_ug)*100 ### File 5:\n Juvenile_heat_tolerance_data.csv This file is for data pertaining to the\n quantification of heat tolerance in Nvec juveniles. Heat tolerance was\n determined in juveniles by exposing them to a short heat ramp (1 h at\n temperatures ranging from 39–43°C) and measuring survival after 48 h.\n Survival was compared for animals exposed to ambient vs. heat shock\n conditions. This file has headers, which are: - Treatment = experimental\n treatment group - Group = unique letter assigned to groups of larvae\n (biological replication within treatments) - Ramp_temperature_C = peak\n temperature of ramp to which juveniles were exposed - Percent_surviving =\n percentage of juveniles surviving exposure to heat ramp -\n Proportion_surviving = proportion of juveniles surviving exposure to heat\n ramp ### File 6: Term2gene_list.csv This file is used for gene ontology\n (GO) analysis of RNA sequencing data, and contains a reference list of\n genes with corresponding GO terms. This file has headers, which are: -\n term = gene ontology term - gene = genes in the Nvec genome ### File 7:\n Term2name_list.csv This file is used for gene ontology (GO) analysis of\n RNA sequencing data, and contains a reference list of GO terms with the\n corresponding names of those terms. This file has headers, which are: -\n term = gene ontology term - name = name of gene ontology term ### File 8:\n Term2ont_list.csv This file is used for gene ontology (GO) analysis of RNA\n sequencing data, and contains a reference list of GO terms with the\n corresponding ontologies (a type of category for GO terms) of those terms.\n This file has headers, which are: - term = gene ontology term - ontology =\n ontology corresponding to GO term ### File 9: Sample_list.csv This file is\n used for analysis of RNA sequencing data, and contains a list of the\n samples from which such data were collected with corresponding metadata.\n This file has headers, which are: - Sample_name = unique identifying name\n for each sample - Sample_ID = same as sample_name - Day_post_fertilization\n = day after fertilization on which samples was collected for later\n processing - Treatment = experimental treatment group from which sample\n was collected - Group = unique letter assigned to groups of animals\n (biological replication within treatments) ### File 10:\n Summary_count_table.txt This file contains the RNA sequencing read counts\n for each gene and sample. The row names for the table are the gene\n identification numbers (see File 9), while the column names are the sample\n identification numbers as defined in the sample list (see File 10). ##\n Sharing/Access information Links to other publicly accessible locations of\n the data: * NA Data was derived from the following sources: * NA ##\n Code/Software ### File 1: Nemat_heat_shock_dev_code.Rmd This file contains\n code in R markdown format for producing analyses and visualizations of the\n data. Most of the script can be run using the files in this repository. To\n run the script in full, one additional file is required, and this file is\n provided as Additional file 3 in Cole et al. (2024) Frontiers in Zoology,\n available at\n [https://doi.org/10.1186/s12983-024-00529-z](https://doi.org/10.1186/s12983-024-00529-z). To run the script, download the file from Cole et al., save Worksheet 1 as a CSV file with the name \"Cole_NV2_genes.csv\", and place the file in the same directory as the script, which should then run in full."}],"geoLocations":[],"fundingReferences":[{"funderIdentifierType":"ROR","funderName":"U.S. National Science Foundation","funderIdentifier":"https://ror.org/021nxhr62","awardTitle":"\n        Postdoctoral Fellowship: OCE-PRF: Effects of dual anthropogenic\n        stressors across life history in a reef-building coral\n      ","awardNumber":"2506815"},{"funderIdentifierType":"ROR","funderName":"Division of Ocean Sciences","funderIdentifier":"https://ror.org/05wqqhv83","awardTitle":"\n        CAREER: Helping or hindering? Determining the influence of repetitive\n        marine heatwaves on acclimatization of reef-building corals across\n        biological scales\n      ","awardNumber":"2237658"},{"funderIdentifierType":"ROR","funderName":"University of Pennsylvania","funderIdentifier":"https://ror.org/00b30xv10","awardTitle":"Dissertation completion fellowship (to B.H.G.)"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.ksn02v7jq","contentUrl":null,"metadataVersion":0,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-21T00:14:02Z","registered":"2026-08-21T00:14:04Z","published":null,"updated":"2026-08-21T00:14:04Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.tqjq2bwf8","type":"dois","attributes":{"doi":"10.5061/dryad.tqjq2bwf8","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["Oregon State University"],"name":"Freedman, Jared","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0002-5430-1947"}]},{"nameType":"Personal","affiliation":["United States Geological Survey"],"name":"Kennedy, Theodore","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Oregon State University"],"name":"Burke, Molly","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Oregon State University"],"name":"Lytle, Dave","nameIdentifiers":[]}],"titles":[{"title":"Data from: Environmental and spatial heterogeneities shape aquatic communities at tributary confluences in a desert river network"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Community ecology","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Freshwater ecology","subjectScheme":"PLOS Subject Area Thesaurus"},{"subject":"aquatic invertebrates"},{"subject":"Environmental DNA"}],"contributors":[],"dates":[{"date":"2026-07-22T01:41:48Z","dateType":"Created"},{"date":"2026-08-17T21:18:26Z","dateType":"Submitted"},{"date":"2026-08-20T00:00:00Z","dateType":"Issued"},{"date":"2026-08-20T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[],"relatedItems":[],"sizes":["488844 bytes"],"formats":[],"version":"4","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Landscape heterogeneities influence dispersal between habitats and\n environmental conditions within habitats, shaping the metacommunity\n dynamics that govern the distribution of species throughout an ecosystem.\n Riverine aquatic communities are sensitive to heterogeneity in flow,\n substrate, and other abiotic conditions arising from interactions between\n aquatic and terrestrial landscapes, as well as from spatial heterogeneity\n in river networks that can cause isolation and fragmentation. Landscape\n characteristics can be particularly variable at tributary confluences,\n where branches with different biotic and abiotic characteristics join\n together. Here, we investigate the effect of tributary confluences on\n aquatic invertebrate community connectivity in Grand Canyon, testing the\n competing hypotheses that a) tributary confluences facilitate community\n connectivity by linking habitats and inducing gradual shifts in habitat\n conditions, or b) differences in adjacent tributary and mainstem habitat\n conditions are sufficient to isolate tributary communities from nearby\n mainstem communities. Using invertebrate-specific environmental DNA (eDNA)\n metabarcoding, we quantified aquatic invertebrate community structure at\n 22 tributary confluences in Grand Canyon, with samples collected at the\n tributary mouths and in the mainstem Colorado River. We identified 667\n aquatic invertebrate taxa in total, with assemblages dominated by Diptera,\n and tributary communities possessed higher taxonomic richness than\n mainstem communities. Community structure at confluences differed between\n tributary and mainstem habitats, with high beta diversity driven by\n species turnover rather than nestedness. Between tributary communities\n throughout Grand Canyon, beta diversity was again driven by high turnover,\n with environmental heterogeneity explaining more of the community\n structure than spatial relationships of the network. The significance of\n turnover at multiple scales suggests that changes in community composition\n are driven by local dynamics of species sorting and habitat filtering, and\n that ecological isolation is an important dynamic in structuring aquatic\n invertebrate communities. Tributaries contribute to the high aquatic\n invertebrate diversity in Grand Canyon, but their isolation relative to\n each other leaves communities vulnerable to habitat degradation and\n stochastic disturbances. Together, these results suggest that landscape\n heterogeneity at tributary-mainstem confluences limits ecological\n connectivity even at short geographical distances, resulting in an\n archipelago-like pattern of disjunct tributary communities despite the\n hydrological connectivity within the river network."},{"descriptionType":"TechnicalInfo","description":"# Data from: Environmental and spatial heterogeneities shape aquatic\n communities at tributary confluences in a desert river network Dataset\n DOI: [10.5061/dryad.tqjq2bwf8](https://doi.org/10.5061/dryad.tqjq2bwf8) ##\n Description of the data and file structure Data and R code provided here\n were used for bioinformatics and analysis of eDNA metabarcoding data from\n Lower Colorado River Basin aquatic habitats collected in spring 2022, as a\n part of a manuscript submitted to the journal Ecography: **The farther you\n are, the closer you get: Environmental similarity connects aquatic\n invertebrate communities across large river networks via ecological\n convergence.** eDNA amplicons of aquatic invertebrate COI gene fragments\n were produced using the fwhF2/EPTDr2n primer set, and sequenced on\n Illumina NextSeq 2000. Raw sequencing reads can be found at NCBI\n BioProject\n PRJNA1347436: [https://dataview.ncbi.nlm.nih.gov/object/PRJNA1347436?reviewer=cf84g9qi1hdfkmj09ffkitlbnm](https://dataview.ncbi.nlm.nih.gov/object/PRJNA1347436?reviewer=cf84g9qi1hdfkmj09ffkitlbnm).  Sequencing reads were processed using the JAMP metabarcoding pipeline ([https://github.com/VascoElbrecht/JAMP](https://github.com/VascoElbrecht/JAMP)). Paired-end reads were merged and primer sequences trimmed, followed by filtering for sequence length the 142bp +/- 10bp. Reads were then filtered for a max EE of 0.5, and denoised with an alpha value of 5. OTUs were then clustered using a 3% similarity threshold, and OTUs were transformed to presence-absence for further community analysis. Taxonomic identifications were made by aligning OTU consensus sequences to the BOLD reference database of COI barcodes ([https://boldsystems.org](https://boldsystems.org/)), using the BOLDigger3 Python package ([https://github.com/DominikBuchner/BOLDigger3](https://github.com/DominikBuchner/BOLDigger3)).  ### Files and variables #### File: 2022_GC_eDNA_Sequencing.zip **Description:** ZIP file containing R code used for processing, filtering, and clustering of eDNA metabarcoding sequencing reads. For details on bioinformatics methods used, see Appendix 1, Section S3. All code is in R version 4.5.3. Required R files include: * 2022_GC_eDNA_combined_JAMP.R -- Code that loads file paths to the eDNA data on Oregon State University's HPC, and calls the JAMP_pipeline() function to run the bioinformatics pipeline. * Required packages: *R.utils* v2.13.0 * JAMP_pipeline_function.R -- Function that calls the steps of JAMP vX.x in the correct order, and renames the folders using JAMP_folder_rename(). More information on the JAMP bioinformatics pipeline can be found here: [https://github.com/VascoElbrecht/JAMP](https://github.com/VascoElbrecht/JAMP).  * Required packages: *JAMP* v0.77 * demulti_file_rename.R -- Helper function to trim demultiplexed file names (as provided by OSU's Center for Quantitative Life Sciences) to sample codes. * JAMP_folder_rename.R -- Helper function to provide a more descriptive folder name than the automated output folders from JAMP. #### File: NCBI_SRA_Metadata.csv **Description:** Metadata file that provides sample IDs, accession numbers, and sequencing information associated with NCBI Sequence Read Archive data for this manuscript.  ##### Variables * Library ID: Unique short identifier assigned to a specific sequencing library * Title: A short descriptive title for each sequencing library * SRA accession: NCBI-assigned unique accession number for each individual library * BioSample accession: NCBI-assigned unique accession number for each physical sample * BioProject accession: NCBI-assigned accession number for the overarching project * Library strategy: Sequencing approach for producing the library * Library source: Type of target material isolated for sequencing * Library selection: Method used to enrich the target nucleic acid * Library layout: Whether sequencing data is single or paired ended * Platform: Sequencing technology brand used to generate the data. * Instrument model: Specific model of sequencing implement used. * Filetype: Format of submitted raw data files. * Filename: exact name of the raw sequence files uploaded to NCBI. #### File: GC_TribMain_Env_Data.csv **Description:** Data sheet containing site metadata and environmental measurements for sampled aquatic habitat in Grand Canyon. Cells containing \"NA\" represent environmental variables that were not measured or obtainable for the given location. This is often due to difficulty in obtaining substrate and land-use data for habitats in the mainstem Colorado River. ##### Variables * sample: Shorthand site code name * site: Name of sampling location * site_pair: Paired sampling site sharing the same tributary confluence  * stream_type: Identity as mainstem or tributary site * river_mile: Distance in miles along Colorado River channel downstream from Lee's Ferry. River mile is a standard river distance measure used in Grand Canyon * Temp: Measured water temperature at sampling location, in celcius * pH: Measured water pH at sampling location * Cond: Measured water specific conductivity at sampling location, in μS/cm * DO: Measured water dissolved oxygen at sampling location, in mg/L * D15: 15th percentile particle size, measured during a Wolman pebble count * D50: 50th percentile particle size, measured during a Wolman pebble count * D85: 85th percentile particle size, measured during a Wolman pebble count * DRNAREA_sqmi: Drainage area of basin upstream from sampling location, in square miles * PRECIP_in: Annual precipitation at site, in inches * RELIEF_ft: Vertical relief from sampling location to highest point in the drainage basin, in feet * FLOOD_50: Magnitude of the estimated 2 year interval flood, in feet^3^/second * FLOOD_1: Magnitude of the estimated 100 year interval flood, in feet^3^/second * FLOOD_002: Magnitude of the estimated 500 year flood interval flood, in feet^3^/second * LC01BAR: Percentage of area barren land, NLCD 2001 category 31 * LC01DEV: Percentage of land-use from NLCD 2001 classes 21-24 * LC01FOREST: Percentage of forest from NLCD 2001 classes 41-43 * LC01HERB: Percentage of herbaceous upland from NLCD 2001 class 71  #### File: combined_edna_dataset_analysis_updated.R **Description:** R script for community analysis included in the manuscript, including all statistical analyses and figures (excluding Figure 1). R version 4.5.3 * Required packages: *betapart  *v1.6.1; *bipartite* v2.24;* cooccur* v1.3; *dplyr* v1.2.1; *ecole* v0.9.2021; *forcats* v1.0.1; *ggpattern* v1.3.1; *ggplot2* v4.0.3; *ggpubr* v1.0.0; *ggsci* v5.1.0; *gridExtra* v2.3.1; *igraph* v2.2.3;* lsd* v1.0.0;* RColorBrewer* v1.1.3; *reshape2* v 1.4.5; *usedist* v0.4.0; *vegan* v2.7.5; *visNetwork* v2.1.4 #### File: network_helper_functions.r **Description:** R script containing custom functions to streamline network modularity and nestedness analysis. R version 4.5.3. * Required packages: *bipartite* v2.24; *ggplot2* v4.0.3; *igraph* v2.2.3;* vegan* v2.7.5 #### File: Taxa_List.csv **Description:** Output from BOLDigger alignment of aquatic invertebrate OTUs (derived from eDNA metabarcoding sequences) to the BOLD reference database of COI barcodes, and OTU detections by site. Only OTUs with \u0026gt;85% similarity to a reference sequence were retained. Binary numbers indicated the presences (1) or absences (0) of OTUs at a given sampling location. ##### Variables * OTU_ID: OTU number * Class: Class of top COI barcode ID * Order: Order of top COI barcode ID (\u0026gt;85%) * Family: Family of top COI barcode ID (\u0026gt;90%) * Genus: Genus of top COI barcode ID (\u0026gt;95%) * Species: Species of top COI barcode ID (\u0026gt;97%) * Similarity: Percent similarity of OTU consensus sequence to top COI barcode * PAR: Paria River * Nank: Nankoweap Creek * LCR: Little Colorado River * CLR: Clear Creak * BA: Bright Angel Creek * Her: Hermit Creek * Bou: Boucher Creek * Cry: Crystal Creek * Shi: Shinumo Creek * RAC: Royal Arch Creek * Tap: Tapeats Creek * Deer: Deer Creek * Kan: Kanab Creek * Mat: Matkatamiba Creek * Hav: Havasu Creek * VW: Vulcan's Well * Spr: Spring Creek * TS: Three Springs Creek * Dia: Diamond Creek * Tra: Travertine Canyon * Spe: Spencer Creek * Col: Columbine Creek * CR01: Colorado River Mile 01 * CR19: Colorado River Mile 19 * CR52: Colorado River Mile 52 * CR61: Colorado River Mile 61 * CR76: Colorado River Mile 76 * CR84: Colorado River Mile 84 * CR88: Colorado River Mile 88 * CR95: Colorado River Mile 95 * CR97: Colorado River Mile 97 * CR98:  Colorado River Mile 98 * CR109: Colorado River Mile 109 * CR117: Colorado River Mile 117 * CR134: Colorado River Mile 134 * CR136: Colorado River Mile 136 * CR144: Colorado River Mile 144 * CR148: Colorado River Mile 148 * CR157: Colorado River Mile 157 * CR175: Colorado River Mile 175 * CR180: Colorado River Mile 180 * CR205: Colorado River Mile 205 * CR215: Colorado River Mile 215 * CR226: Colorado River Mile 226 * CR229: Colorado River Mile 229 * CR246: Colorado River Mile 246 * CR275: Colorado River Mile 275 #### File: GC_TribMain_OTU_Table.csv **Description:** OTU x Site community matrix, for each individual PCR replicate. Each site is replicated twice. Numbers correspond to metabarcoding read count.   ##### Variables * OTU_ID: Operational Taxonomic Unit number * ESV_count: Number of exact sequence variants within each OTU * haplotype: Top ESV associated with each OTU * OTU: Operational Taxonomic Unit * BA_1_S9_PE: Bright Angel Creek, Replicate 1 * BA_2_S10_PE: Bright Angel Creek, Replicate 2 * Bou_1_S13_PE: Boucher Creek, Replicate 1 * Bou_2_S14_PE: Boucher Creek, Replicate 2 * CLR_1_S7_PE: Clear Creek, Replicate 1 * CLR_2_S8_PE: Clear Creek, Replicate 2 * Col_1_S43_PE: Columbine Creek, Replicate 1 * Col_2_S44_PE: Columbine Creek, Replicate 2 * CR01_1_S45_PE: Colorado River Mile 01, Replicate 1 * CR01_2_S46_PE: Colorado River Mile 01, Replicate 2 * CR109_1_S65_PE: Colorado River Mile 109, Replicate 1 * CR109_2_S66_PE: Colorado River Mile 109, Replicate 2 * CR117_1_S67_PE: Colorado River Mile 117, Replicate 1 * CR117_2_S68_PE: Colorado River Mile 117, Replicate 2 * CR134_1_S69_PE: Colorado River Mile 134, Replicate 1 * CR134_2_S70_PE: Colorado River Mile 134, Replicate 2 * CR136_1_S71_PE: Colorado River Mile 136, Replicate 1 * CR136_2_S72_PE: Colorado River Mile 136, Replicate 2 * CR144_1_S73_PE: Colorado River Mile 144, Replicate 1 * CR144_2_S74_PE: Colorado River Mile 144, Replicate 2 * CR148_1_S75_PE: Colorado River Mile 148, Replicate 1 * CR148_2_S76_PE: Colorado River Mile 148, Replicate 2 * CR157_1_S77_PE: Colorado River Mile 157, Replicate 1 * CR157_2_S78_PE: Colorado River Mile 157, Replicate 2 * CR175_1_S79_PE: Colorado River Mile 175, Replicate 1 * CR175_2_S80_PE: Colorado River Mile 175, Replicate 2 * CR180_1_S81_PE: Colorado River Mile 180, Replicate 1 * CR180_2_S82_PE: Colorado River Mile 180, Replicate 2 * CR19_1_S47_PE: Colorado River Mile 19, Replicate 1 * CR19_2_S48_PE: Colorado River Mile 19, Replicate 2 * CR205_1_S83_PE: Colorado River Mile 205, Replicate 1 * CR205_2_S84_PE: Colorado River Mile 205, Replicate 2 * CR215_1_S85_PE: Colorado River Mile 215, Replicate 1 * CR215_2_S86_PE: Colorado River Mile 215, Replicate 2 * CR226_1_S87_PE: Colorado River Mile 226, Replicate 1 * CR226_2_S88_PE: Colorado River Mile 226, Replicate 2 * CR229_1_S89_PE: Colorado River Mile 229, Replicate 1 * CR229_2_S90_PE: Colorado River Mile 229, Replicate 2 * CR246_1_S91_PE: Colorado River Mile 246, Replicate 1 * CR246_2_S92_PE: Colorado River Mile 246, Replicate 2 * CR275_1_S93_PE: Colorado River Mile 275, Replicate 1 * CR275_2_S94_PE: Colorado River Mile 275, Replicate 2 * CR52_1_S49_PE: Colorado River Mile 52, Replicate 1 * CR52_2_S50_PE: Colorado River Mile 52, Replicate 2 * CR61_1_S51_PE: Colorado River Mile 61, Replicate 1 * CR61_2_S52_PE: Colorado River Mile 61, Replicate 2 * CR76_1_S53_PE: Colorado River Mile 76, Replicate 1 * CR76_2_S54_PE: Colorado River Mile 76, Replicate 2 * CR84_1_S55_PE: Colorado River Mile 84, Replicate 1 * CR84_2_S56_PE: Colorado River Mile 84, Replicate 2 * CR88_1_S57_PE: Colorado River Mile 88, Replicate 1 * CR88_2_S58_PE: Colorado River Mile 88, Replicate 2 * CR95_1_S59_PE: Colorado River Mile 95, Replicate 1 * CR95_2_S60_PE: Colorado River Mile 95, Replicate 2 * CR97_1_S61_PE: Colorado River Mile 97, Replicate 1 * CR97_2_S62_PE: Colorado River Mile 97, Replicate 2 * CR98_1_S63_PE: Colorado River Mile 98, Replicate 1 * CR98_2_S64_PE: Colorado River Mile 98, Replicate 2 * Cry_1_S15_PE: Crystal Creek, Replicate 1 * Cry_2_S16_PE: Crystal Creek, Replicate 2 * Deer_1_S23_PE: Deer Creek, Replicate 1 * Deer_2_S24_PE: Deer Creek, Replicate 2 * Dia_1_S37_PE: Diamond Creek, Replicate 1 * Dia_2_S38_PE: Diamond Creek, Replicate 2 * Hav_1_S29_PE: Havasu Creek, Replicate 1 * Hav_2_S30_PE: Havasu Creek, Replicate 2 * Her_1_S11_PE: Hermit Creek, Replicate 1 * Her_2_S12_PE: Hermit Creek, Replicate 2 * Kan_1_S25_PE: Kanab Creek, Replicate 1 * Kan_2_S26_PE: Kanab Creek, Replicate 2 * LCR_1_S5_PE: Little Colorado River, Replicate 1 * LCR_2_S6_PE: Little Colorado River, Replicate 2 * Mat_1_S27_PE: Matkatamiba Creek, Replicate 1 * Mat_2_S28_PE: Matkatamiba Creek, Replicate 2 * Nank_1_S3_PE: Nankoweap Creek, Replicate 1 * Nank_1_S3_PE.1: Nankoweap Creek, Replicate 1 (duplicate/technical replicate) * Nank_2_S4_PE: Nankoweap Creek, Replicate 2 * P_con_S95_PE: Positive Control * PAR_1_S1_PE: Paria River, Replicate 1 * PAR_2_S2_PE: Paria River, Replicate 2 * RAC_1_S19_PE: Royal Arch Creek, Replicate 1 * RAC_2_S20_PE: Royal Arch Creek, Replicate 2 * Shi_1_S17_PE: Shinumo Creek, Replicate 1 * Shi_2_S18_PE: Shinumo Creek, Replicate 2 * Spe_1_S41_PE: Spencer Creek, Replicate 1 * Spe_2_S42_PE: Spencer Creek, Replicate 2 * Spr_1_S33_PE: Spring Creek, Replicate 1 * Spr_2_S34_PE: Spring Creek, Replicate 2 * Tap_1_S21_PE: Tapeats Creek, Replicate 1 * Tap_2_S22_PE: Tapeats Creek, Replicate 2 * Tra_1_S39_PE: Travertine Canyon, Replicate 1 * Tra_2_S40_PE: Travertine Canyon, Replicate 2 * TS_1_S35_PE: Three Springs Creek, Replicate 1 * TS_2_S36_PE: Three Springs Creek, Replicate 2 * VW_1_S31_PE: Vulcan's Well, Replicate 1 * VW_2_S32_PE: Vulcan's Well, Replicate 2 * sequences: nucleotide sequence for each OTU ## Access information Other publicly accessible locations of the data: * Raw sequencing reads can be found at NCBI BioProject **PRJNA1347436: **[https://dataview.ncbi.nlm.nih.gov/object/PRJNA1347436?reviewer=cf84g9qi1hdfkmj09ffkitlbnm](https://dataview.ncbi.nlm.nih.gov/object/PRJNA1347436?reviewer=cf84g9qi1hdfkmj09ffkitlbnm). "}],"geoLocations":[],"fundingReferences":[{"funderIdentifierType":"ROR","funderName":"United States Bureau of Reclamation","funderIdentifier":"https://ror.org/00ezrrm21","awardTitle":"Glen Canyon Dam Adaptive Management Program"},{"funderIdentifierType":"ROR","funderName":"United States Geological Survey","funderIdentifier":"https://ror.org/035a68863","awardTitle":"Water Quality Partnership Program"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.tqjq2bwf8","contentUrl":null,"metadataVersion":0,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-20T23:13:47Z","registered":"2026-08-20T23:13:48Z","published":null,"updated":"2026-08-20T23:13:48Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.98sf7m0xg","type":"dois","attributes":{"doi":"10.5061/dryad.98sf7m0xg","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["University of Oxford"],"name":"Hending, Daniel","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0003-0609-4354"}]},{"nameType":"Personal","affiliation":["University of Oxford"],"name":"Stewart-Roberts, Sky","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Oxford"],"name":"Rocha, Ricardo","nameIdentifiers":[]}],"titles":[{"title":"Environmental drivers of bat distribution in Madagascar"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Climate change","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Bats","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Geographic distribution","subjectScheme":"PLOS Subject Area Thesaurus"}],"contributors":[],"dates":[{"date":"2025-11-18T07:40:09Z","dateType":"Created"},{"date":"2025-11-18T07:40:21Z","dateType":"Submitted"},{"date":"2026-08-20T00:00:00Z","dateType":"Issued"},{"date":"2026-08-20T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[{"relationType":"IsCitedBy","relatedIdentifier":"10.1111/ddi.70254","relatedIdentifierType":"DOI"}],"relatedItems":[],"sizes":["186264973 bytes"],"formats":[],"version":"5","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"This dataset contains 1) an excel file containing occurrence points\n (updated as of January 2025) for all bat species in Madagascar extracted\n from published papers and online data repositories (iNaturalist, GBIF),\n and 2) raster data for the 19 BioClim variables at the current time\n (averaged for seven climate models), and for the year 2080 at a\n \"business-as-usual\" climate trajectory. Occurrence data is\n ordered by taxonomic family, and is presented at the same resolution as\n the original data source (i.e., all occurrence points are available at the\n exact same resolution in the original publications/repositories from which\n they were obtained). Raster data is clipped to cover the geographic area\n of Madagascar only. "},{"descriptionType":"TechnicalInfo","description":"# Environmental drivers of bat distribution in Madagascar Dataset DOI:\n [10.5061/dryad.98sf7m0xg](https://doi.org/10.5061/dryad.98sf7m0xg) ##\n Description of the data and file structure All occurrence point data was\n obtained from either the published literature, or online data repositories\n (GBIF, iNaturalist). X and Y coordinates were saved in the same resolution\n as the original data source. Climate data was downloaded from WorldClim.\n ### Files and variables #### File: Occurence_Point_Database.csv\n **Description:** A database of occurrence points for bat species present\n in Madagascar. ##### Variables * Taxonomic family = details on the\n taxonomic family of the corresponding species. * Scientific name = details\n of the scientific name of the corresponding species. * Common name\n = details of the common name of the corresponding species. * Description =\n the year that the species was originally described. * IUCN = the species\n current IUCN Red List classification status. * Endemic = information on\n whether the species is endemic to Madagascar or not. * Latitude = the Y\n coordinate of the occurrence point (at the same resolution as the original\n data source). * Longitude = the X coordinate of the occurrence point (at\n the same resolution as the original data source). * Ref = the reference\n data source. #### File: Current_copy.tif **Description:** A raster layer\n containing data for 19 bioclimatic variables for Madagascar for the\n current time period. Resolution = 1km x 1km. Each individual bioclimatic\n variable is saved as a separate band in the raster layer (19 bands total).\n The variables are as follows: BIO1 = Annual Mean Temperature (degrees)\n BIO2 = Mean Diurnal Range (Mean of monthly (max temp - min\n temp)) (degrees) BIO3 = Isothermality (BIO2/BIO7) (×100) BIO4 =\n Temperature Seasonality (standard deviation ×100) BIO5 = Max Temperature\n of Warmest Month (degrees) BIO6 = Min Temperature of Coldest\n Month (degrees) BIO7 = Temperature Annual Range (BIO5-BIO6) (degrees) BIO8\n = Mean Temperature of Wettest Quarter (degrees) BIO9 = Mean Temperature of\n Driest Quarter (degrees) BIO10 = Mean Temperature of Warmest\n Quarter (degrees) BIO11 = Mean Temperature of Coldest Quarter (degrees)\n BIO12 = Annual Precipitation (mm) BIO13 = Precipitation of Wettest\n Month (mm) BIO14 = Precipitation of Driest Month (mm) BIO15 =\n Precipitation Seasonality (Coefficient of Variation) BIO16 = Precipitation\n of Wettest Quarter (mm) BIO17 = Precipitation of Driest Quarter (mm) BIO18\n = Precipitation of Warmest Quarter (mm) BIO19 = Precipitation of Coldest\n Quarter (mm) #### File: 2080_BAU.zip **Description:** A zip folder\n containing seven raster layers (from seven different future climate\n models) with data for 19 biolcimatic variables for Madagascar for the year\n 2080 using a \"business-as-usual\" RCP 8.5 trajectory. Resolution\n = 1km x 1km. Each individual bioclimatic variable is saved as a separate\n band in the raster layer (19 bands total for each of the seven raster\n layers). Each raster layer contains 70 bands in total, with bands 37-55\n corresponding to the bioclimatic variables 1-19. Bands 1-12 are monthly\n minimum temperatures, bands 13-24 are monthly maximum temperatures, bands\n 25-36 are mean monthly precipitation, bands 56-67 are monthly potential\n evapotranspiration, band 68 represents annual potential\n evapotranspiration, band 69 represents annual climate water deficit, and\n band 70 represents number of dry months. The 19 bioclimatic variables are\n the same as described for the Current_copy.tif. The naming convention for\n the seven raster layers is \"modelname_RCPpathway_year.tif\". All\n temperature values are in degrees Celsius, while all precipitation values\n are in millimeters. ## Code/software Microsoft excel GIS (Arc or Q) ##\n Access information Other publicly accessible locations of the data: * GBIF\n * iNaturalist * Worldclim/Madaclim Data was derived from the following\n sources: * GBIF * iNaturalist * Worldclim/Madaclim"}],"geoLocations":[],"fundingReferences":[],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.98sf7m0xg","contentUrl":null,"metadataVersion":0,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":3,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-20T23:01:28Z","registered":"2026-08-20T23:01:30Z","published":null,"updated":"2026-08-20T23:01:30Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.5mkkwh7jf","type":"dois","attributes":{"doi":"10.5061/dryad.5mkkwh7jf","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["ETH Zurich"],"name":"Wang, Longlong","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0001-5279-1885"}]}],"titles":[{"title":"Data from: A designed ubiquitin-binding protein selectively enriches and visualizes unanchored ubiquitin"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"subject":"ubiquitin"},{"subject":"protein design"},{"subject":"unanchored ubiquitin"}],"contributors":[],"dates":[{"date":"2025-09-07T09:56:31Z","dateType":"Created"},{"date":"2026-08-14T12:08:29Z","dateType":"Submitted"},{"date":"2026-08-20T00:00:00Z","dateType":"Issued"},{"date":"2026-08-20T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[],"relatedItems":[],"sizes":["2423339802 bytes"],"formats":[],"version":"9","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"The datasets included here are the raw data for the accepted manuscript\n \" A designed ubiquitin-binding protein selectively enriches\n and visualizes unanchored ubiquitin\" in Science Advances (First\n author: Shihua Shi). Each folder corresponds to the figures/tables in the\n manuscript. Manuscript abstract: While protein ubiquitination has been\n extensively studied, the roles of unanchored ubiquitin and its chains\n remain less understood, largely due to a lack of specific, high-affinity\n tools for their enrichment and visualization. To address this, we employed\n high-throughput affinity maturation and computational protein design to\n engineer novel proteins that selectively bind unanchored ubiquitin. We\n first enhanced the binding affinity of the zinc finger domain of HDAC6, a\n natural unanchored ubiquitin binder, by including its unstructured\n N-terminal loop and introducing the V1091L mutation. Using RF-diffusion,\n we then designed 11 novel ubiquitin binding proteins (UBiPs), among which\n UBiP10 showed remarkable specificity for unanchored ubiquitin. This\n protein enriched unanchored ubiquitin from cells and influenza A virions\n and served as a probe to determine its cellular levels, supporting a\n relative decrease upon proteasome inhibition. Imaging studies with UBiP10\n further revealed a novel unanchored ubiquitin coat around aggresomes,\n illustrating the potential of UBiP10 as a unique tool to study unanchored\n ubiquitin in cellular regulation."},{"descriptionType":"TechnicalInfo","description":"# Data from: A designed ubiquitin-binding protein selectively enriches and\n visualizes unanchored ubiquitin Dataset DOI:\n [10.5061/dryad.5mkkwh7jf](https://doi.org/10.5061/dryad.5mkkwh7jf) ##\n Description of the data and file structure This dataset includes the data\n and quantification files for generating the figures in the related\n manuscript. All the files generated from RF-diffusion peptide are also\n included to help reproducing. Microscopy files that used for generate\n example figures in the manuscript (aec0818, in press) a also included, but\n for the other raw microscopy files, due to the size, are not uploaded. For\n requesting them, please write to\n [longlong.wang@pharma.ethz.ch](mailto:longlong.wang@pharma.ethz.ch) or\n [longlongwang@outlook.com](mailto:longlongwang@outlook.com) ### Files and\n variables ### File: output_of_UBiP_design.zip **Description:** file for\n evaluating the design; this file contains the selected 11 novel ubiquitin\n binding protein (UBiP) candidates, designed by RF-diffusion/ProteinMPNN.\n The UBiPs are referred to as \"UBP01\" through \"UBP11\"\n in folder/file names. For each UBiPs, the folder contains the command for\n design in the included .sh files, and also the generated Excel file for\n evaluating the design. The variables in the \"results\" CSV files\n are described below, in the section titled \"Variables in tabular\n files\". Relevant FASTA and/or molecular model (.pdb, .pse) files are\n included. Where relevant, .png images of molecular models are also\n provided. If multiple models or designs were tested for a single UBiP, the\n related files are stored within subfolders. Starting models have\n \"starting model\" in the file name. #### File:\n AF2_prediction_for_ipTM_in_table_1.zip **Description:** AlphaFold2\n predictions of 10 of the 11 Ub–UBiP complexes (excluding UBP05), from\n which the ipTM values in Table 1 were taken. Organised by UBiP; each\n folder contains the predicted complex structures (.pdb/.cif), the input\n sequences/alignments (.fasta, .sto), and a \"result\" subfolder\n with the confidence metrics (.json / summary table). Higher ipTM (\u0026gt;0.8)\n indicates a confidently predicted Ub–UBiP interface. Additional\n description of predictions, analyses, and results can be found in the\n \"results.html\" files in each subdirectory. These files can be\n found within the \"predictions\" subfolder of each UBiP-named\n directory. #### File: Fig_1.zip **Description: Affinity maturation of the\n HDAC6 ZnF domain for mono-Ub.** Subfolders by panel: •    **Fig_1c** —\n isothermal titration calorimetry (ITC) of mHDAC6_(1007–1149) (and mutants)\n against mono-Ub: raw .itc plus exported fitting reports. •    **Fig_1f to\n h/PCA analysis** — processed yeast deep-mutational-scanning tables, one\n per assay (BindingPCA, AbundancePCA), one row per variant. The variables\n in these files are described below, in the section titled \"Variables\n in tabular files\". The file\n \"Collapsed_reads_per_sample_abundance_screen_meta.csv\" contains\n metadata about each input and output variable (the counts of sequencing\n reads before and after selection). •    **Fig_1i** — ITC of the V1091L\n mutant against mono-Ub. In the folders for Fig 1c and Fig 1i, associated\n Origin Project (.opj) files are provided alongside the ITC raw files.\n These can be opened in Origin Lab's free viewer software. #### File:\n Fig_2.zip **Description: In silico design of the UBiPs and their affinity\n for mono-Ub.** Subfolders by panel: •    **Fig_2c, Fig_2e, Fig_2f** — ITC\n raw .itc files and exported fitting reports for UBiP05, UBiP10 (folders\n Fig 2c and 2e), and the two UBiP10 mutants (folder Fig 2f). Associated\n Origin Project (.opj) files are provided alongside the ITC raw files.\n These can be opened in Origin Lab's free viewer software. #### File:\n Fig_3.zip **Description:** **selectivity of UBiP10 for unanchored Ub**\n Subfolders by panel: •    **Fig_3b** — raw mass-spectrometry files (in .xy\n format) and exported peak tables of the Lb^pro clipping experiment (PDFs).\n •    **Fig_3c, 3d, 3f, 3g, 3h, 3i, 3j, 3k** — uncropped, unadjusted scans\n of every western blot (.tif/.pdf), one file per membrane/antibody, with\n lane and antibody key. •    **Fig_3e** — ITC of UBiP10 against K48- and\n K63-linked di-Ub (raw .itc files + reports in OPJ format, with relevant\n plots included as TIFFs). •    **Fig_3l** — AP-MS search result tables\n used for the volcano plot. #### File: Fig_4.zip **Description: detection\n and imaging of unanchored Ub in cells.** Subfolders by panel: •    **Fig\n 4a and b** — microscopy acquisitions and exported example images for the\n four cell lines and three treatments, with the ImageJ quantification\n tables and the Prism file used for the bar plot and statistics. Folders\n are named according to cell line (A549, Hela, MEF, and SW1353). •    **Fig\n 4c** — microscopy acquisitions and example images of influenza virions\n (channels: viral HA, unanchored Ub, merge). •    **Fig 4d** and **Fig 4e**\n — microscopy acquisitions for GST-UBiP05, GST-UBiP10 and GST staining,\n plus intensity tables and the Prism file. #### File: Fig_5_new.zip\n **Description:** **an unanchored Ub coat around the aggresome.**\n Subfolders by panel: •    **Fig 5a** and **Fig 5b** — microscopy\n acquisitions and example images, plus the ImageJ line-profile exports\n (distance vs grey value) for the two enlarged aggresomes. The distance vs\n grey value exports are stored in the four CSV files in this folder. For\n each, the two variables are *Distance_(microns)* and *Gray_Value*. •   \n **Fig 5c and g** — free (unanchored) Ub intensity quantification at the\n three positions, as tables plus the Prism file with the plot and ANOVA.\n •    **Fig 5d** — uncropped western blot scans for the myosin-10 KD and\n HDAC6 KO validation. •    **Fig 5e and Extendend Data Fig 6c** —\n microscopy acquisitions and example images of scRNA, siMyo and HDAC6-KO\n cells, with the ImageJ line-profile and intensity tables (in CSV format;\n variables are *Distance_(microns)* and *Gray_Value*). #### File:\n supplementary_figures_new.zip **Description:** One folder per\n supplementary figure (Fig_S1 … Fig_S7), with panel subfolders inside. •   \n **Fig_S1** — affinities of the various HDAC6 constructs for mono-Ub.\n Structural detail of ZnF selectivity (A) and the scheme of the yeast\n mutagenesis assay (D); ITC of hHDAC6_(1073–1215) (B), full-length hHDAC6\n (C) and the mouse mutants Y1039R, A1040K, S1064H and S1064H/V1091L (E–H).\n Files: raw .itc and exported reports (OPJ format) per construct, plus the\n structure files for A and I. •    **Fig S2 and Table S1** — summary of the\n design campaign: six starting models were used and UBiP01–UBiP11 selected\n by RMSD to the RF-diffusion design. Files: one folder per UBiP with the\n .sh command, the RF-diffusion/ProteinMPNN result .csv and the models. The\n variables in the \"results\" .csv files are described below, in\n the section titled \"Variables in tabular files\". •    **Fig S3**\n — purification of UBiP01–UBiP10 and their affinities for mono-Ub (with tUI\n for comparison), the modelled UBiP05–Ub interface, and purification of the\n UBiP10 mutants. Files: one folder per UBiP with raw .itc and reports (OPJ\n format), plus gel scans and structure files. •    **Fig_S4** — interaction\n of UBiP10 with unanchored Ub chains: AlphaFold3 models of UBiP05/UBiP10\n with Ub lacking its C-terminal tail, with linear di-Ub, and of tUI–Ub (A,\n B, F, G); purification of K48/K63 di-Ub (C); superposition with di-Ub\n structures (D, E); immunoblots of HA-tagged construct pulldowns and of GST\n pulldowns from di-Ub-supplemented and Poly (I:C)-treated lysates (H–J);\n IAV uncoating assay measured as the M1-spreading cell ratio (K). Files:\n structures, uncropped blot scans, and the uncoating quantification table\n with its Prism file. •    **Fig_S5** — validation of the UBiPs for\n immunofluorescence: staining workflow (A), background of the GST control\n (B), absence of aggregate staining by UBiP10 (C), ELISA on immobilised\n 3xUb-RARA/mono-Ub mixtures probed with GST-UBiP05 or GST-UBiP10 (D),\n linear fluorescence response up to 10 µg/mL Ub and its calibration (E, F),\n quantification of the Ub-antibody signal from Fig. 4D (G), and the\n estimated unanchored fraction of total Ub (H). Files: microscopy\n acquisitions and example images, the ELISA raw plate readings, and the\n quantification tables with Prism files. •    **Fig_S6** — unanchored Ub\n during aggresome formation: a bortezomib time course imaged with anti-Ub\n and GST-UBiP10 (A); the Fiji mask workflow used for the three-position\n quantifications of Fig. 5C and 5G (B: aggresome = Mask 1, peri-aggresomal\n ring = dilated minus original, remote area = cell outline minus eroded\n outline); and the control scRNA condition (C). Files: microscopy\n acquisitions, example images, and quantification tables reporting distance\n (microns) and gray value. •    **Fig_S7** — the aggresome experiment\n repeated in differentiated SH-SY5Y cells (bortezomib 2 µM, 24 h): imaging\n (A), line profiles (B), three-position quantification of ca. 40 aggresomes\n (C), and a comparison of UBiP10 with a TUBE reagent (D). Files: microscopy\n acquisitions, example images, quantification tables and the Prism file.\n **Variables in tabular files:** **Design / prediction score tables**\n (design output, fig. S2, Table 1): * **design/model** = design identifier;\n * **seq** = designed sequence; * **mpnn** = ProteinMPNN sequence score\n (dimensionless, lower = better); * **plddt** = AlphaFold per-residue\n confidence (0–100, some files 0–1); * **ptm** = fold confidence (0–1); *\n **i_ptm/iptm** = interface confidence (0–1, the value in Table 1); *\n **pae** = predicted aligned error (Å, lower = better); * **rmsd** =\n deviation between the AlphaFold2 prediction and the RF-diffusion design\n (Å, the ranking criterion). * Remaining columns (seed, contig, sample\n number, recycles) are pipeline bookkeeping. **BindingPCA / AbundancePCA\n tables** (Fig. 1F–H): one row per single amino-acid variant — * **Input1–6\n / Input1–4** — read counts before selection (the library as built) *\n **Output1–6 / Output1–4** — counts after selection; the ratio to input is\n the signal * **Zinc finger variant** — the 216 nt (72 codon) mutagenised\n region carried by that variant * **Barcode** — 20 nt tag; 41,722 unique\n barcodes, median ~10 per variant * **Variant Name** — WT amino acid +\n position + mutant, `*` = stop codon * **logFC** — log2(output/input), the\n raw PCA fitness score. Negative = depleted = destabilised (abundance) or\n reduced binding (interaction). * **Relative logFC (dLFC)** —It's\n `logFC` minus the wild-type's logFC, so **WT sits at 0**. * **P-Value\n / Adjusted P-Value** — test of logFC ≠ 0, and its FDR-corrected version *\n **Standard Error logFC** — uncertainty on the estimate; use it to filter\n or weight low-confidence variants * **Wildtype Amino Acid / Mutant Amino\n Acid / Protein Position** — just parsed out of the mutation name for\n convenience ## Code/software No code was generated, and for graph\n plotting, we used graphpad prism 10, and for microscopy picture\n processing, we used ImageJ. ver 1.54p. `.sto` files: Stockholm\n multiple-sequence alignment format; can be read using HMMER, Jalview,\n UGENE, or a text editor. UGENE explicitly supports Stockholm `.sto` files.\n `.itc` files: are raw isothermal titration calorimetry (ITC) data files\n and can be opened and analyzed using the MicroCal PEAQ-ITC Analysis\n Software or compatible MicroCal ITC analysis software. All the code used\n in this manuscript have been cited."}],"geoLocations":[],"fundingReferences":[{"funderIdentifierType":"ROR","funderName":"European Research Council","funderIdentifier":"https://ror.org/0472cxd90","awardNumber":"856581"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.5mkkwh7jf","contentUrl":null,"metadataVersion":0,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-20T21:45:00Z","registered":"2026-08-20T21:45:01Z","published":null,"updated":"2026-08-20T21:45:01Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.0cfxpnwg2","type":"dois","attributes":{"doi":"10.5061/dryad.0cfxpnwg2","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["Environment and Climate Change Canada"],"name":"Fiorino, Giuseppe","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0002-1569-0767"}]},{"nameType":"Personal","affiliation":["Environment and Climate Change Canada"],"name":"Denomme-Brown, Simon","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Environment and Climate Change Canada"],"name":"Rogers, Hayley","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Environment and Climate Change Canada"],"name":"Smith, Ian","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["SUNY Brockport"],"name":"Amatangelo, Kathryn","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Grand Valley State University"],"name":"Cooper, Matthew","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Wisconsin–Superior"],"name":"Danz, Nicholas","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Wisconsin–Green Bay"],"name":"Howe, Robert","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["SUNY Brockport"],"name":"Schultz, Rachel","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Central Michigan University"],"name":"Wheelock, Bridget","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["SUNY Brockport"],"name":"Wilcox, Douglas","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Environment and Climate Change Canada"],"name":"Grabas, Greg","nameIdentifiers":[]}],"titles":[{"title":"Data from: Beyond floristic quality: A data-driven approach to assessing coastal wetland condition using plants"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Earth and related environmental sciences","subjectScheme":"fos"},{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Natural sciences","subjectScheme":"fos"},{"subject":"anthropogenic disturbance"},{"subject":"index of biotic condition"},{"subject":"Laurentian Great Lakes"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Marshes","subjectScheme":"PLOS Subject Area Thesaurus"},{"subject":"Monitoring"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Plant communities","subjectScheme":"PLOS Subject Area Thesaurus"}],"contributors":[],"dates":[{"date":"2025-12-02T14:37:49Z","dateType":"Created"},{"date":"2026-07-02T19:59:28Z","dateType":"Submitted"},{"date":"2026-08-20T00:00:00Z","dateType":"Issued"},{"date":"2026-08-20T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[],"relatedItems":[],"sizes":["303943 bytes"],"formats":[],"version":"10","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Plant communities are commonly used to assess ecological condition due to\n their sensitivity to local environmental factors and the practicality of\n conducting vegetation surveys. Using a large dataset (41,314 survey plots\n from 563 wetlands across 12 years), we developed and tested a new\n plant-based Index of Biotic Condition (pIBC) for coastal wetlands for each\n of the five Laurentian Great Lakes. The pIBC shares conceptual\n similarities with Floristic Quality Assessment (FQA) metrics but\n incorporates species-specific sensitivity and responsiveness to\n anthropogenic disturbance based on modeled probabilities of occurrence\n derived from field data. This distinguishes it from traditional FQA\n metrics, which rely on expert-assigned Coefficients of Conservatism and do\n not incorporate occurrence probability. We found that each lake-specific\n version of the pIBC consistently predicted anthropogenic disturbance, as\n indicated by water quality and surrounding land use. Additionally,\n comparisons with seven other plant-based metrics showed that the pIBC\n generally provides greater within-lake sensitivity to water-quality and\n land-use impacts, providing higher-resolution insight into spatial and\n temporal variation at the lake scale, which is often the most relevant\n scale for management decisions. The pIBC was also applicable across\n wetlands of different hydrogeomorphic types and under different\n water-level conditions. The pIBC’s strong performance suggests that it is\n well-suited for assessing coastal wetland condition across sites and\n within sites through time. Overall, this new index is a conceptually\n grounded and statistically robust tool for conservation practitioners that\n is easy to calculate and interpret."},{"descriptionType":"Methods","description":"\u003cem\u003eStudy Area\u003c/em\u003e We used data\n collected from 2011 to 2022 by the Great Lakes Coastal Wetland Monitoring\n Program (https://greatlakeswetlands.org/; Uzarski et al. 2017; Uzarski et al. 2019). The pool of study sites was all coastal wetlands in the Great Lakes basin greater than 4 ha in size with a permanent or periodic surface-water connection to an adjacent Great Lake or connecting river system. Coastal wetlands were selected for sampling using a random sampling protocol stratified by (1) wetland hydrogeomorphic type (lacustrine, riverine, barrier protected; Albert et al. 2005), (2) region (northern or southern; Danz et al. 2005), and (3) lake (with connecting channels included as part of the downstream Great Lake) (Uzarski et al. 2017; Uzarski et al. 2019). We sampled roughly 20% of all wetlands in each stratum each year, so that nearly all coastal wetlands within the Great Lakes basin meeting the selection criteria were sampled at least once every five years. In addition, we resampled 10% of wetlands according to a rotating panel design. Sampled wetlands were dominated by emergent, herbaceous vegetation and shallow water (\u0026lt;2 m deep) containing floating and/or submerged vegetation. We used data from 563 wetlands that were sampled an average of 2.3 times over the 12-year study period, representing 1,273 wetland-year sampling events. The number of wetlands sampled for vegetation in each lake was 73 for Lake Superior (155 sampling events), 87 for Lake Michigan (197), 194 for Lake Huron (456), 66 for Lake Erie (150), and 143 for Lake Ontario (315). \u003cem\u003ePlant Surveys\u003c/em\u003e Vegetation sampling for each wetland occurred between June and August along three transects established perpendicular to the depth contours of the wetland that crossed through selected vegetation zones. Major vegetation zones were defined as the wet meadow zone, emergent vegetation zone, and the submerged and floating aquatic vegetation zone. Each transect consisted of five quadrat sampling points (i.e., survey plots) per vegetation zone that were evenly spaced and centered between the boundaries of each vegetation zone. Up to 15 quadrat sampling points per transect were sampled if all vegetation zones were present. Vegetation was surveyed in 1m\u003csup\u003e2\u003c/sup\u003e quadrats at each sampling point along each transect for a total of 15–45 quadrats per wetland (depending on number of zones present). A width of 11m was used as a zone width threshold because that was the smallest width to accommodate five 1m\u003csup\u003e2\u003c/sup\u003e quadrats with a 1m distance between quadrats and the zone boundaries. If the width of a vegetation zone was less than 11m, a perpendicular transect was established at the midpoint of the zone along the original transect, and quadrats were placed at 5m intervals along the perpendicular transect in the narrow vegetation zone. The presence and percent cover for each plant species were recorded for each quadrat. Plant species were identified to the species-level where possible, with the exception of certain non-vascular species (e.g., \u003cem\u003eChara\u003c/em\u003e spp. and \u003cem\u003eNitella\u003c/em\u003e spp.). \u003cem\u003eEnvironmental Condition Gradient\u003c/em\u003e The first step in developing the pIBC was to quantify the response of individual plant species to an independently-derived gradient of environmental condition (Howe et al. 2007b; Howe et al. 2023). The environmental condition gradient was defined by a collection of independent environmental variables associated with the wetland sampling points, yielding an objective score ranging from very low environmental quality to very high environmental quality, but not based at all on wetland plant occurrences. We used the Coastal Wetland Monitoring Program’s Water Quality and Land Use (WQLU) Index (formerly called Sum-Rank; Uzarski et al. 2017; Harrison et al. 2020) as our measure of environmental condition because it incorporates components of water quality and land cover that are likely to influence the structure of coastal wetland plant communities at different spatial scales directly and indirectly. This index is also strongly correlated with other common measures of disturbance for Great Lakes coastal wetlands, including AgDev (r = 0.80), a stress index based on the percentage of agricultural land within a watershed, the percentage of urban land use within a watershed, population density, and road density (Host et al. 2019), and the Water Quality Index (r = 0.89), a stress index based solely on water chemistry (Chow-Fraser 2006; equation 8). The WQLU Index was calculated based on ten \u003cem\u003ein situ\u003c/em\u003e water quality variables (water clarity, specific conductance, total nitrogen, nitrate-nitrite-N, ammonium-N, total phosphorus, soluble reactive phosphorus, chlorophyll-a, dissolved oxygen, and pH), eight land-cover variables (proportions of agriculture, development, natural vegetation, and wetlands within 1km and 20km buffers from each wetland), and one principal component (PC1) derived from a principal component analysis of all water quality and land-cover variables (scaled such that higher values indicated better condition) (Uzarski et al. 2005). Values for each variable were rank‐transformed such that the rank order was ordinal to the inferred degree of anthropogenic disturbance (with the greatest rank for each variable indicating the least disturbance). Ranks from all 19 variables were then summed, and each summed WQLU Index value was scaled from 0 to 10, with higher values indicating better environmental condition. Water quality sampling was conducted at three replicate locations within dominant plant growth forms in each wetland where vegetation sampling occurred. Plant growth forms were defined as patches of vegetation dominated by species sharing similar physical structure (e.g., a mixed patch of \u003cem\u003eNuphar variegata\u003c/em\u003e and \u003cem\u003eNymphaea odorata\u003c/em\u003e were considered a single growth form). The most common growth forms sampled were cattail (\u003cem\u003eTypha\u003c/em\u003e spp.), submerged aquatic vegetation (e.g., \u003cem\u003ePotamogeton, Vallisneria, Ceratophyllum\u003c/em\u003e spp.), lily (e.g., \u003cem\u003eNymphaea, Nuphar, Nelumbo\u003c/em\u003e spp.), dense bulrush (\u003cem\u003eSchoenoplectus\u003c/em\u003e spp.), sparse bulrush, wet meadow (e.g., \u003cem\u003eCarex, Juncus\u003c/em\u003e spp.), and common reed (\u003cem\u003ePhragmites australis\u003c/em\u003e). \u003cem\u003eIn situ\u003c/em\u003e water quality variables (dissolved oxygen, pH, and specific conductance) were measured at the mid-depth of the water column using a water quality sonde. Water samples were collected from each sample location and stored for laboratory analysis of total nitrogen, total phosphorus, ammonium-N, nitrate-nitrite- N, soluble reactive phosphorus, and chlorophyll-a following standard analytical methods (APHA 2005). A transparency tube with Secchi disk was used to measure water clarity (see Uzarski et al. 2017 for additional detail). For wetlands where multiple plant growth forms were sampled, water quality parameters were averaged across growth forms, and these average values were used to calculate site-level WQLU Index values for each year a wetland was sampled. Land-cover variables were calculated using the 2010 North American Land Cover 30m dataset from the North American Land Change Monitoring System (NALCMS) (https://www.cec.org/north-american-environmental-atlas/land-cover-2010-landsat-30m/). \u003cem\u003eLinking Species Occurrences to Environmental Condition\u003c/em\u003e For each plant species, we generated best-fit “curves,” referred to as biotic response (BR) functions, relating the probability of occurrence of each species to the WQLU Index environmental condition gradient. Shapes of the curves across the environmental gradient were either monotonically increasing (the upward part of the normal curve), monotonically decreasing (the downward part of the normal curve), or unimodal with complete or truncated tails. To quantify these responses, we applied an algorithm developed by Howe et al. (2023) to select the best-fit parameters of a normal (bell-shaped) curve. The curve was defined using the R function \u003cem\u003ednorm\u003c/em\u003e (R Core Team 2023), and the parameters were estimated via the \u003cem\u003enlminb\u003c/em\u003e function (Gay, 1990). We divided our dataset into lake-specific training datasets (to develop/calibrate the index) and testing datasets (to validate the index on a dataset that was separate from the dataset used for development/calibration). Our training datasets consisted of 75% of sites that were sampled at least two times over the study period; if these sites were sampled more than twice, then two sample years were randomly selected to ensure that sites that were sampled more frequently were not overrepresented (36 sites [72 sampling events] for Lake Superior, 42 [84] for Lake Michigan, 104 [208] for Lake Huron, 39 [78] for Lake Erie, 76 [152] for Lake Ontario). Our testing datasets consisted of the remaining 25% of sites that were sampled at least two times over the study period, as well as any sites that were sampled only once during the study period (37 sites [59 sampling events] for Lake Superior, 45 [74] for Lake Michigan, 90 [159] for Lake Huron, 27 [48] for Lake Erie, 67 [107] for Lake Ontario). The training datasets were used to generate lake-specific BR functions for all native species that occurred at least five times in those datasets (124 species for Lake Superior, 141 for Lake Michigan, 251 for Lake Huron, 59 for Lake Erie, 124 for Lake Ontario; Fig. 1, Online Resource 1). \u003cem\u003ePlant Index of Biotic Condition (pIBC)\u003c/em\u003e Site-level pIBC values were computed by aggregating species “weights” (\u003cem\u003ew\u003csub\u003ei\u003c/sub\u003e\u003c/em\u003e), which were calculated using parameters derived from BR functions for species encountered during standardized field surveys (see the files in this dataset, https://doi.org/10.5061/dryad.0cfxpnwg2, for weights for all species for each lake and a pIBC calculator). A weight for each native species was calculated as the product of two parameters: 1) the species optimum, \u003cem\u003e\u003csub\u003ei\u003c/sub\u003e\u003c/em\u003e (the value of WQLU at which probability of occurrence is highest; the peak of the BR function), and 2) the difference between \u003cem\u003ep(\u003c/em\u003e\u003cem\u003e\u003csub\u003ei\u003c/sub\u003e)\u003csub\u003ei\u003c/sub\u003e\u003c/em\u003e (the predicted probability of occurrence at the optimum) and \u003cem\u003ep(0)\u003csub\u003ei\u003c/sub\u003e\u003c/em\u003e (the predicted probability of occurrence under the most degraded conditions [WQLU = 0]). The first parameter (\u003cem\u003e\u003csub\u003ei\u003c/sub\u003e\u003c/em\u003e) reflects the sensitivity of the species to the environmental condition gradient; in general, the more sensitive the species, the larger the optimum value. The second parameter [\u003cem\u003ep(\u003c/em\u003e\u003cem\u003e\u003csub\u003ei\u003c/sub\u003e)\u003csub\u003eI\u003c/sub\u003e\u003c/em\u003e \u003cem\u003e– p(0)\u003csub\u003ei\u003c/sub\u003e\u003c/em\u003e] reflects the responsiveness of the species to change in the environmental condition gradient; in general, the more responsive the species, the larger the difference between the maximum probability of occurrence and the probability of occurrence when environmental condition is most degraded. Hypothetically, a species with a maximum weight (\u003cem\u003ew\u003csub\u003ei\u003c/sub\u003e\u003c/em\u003e = 10) would have peak probability of occurrence when WQLU = 10 (\u003cem\u003e\u003csub\u003ei\u003c/sub\u003e\u003c/em\u003e = 10) and would occur at 100% of sites when WQLU = 10 and 0% of sites when WQLU = 0 [\u003cem\u003ep(\u003c/em\u003e\u003cem\u003e\u003csub\u003ei\u003c/sub\u003e)\u003csub\u003ei\u003c/sub\u003e\u003c/em\u003e \u003cem\u003e– p(0)\u003csub\u003ei\u003c/sub\u003e\u003c/em\u003e  = 1]. All non-native species were assigned a weight of zero, as their presence inherently reflects anthropogenic introduction and departure from reference conditions, which was consistent with the conceptual basis of the index. We considered retaining empirically derived weights for non-native species. Although preliminary calculations of weights for non-natives were generally low (median = 0.50), a small subset received comparatively higher weights, reflecting that some non-native species occur more frequently in sites otherwise characterized by better conditions. However, these cases likely reflect factors such as introduction history and dispersal pathways that are not fully captured by the WQLU environmental condition gradient. In addition, retaining empirically derived weights for such taxa could give the misleading impression that their presence is desirable. Importantly, preliminary testing (not shown here) indicated that this decision to assign non-native species a weight of zero did not meaningfully affect the performance of the pIBC. The pIBC value for a given wetland () was calculated as the mean weight of all plant species observed multiplied by the square-root of the number of species with weights, following the same general structure of the Floristic Quality Index (mean C multiplied by √N; Swink and Wilhelm 1994):   where  is the mean weight and  is the number of species with weights for wetland \u003cem\u003ei\u003c/em\u003e within a given lake. We note that this computation differs from Howe et al. (2023), who calculated their bird IBC as the sum of weights for all species present at a site. Calculating the pIBC in this modified manner integrates species richness but avoids scenarios where sites with numerous low-weight species outscore sites with fewer high-weight species. This was particularly important for the pIBC because of the substantial number of species included in the different versions of the index (59 to 251 depending on the lake, compared to 14 species used for the bird IBC). \u003cem\u003eIndex Validation\u003c/em\u003e To validate each lake-specific version of the pIBC, we fit linear mixed-effects models (Gaussian family) with the WQLU Index as the response variable, pIBC, dataset (whether the site/sampling event belonged to the training or testing dataset), an interaction between both terms as fixed effects, and site and year as random effects. Each lake-specific pIBC was considered validated if there was (1) a non-statistically significant interaction between pIBC and dataset (indicating that the relationship between pIBC and WQLU was not significantly different between training and testing datasets), and (2) a statistically significant positive relationship between pIBC and WQLU. Although the models could have been fitted with WQLU Index as the predictor and pIBC as the response, we modelled WQLU as the response because our objective was to validate the pIBC rather than to model the ecological process (that disturbance shapes plant communities). In this validation framework, we assessed whether the pIBC, derived solely from plant community data, recovers known gradients of anthropogenic disturbance measured independently through water quality and land use. The inclusion of the dataset predictor and its interaction with pIBC was intended to ensure that validation was independent of the weight calculation process and to evaluate whether the pIBC performed consistently on previously unseen sites. \u003cem\u003eComparison to Other Metrics\u003c/em\u003e We compared the performance of each lake-specific version of the pIBC to seven other vegetation metrics by evaluating their relationships with the WQLU Index (Uzarski et al. 2017; Harrison et al. 2020) within each lake. The vegetation metrics were: (1) mean coefficient of conservatism (mean C), (2) floristic quality index (FQI) (Swink and Wilhelm 1994), (3) cover-weighted mean C, (4) cover-weighted FQI (Bourdaghs et al. 2006), (5) a recently developed multi-metric index of biotic integrity (IBI) for Great Lakes coastal wetlands (Dybiec et al. 2020), (6) native species richness, and (7) invasive species cover. All metrics were calculated by aggregating survey data from all quadrats sampled at a site, except for invasive species cover, which was calculated as a quadrat-level average because preliminary analyses showed a stronger relationship between quadrat-level average cover and WQLU than site-level cumulative cover. For the FQA metrics, non-native species were included and assigned C values of zero. We elected to include non-native species in these metrics because many are widespread and often among the dominant taxa in many Great Lakes coastal wetlands (Trebitz and Taylor 2007). This approach was also consistent with our pIBC methodology (assigning weights of zero to non-native species). C values assigned to native species were based on the Floristic Quality Assessment System for Southern Ontario (for lakes Erie and Ontario) (Oldham et al. 1995) and the Michigan Floristic Quality Assessment Database (for lakes Superior, Michigan, and Huron) (Reznicek et al. 2014). Using the full dataset (training, testing, and sampling events originally omitted from the training dataset), we fit a series of linear mixed-effects models for each lake (Gaussian family) with the WQLU Index as the response variable, one of the eight vegetation metrics as a fixed effect, and site and year as random effects. The full dataset was used to maximize statistical power, which was justified because, to validate the index, we demonstrated that the relationship between each lake-specific version of the pIBC and WQLU was not significantly different between training and testing datasets. We then assessed the relative performance of the eight candidate models (representing each vegetation metric) for each lake by comparing Akaike Information Criterion corrected for small sample size (ΔAICc) and marginal R\u003csup\u003e2\u003c/sup\u003e (R\u003csup\u003e2\u003c/sup\u003e\u003csub\u003em\u003c/sub\u003e; Nakagawa and Schielzeth 2013) values. This mixed-effects modelling framework provided an advantage over simpler cross-sectional approaches that use only one observation per site. By incorporating site-level sampling events in different years and including site and year as random effects, our models accounted for spatial and temporal structure in the data. This allowed us to evaluate each vegetation metric’s ability to track variation in disturbance across sites and within sites over time, which is fundamental for long-term ecological monitoring. \u003cem\u003eWater Levels and Hydrogeomorphic Types\u003c/em\u003e We also tested whether each lake-specific pIBC was suitable for use under different water-level conditions, and for wetlands of different hydrogeomorphic types. To achieve this, for each lake version of the pIBC, we fit a linear mixed-effects model (Gaussian family) with the WQLU Index as the response variable, two interaction terms (pIBC × annual water level [high or low] and pIBC × hydrogeomorphic type [lacustrine, riverine, or barrier-protected; Albert et al. 2005]) as fixed effects, and site and year as random effects. For all lakes except Lake Ontario, annual water level was classified as either high or low depending on whether the mean water level during the growing season (May to October) for a given year was above or below the median from 2011 to 2022 (water-level data were from the Great Lakes Coordinating Committee; https://www.greatlakescc.org/en/coordinating-committee-products-and-datasets). Because water levels on Lake Ontario generally experience little interannual variation as a result of outflow regulation, only the extreme water levels of 2017 and 2019 were considered high-water-level years (Smith et al. 2021). For Lake Superior, “low” annual water levels ranged from 183.19 to 183.59m relative to the International Great Lakes Datum 1985 (IGLD) (in five of the twelve years of the study) and “high” annual water levels ranged from 183.64 to 183.85m IGLD (seven years); for Lake Michigan and Lake Huron (which are hydrologically connected), “low” annual water levels ranged from 175.96 to 176.74m IGLD (six years) and “high” annual water levels ranged from 176.80 to 177.38m IGLD (six years); for Lake Erie, “low” annual water levels ranged from 174.07 to 174.51m IGLD (six years) and “high” annual water levels ranged from 174.55 to 174.99m IGLD (six years); and for Lake Ontario, “low” annual water levels ranged from 74.68 to 75.07m IGLD (10 years) and “high” annual water levels ranged from 75.45 to 75.54m IGLD (two years). For the models, non-statistically significant interactions indicated that the relationship between pIBC and the WQLU Index was similar between high- and low-water-level years, and among hydrogeomorphic wetland types. For this analysis, we also considered using raw annual water-level values and a three-category classification (high, moderate, or low), but these alternative approaches did not meaningfully change the interpretation of the results and are not presented here. All analyses were completed in R (Version 4.3.1; R Core Team 2023). Mixed models were fit using the \u003cem\u003elme4\u003c/em\u003e package in R (Bates et al. 2015), \u003cem\u003eP\u003c/em\u003e-values were obtained using the \u003cem\u003elmerTest\u003c/em\u003e package (Kuznetsova et al. 2017), and models were compared using the \u003cem\u003eperformance\u003c/em\u003e package (Lüdecke et al. 2021). Statistical assumptions of linear mixed effects models were assessed using standard diagnostic plots and were adequately met in all cases."},{"descriptionType":"TechnicalInfo","description":"# Data from: Beyond floristic quality: A data-driven approach to assessing\n coastal wetland condition using plants Dataset DOI:\n [10.5061/dryad.0cfxpnwg2](https://doi.org/10.5061/dryad.0cfxpnwg2) ##\n Description of the data and file structure Using a massive dataset (41,314\n survey plots from 563 wetlands across 12 years), we developed and tested a\n new plant-based Index of Biotic Condition (pIBC) for coastal wetlands of\n each of the five Laurentian Great Lakes. See \"Methods\" for\n details. ### Files and variables #### File: dataset.csv\n **Description:** Dataset generated and analyzed for the study. Each row in\n the table is a wetland sampling event (a site sampled in a given year).\n ##### Variables * Site: Wetland site number. * Lake: Great Lake where the\n wetland site is located. * Year: Year sampled. * Dataset: Whether the\n sampling event was part of the lake-specific training datasets (to\n develop/calibrate the index) or testing datasets (to validate the index on\n a dataset that was separate from the dataset used for\n development/calibration). NOTE: “NA,” indicating “Not Applicable,” was\n applied to sampling events that were not included in training or testing\n datasets. * pIBC_LE: Lake Erie version of the pIBC. NOTE: “NA,” indicating\n “Not Applicable,” was applied to sites on lakes other than Lake Erie. *\n pIBC_LH: Lake Huron version of the pIBC. NOTE: “NA,” indicating “Not\n Applicable,” was applied to sites on lakes other than Lake Huron. *\n pIBC_LM: Lake Michigan version of the pIBC. NOTE: “NA,” indicating “Not\n Applicable,” was applied to sites on lakes other than Lake Michigan. *\n pIBC_LO: Lake Ontario version of the pIBC. NOTE: “NA,” indicating “Not\n Applicable,” was applied to sites on lakes other than Lake Ontario. *\n pIBC_LS: Lake Superior version of the pIBC. NOTE: “NA,” indicating “Not\n Applicable,” was applied to sites on lakes other than Lake Superior. *\n FQA_MeanC: Mean Coefficient of Conservatism. * FQA_FQI: Floristic Quality\n Index. * FQA_CoverWeightedMeanC: Cover-weighted mean Coefficient of\n Conservatism. * FQA_CoverWeightedFQI: Cover-weighted Floristic Quality\n Index. * IBI: Multi-metric index of biotic integrity (IBI) for Great Lakes\n coastal wetlands. * InvCovPerQuad: Average invasive species cover per\n quadrat. NOTE: “NR,” indicating “Not Reported,” was applied to five\n sampling events due to a methodological deviation that limited\n comparability with other sampling events. * NativeRich: Native species\n richness. * WQLU: The Coastal Wetland Monitoring Program’s Water Quality\n and Land Use (WQLU) Index. * WaterLevelCategory: Annual water level\n classified as either high or low depending on whether the mean water level\n during the growing season (May to October) for a given year was above or\n below the median from 2011 to 2022. * HydrogeomorphicType:\n Wetland hydrogeomorphic type (lacustrine, riverine, or barrier-protected).\n #### File: weights_LS.csv **Description:** List of species and the\n associated weights used to calculate Lake Superior-specific pIBC values.\n The two parameters derived from biotic response (BR) functions that were\n used to calculate the weights are also provided: (1) the optimum of the BR\n function (also referred to as the mean), and (2) the difference between\n the predicted probability of occurrence at the optimum (mean) of the BR\n function and the predicted probability of occurrence when WQLU = 0. #####\n Variables * species: Scientific name. * nativeness: Whether a species was\n native or non-native. * mean: The optimum of the BR function (also\n referred to as the mean). NOTE: “NA,” indicating “Not Applicable,” was\n applied to non-native species, which were assigned a weight of zero rather\n than an empirically derived weight. * p(mean)-p(0): The difference between\n the predicted probability of occurrence at the optimum (mean) of the BR\n function and the predicted probability of occurrence when WQLU = 0. NOTE:\n “NA,” indicating “Not Applicable,” was applied to non-native species,\n which were assigned a weight of zero rather than an empirically derived\n weight. * weight: Species weight. #### File: weights_LM.csv\n **Description:** List of species and the associated weights used to\n calculate Lake Michigan-specific pIBC values. The two parameters derived\n from biotic response (BR) functions that were used to calculate the\n weights are also provided: (1) the optimum of the BR function (also\n referred to as the mean), and (2) the difference between the predicted\n probability of occurrence at the optimum (mean) of the BR function and the\n predicted probability of occurrence when WQLU = 0. ##### Variables *\n species: Scientific name. * nativeness: Whether a species was native or\n non-native. * mean: The optimum of the BR function (also referred to as\n the mean). NOTE: “NA,” indicating “Not Applicable,” was applied to\n non-native species, which were assigned a weight of zero rather than an\n empirically derived weight. * p(mean)-p(0): The difference between the\n predicted probability of occurrence at the optimum (mean) of the BR\n function and the predicted probability of occurrence when WQLU = 0. NOTE:\n “NA,” indicating “Not Applicable,” was applied to non-native species,\n which were assigned a weight of zero rather than an empirically derived\n weight. * weight: Species weight. #### File: weights_LH.csv\n **Description:** List of species and the associated weights used to\n calculate Lake Huron-specific pIBC values. The two parameters derived from\n biotic response (BR) functions that were used to calculate the weights are\n also provided: (1) the optimum of the BR function (also referred to as the\n mean), and (2) the difference between the predicted probability of\n occurrence at the optimum (mean) of the BR function and the predicted\n probability of occurrence when WQLU = 0. ##### Variables * species:\n Scientific name. * nativeness: Whether a species was native or non-native.\n * mean: The optimum of the BR function (also referred to as the\n mean). NOTE: “NA,” indicating “Not Applicable,” was applied to non-native\n species, which were assigned a weight of zero rather than an empirically\n derived weight. * p(mean)-p(0): The difference between the predicted\n probability of occurrence at the optimum (mean) of the BR function and the\n predicted probability of occurrence when WQLU = 0. NOTE: “NA,” indicating\n “Not Applicable,” was applied to non-native species, which were assigned a\n weight of zero rather than an empirically derived weight. * weight:\n Species weight. #### File: weights_LE.csv **Description:** List of species\n and the associated weights used to calculate Lake Erie-specific pIBC\n values. The two parameters derived from biotic response (BR) functions\n that were used to calculate the weights are also provided: (1) the optimum\n of the BR function (also referred to as the mean), and (2) the difference\n between the predicted probability of occurrence at the optimum (mean) of\n the BR function and the predicted probability of occurrence when WQLU = 0.\n ##### Variables * species: Scientific name. * nativeness: Whether a\n species was native or non-native. * mean: The optimum of the BR function\n (also referred to as the mean). NOTE: “NA,” indicating “Not Applicable,”\n was applied to non-native species, which were assigned a weight of zero\n rather than an empirically derived weight. * p(mean)-p(0): The difference\n between the predicted probability of occurrence at the optimum (mean) of\n the BR function and the predicted probability of occurrence when WQLU =\n 0. NOTE: “NA,” indicating “Not Applicable,” was applied to non-native\n species, which were assigned a weight of zero rather than an empirically\n derived weight. * weight: Species weight. #### File: weights_LO.csv\n **Description:** List of species and the associated weights used to\n calculate Lake Ontario-specific pIBC values. The two parameters derived\n from biotic response (BR) functions that were used to calculate the\n weights are also provided: (1) the optimum of the BR function (also\n referred to as the mean), and (2) the difference between the predicted\n probability of occurrence at the optimum (mean) of the BR function and the\n predicted probability of occurrence when WQLU = 0. ##### Variables *\n species: Scientific name. * nativeness: Whether a species was native or\n non-native. * mean: The optimum of the BR function (also referred to as\n the mean). NOTE: “NA,” indicating “Not Applicable,” was applied to\n non-native species, which were assigned a weight of zero rather than an\n empirically derived weight. * p(mean)-p(0): The difference between the\n predicted probability of occurrence at the optimum (mean) of the BR\n function and the predicted probability of occurrence when WQLU = 0. NOTE:\n “NA,” indicating “Not Applicable,” was applied to non-native species,\n which were assigned a weight of zero rather than an empirically derived\n weight. * weight: Species weight. #### File: pIBCcalculator.xlsx\n **Description:** Each sheet in this workbook contains a list of species\n and the associated weights used to calculate lake-specific pIBC values.\n The two parameters derived from biotic response (BR) functions that were\n used to calculate the weights are also provided: (1) the optimum of the BR\n function (also referred to as the mean), and (2) the difference between\n the predicted probability of occurrence at the optimum (mean) of the BR\n function and the predicted probability of occurrence when WQLU = 0 (NOTE:\n “NA,” indicating “Not Applicable,” was applied to non-native species for\n these two parameters because they were assigned a weight of zero rather\n than an empirically derived weight). The pIBC value for a given wetland is\n calculated as the mean weight of all plant species observed during\n standardized field surveys multiplied by the square-root of the number of\n species with weights. **To generate a site-level pIBC score using these\n sheets, indicate whether each species was present (Y) or absent (N) during\n standardized field surveys in column G of the sheet for a given lake. The\n pIBC score is shown in cell K2.** ## Code/software Microsoft Excel is\n required to use the pIBC calculator. The calculator is implemented as an\n Excel workbook containing embedded formulas that compute pIBC values after\n users input species occurrence data. Underlying data files are provided in\n both Excel and CSV format for accessibility. If users require assistance\n with running the calculator or understanding the workflow, they may\n contact the lead author. ## Access information Other publicly accessible\n locations of the data: * Raw plant survey data and Water Quality and Land\n Use Index data are available upon request from the Great Lakes Coastal\n Wetlands Monitoring\n Program: [https://www.greatlakeswetlands.org/Account/Request.vbhtml](https://www.greatlakeswetlands.org/Account/Request.vbhtml). * Water-level data were from the Great Lakes Coordinating Committee: [https://www.greatlakescc.org/en/coordinating-committee-products-and-datasets.](https://www.greatlakescc.org/en/coordinating-committee-products-and-datasets/) * Hydrogeomorphic classifications are based on Albert, D. A., D. A. Wilcox, J. W. Ingram, and T. A. Thompson. 2005. Hydrogeomorphic Classification for Great Lakes Coastal Wetlands. Journal of Great Lakes Research 31:129–146."}],"geoLocations":[],"fundingReferences":[{"funderIdentifierType":"ROR","funderName":"Environmental Protection Agency","funderIdentifier":"https://ror.org/03tns0030","awardNumber":"GL-00E00612-0"},{"funderIdentifierType":"ROR","funderName":"Environmental Protection Agency","funderIdentifier":"https://ror.org/03tns0030","awardNumber":"00E01567"},{"funderIdentifierType":"ROR","funderName":"Environmental Protection Agency","funderIdentifier":"https://ror.org/03tns0030","awardNumber":"00E02956"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.0cfxpnwg2","contentUrl":null,"metadataVersion":0,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-20T21:30:33Z","registered":"2026-08-20T21:30:34Z","published":null,"updated":"2026-08-20T21:30:34Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.pk0p2nh55","type":"dois","attributes":{"doi":"10.5061/dryad.pk0p2nh55","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["University of Alicante"],"name":"Mellone, Ugo","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0001-9665-505X"}]},{"nameType":"Personal","affiliation":["University of Veterinary Sciences Brno"],"name":"Literak, Ivan","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Milan"],"name":"Berlusconi, Alessandro","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Pavia"],"name":"Bogliani, Giuseppe","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Ornis Italica"],"name":"Dell'Omo, Giacomo","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["National Research Council"],"name":"Morganti, Michelangelo","nameIdentifiers":[]}],"titles":[{"title":"Migratory divides highlight intraspecific flexibility in the stop-over tactic of black kites"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Ornithology","subjectScheme":"PLOS Subject Area Thesaurus"},{"subject":"migration"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Raptors","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Natural sciences","subjectScheme":"fos"},{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Animal and dairy science","subjectScheme":"fos"}],"contributors":[],"dates":[{"date":"2026-07-28T15:24:30Z","dateType":"Created"},{"date":"2026-07-28T15:24:31Z","dateType":"Submitted"},{"date":"2026-08-20T00:00:00Z","dateType":"Issued"},{"date":"2026-08-20T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[{"relationType":"IsCitedBy","relatedIdentifier":"10.1002/jav.03617","relatedIdentifierType":"DOI"}],"relatedItems":[],"sizes":["8725 bytes"],"formats":[],"version":"4","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"This dataset contains the data used to investigate variation in migratory\n performance and stopover use among juvenile black kites (Milvus migrans)\n originating from different European breeding populations and migrating\n along three Mediterranean flyways. It provides information on individual\n birds, including breeding population, natal area, assigned migratory\n flyway, and migration performance metrics derived from GPS telemetry,\n including migration duration, travel speed, stopover frequency and\n duration, and the use of different stopover habitats. These data were used\n to assess how migratory origin and route choice are associated with\n variation in migration strategies and habitat use."},{"descriptionType":"TechnicalInfo","description":"# Migratory divides highlight intraspecific flexibility in the stop-over\n tactic of black kites Dataset DOI:\n [10.5061/dryad.pk0p2nh55](https://doi.org/10.5061/dryad.pk0p2nh55) ##\n Description of the data and file structure ### Files and variables Note:\n Decimal separators are commas in this dataset. #### File:\n TABLE_3RD_MODEL.csv ##### Variables * Individual: tag number of each bird\n * Type: if the stopover was in a \"landfill\" or in\n \"other\" habitat * StopoverDuration: days spent in each event *\n Population: hatching area (4 levels) * Route: used flyway (3 levels) *\n Region: geographical area of the stopover event (2 levels) * DepartureDay:\n Julian day #### File: TABLE_1_AND_2ND_MODEL.csv ##### Variables *\n Individual: tag number of each bird * Population: hatching area (4 levels)\n * Route: used flyway (3 levels) * Year: year of the migration event *\n DepartureDay: Julian day * ArrivalDay: Julian day * Duration: difference\n in days between arrival and departure * Stopdays: overall number of\n stopover days * DaysLandfills: stopover days spent in landfills *\n N-Landfills: number of used landfills * Distance: overall distance of the\n journey in km (summing daily segments) NA/blank cells mean those data are\n not available because the arrival and/or the departure of a given\n individual was not recorded by the GPS transmitter."}],"geoLocations":[],"fundingReferences":[{"funderIdentifierType":"ROR","funderName":"Institute for Complex Systems","funderIdentifier":"https://ror.org/05rcgef49"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.pk0p2nh55","contentUrl":null,"metadataVersion":0,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":1,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-20T20:02:30Z","registered":"2026-08-20T20:02:31Z","published":null,"updated":"2026-08-20T20:02:31Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.8sf7m0d53","type":"dois","attributes":{"doi":"10.5061/dryad.8sf7m0d53","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["University of Missouri–St. Louis","Saint Louis Zoo"],"name":"Tobler, Michael","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0002-0326-0890"}]},{"nameType":"Personal","affiliation":["Max Planck Institute of Animal Behavior"],"name":"Greenway, Ryan","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of California, Santa Cruz"],"name":"Kelley, Joanna","nameIdentifiers":[]}],"titles":[{"title":"Ecology drives the degree of convergence in the gene expression of extremophile fishes"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"subject":"Poeciliidae"},{"subject":"hydrogen sulfide"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Gene expression","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Convergent evolution","subjectScheme":"PLOS Subject Area Thesaurus"}],"contributors":[],"dates":[{"date":"2026-08-05T22:43:19Z","dateType":"Created"},{"date":"2026-08-05T22:43:21Z","dateType":"Submitted"},{"date":"2026-08-07T00:00:00Z","dateType":"Issued"},{"date":"2026-08-07T00:00:00Z","dateType":"Available"},{"date":"2026-08-20T00:00:00Z","dateType":"Updated"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[{"relationType":"IsSupplementedBy","relatedIdentifier":"https://www.ncbi.nlm.nih.gov/bioproject/PRJNA473350/","relatedIdentifierType":"URL"},{"relationType":"IsSupplementedBy","relatedIdentifier":"https://www.ncbi.nlm.nih.gov/bioproject/PRJNA608180/","relatedIdentifierType":"URL"}],"relatedItems":[],"sizes":["28953518 bytes"],"formats":[],"version":"7","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"This repository contains the data and code used to analyze patterns of\n convergent gene expression in hydrogen sulfide-adapted poeciliid fishes.\n The dataset includes: (1) a gene expression matrix of fragments per\n kilobase of transcript per million mapped reads (FPKM) for\n 31,805  genes across 118 gill transcriptomes representing 20 fish\n lineages from sulfidic and nonsulfidic habitats; (2) environmental\n measurements for sulfide spring localities, including water chemistry\n variables (temperature, pH, conductivity, dissolved oxygen, and hydrogen\n sulfide concentration); (3) the maximum-likelihood phylogenetic tree used\n for comparative analyses; and (4) R scripts used to reproduce the\n analyses, including environmental principal component analyses, expression\n variance and evolution (EVE) models, phylogenetic mixed models, and figure\n generation. The data enable replication of all statistical analyses. Raw\n RNA-sequencing reads are available separately through the NCBI Sequence\n Read Archive under BioProject accessions PRJNA473350 and PRJNA608180. The\n data contain no personally identifiable information or human subjects and\n have no known ethical or legal restrictions beyond the terms of this\n repository."},{"descriptionType":"TechnicalInfo","description":"# Ecology drives the degree of convergence in the gene expression of\n extremophile fishes Dataset DOI:\n [10.5061/dryad.8sf7m0d53](https://doi.org/10.5061/dryad.8sf7m0d53) ##\n Description of the data and file structure ### Files and variables ####\n File: PredictorsConvergenceFinalAnalysis.Rmd **Description:** Markdown\n file contains all R code to reproduce analyses presented in the paper\n based on the input files below. #### File: recodedTreeNamed.tre\n **Description:** Best-scoring ML tree depicting the phylogenetic\n relationships of taxa included in this study. Tip labels correspond to\n lineages as listed in ESM Table S1. #### File: FPKM_AllGenes.csv\n **Description:** Gene expression levels in FPKM (Fragments Per Kilobase of\n transcript per Million mapped reads) for 31,804 genes in 118 samples\n analyzed for this study. Column labels correspond to lineages as listed in\n ESM Table S1 in the associated manuscript. ##### Variables * gene_id: Gene\n ID from mapping against the P. mexicana reference genome. * LsulS_1: Limia\n sulphurophila, sulfidic, sample 1 * LsulS_2: Limia sulphurophila,\n sulfidic, sample 2 * LsulS_3: Limia sulphurophila, sulfidic, sample 3 *\n LsulS_4: Limia sulphurophila, sulfidic, sample 4 * LsulS_5: Limia\n sulphurophila, sulfidic, sample 5 * LsulS_6: Limia sulphurophila,\n sulfidic, sample 5 * LperNS_1: Limia perugiae, nonsulfidic, sample 1 *\n LperNS_2: Limia perugiae, nonsulfidic, sample 2 * LperNS_3: Limia\n perugiae, nonsulfidic, sample 3 * LperNS_4: Limia perugiae, nonsulfidic,\n sample 4 * LperNS_5: Limia perugiae, nonsulfidic, sample 5 *\n LperNS_6: Limia perugiae, nonsulfidic, sample 6 * GholS_1: Gambusia\n holbrooki, sulfidic, sample 1 * GholS_2: Gambusia holbrooki, sulfidic,\n sample 2 * GholS_3: Gambusia holbrooki, sulfidic, sample 3 *\n GholS_4: Gambusia holbrooki, sulfidic, sample 4 * GholS_5: Gambusia\n holbrooki, sulfidic, sample 5 * GholS_6: Gambusia holbrooki, sulfidic,\n sample 6 * PlatS_1: Poecilia latipinna, sulfidic, sample 1 *\n PlatS_2: Poecilia latipinna, sulfidic, sample 2 * PlatS_3: Poecilia\n latipinna, sulfidic, sample 3 * PlatS_4: Poecilia latipinna, sulfidic,\n sample 4 * PlatS_5: Poecilia latipinna, sulfidic, sample 5 *\n PlatS_6: Poecilia latipinna, sulfidic, sample 6 * PlatNS_1: Poecilia\n latipinna, nonsulfidic, sample 1 * PlatNS_2: Poecilia latipinna,\n nonsulfidic, sample 2 * PlatNS_3: Poecilia latipinna, nonsulfidic, sample\n 3 * PlatNS_4: Poecilia latipinna, nonsulfidic, sample 4 *\n PlatNS_5: Poecilia latipinna, nonsulfidic, sample 5 * PlatNS_6: Poecilia\n latipinna, nonsulfidic, sample 6 * GholNS_1: Gambusia holbrooki,\n nonsulfidic, sample 1 * GholNS_2: Gambusia holbrooki, nonsulfidic, sample\n 2 * GholNS_3: Gambusia holbrooki, nonsulfidic, sample 3 *\n GholNS_4: Gambusia holbrooki, nonsulfidic, sample 4 * GholNS_5: Gambusia\n holbrooki, nonsulfidic, sample 5 * GholNS_6: Gambusia holbrooki,\n nonsulfidic, sample 6 * PmexPichNS_1: Poecilia mexicana Pichucalco,\n nonsulfidic, sample 1 * PmexPichNS_2: Poecilia mexicana Pichucalco,\n nonsulfidic, sample 2 * PmexPichNS_3: Poecilia mexicana Pichucalco,\n nonsulfidic, sample 3 * PmexPichNS_4: Poecilia mexicana Pichucalco,\n nonsulfidic, sample 4 * PmexPichNS_5: Poecilia mexicana Pichucalco,\n nonsulfidic, sample 5 * PmexPichNS_6: Poecilia mexicana Pichucalco,\n nonsulfidic, sample 6 * XhelS_1: Xiphophorus hellerii, sulfidic, sample 1\n * XhelS_2: Xiphophorus hellerii, sulfidic, sample 2 * XhelS_3: Xiphophorus\n hellerii, sulfidic, sample 3 * XhelS_4: Xiphophorus hellerii, sulfidic,\n sample 4 * XhelS_5: Xiphophorus hellerii, sulfidic, sample 5 *\n XhelS_6: Xiphophorus hellerii, sulfidic, sample 6 * GeurS_1: Gambusia\n eurystoma, sulfidic, sample 1 * GeurS_2: Gambusia eurystoma, sulfidic,\n sample 2 * GeurS_3: Gambusia eurystoma, sulfidic, sample 3 *\n GeurS_4: Gambusia eurystoma, sulfidic, sample 4 * GeurS_5: Gambusia\n eurystoma, sulfidic, sample 5 * GeurS_6: Gambusia eurystoma, sulfidic,\n sample 6 * PbimNS_1: Pseudoxiphophorus, nonsulfidic, sample 1 *\n PbimNS_2: Pseudoxiphophorus, nonsulfidic, sample 2 *\n PbimNS_3: Pseudoxiphophorus, nonsulfidic, sample 3 *\n PbimNS_4: Pseudoxiphophorus, nonsulfidic, sample 4 *\n PbimNS_5: Pseudoxiphophorus, nonsulfidic, sample 5 *\n PbimNS_6: Pseudoxiphophorus, nonsulfidic, sample 6 *\n PbimS_1: Pseudoxiphophorus, sulfidic, sample 1 *\n PbimS_2: Pseudoxiphophorus, sulfidic, sample 2 *\n PbimS_3: Pseudoxiphophorus, sulfidic, sample 3 *\n PbimS_4: Pseudoxiphophorus, sulfidic, sample 4 *\n PbimS_5: Pseudoxiphophorus, sulfidic, sample 5 *\n PbimS_6: Pseudoxiphophorus, sulfidic, sample 6 * XhelNS_1: Xiphophorus\n hellerii, nonsulfidic, sample 1 * XhelNS_2: Xiphophorus hellerii,\n nonsulfidic, sample 2 * XhelNS_3: Xiphophorus hellerii, nonsulfidic,\n sample 3 * XhelNS_4: Xiphophorus hellerii, nonsulfidic, sample 4 *\n XhelNS_5: Xiphophorus hellerii, nonsulfidic, sample 5 *\n XhelNS_6: Xiphophorus hellerii, nonsulfidic, sample 6 * GsexNS_1: Gambusia\n sexradiata, nonsulfidic, sample 1 * GsexNS_2: Gambusia sexradiata,\n nonsulfidic, sample 2 * GsexNS_3: Gambusia sexradiata, nonsulfidic, sample\n 3 * GsexNS_4: Gambusia sexradiata, nonsulfidic, sample 4 *\n GsexNS_5: Gambusia sexradiata, nonsulfidic, sample 5 * GsexNS_6: Gambusia\n sexradiata, nonsulfidic, sample 6 * GsexS_1: Gambusia sexradiata,\n sulfidic, sample 1 * GsexS_2: Gambusia sexradiata, sulfidic, sample 2 *\n GsexS_3: Gambusia sexradiata, sulfidic, sample 3 * GsexS_4: Gambusia\n sexradiata, sulfidic, sample 4 * GsexS_5: Gambusia sexradiata, sulfidic,\n sample 5 * GsexS_6: Gambusia sexradiata, sulfidic, sample 6 *\n PmexPichS_1: Poecilia mexicana Pichucalco, sulfidic, sample 1 *\n PmexPichS_2: Poecilia mexicana Pichucalco, sulfidic, sample 2 *\n PmexPichS_3: Poecilia mexicana Pichucalco, sulfidic, sample 3 *\n PmexPichS_4: Poecilia mexicana Pichucalco, sulfidic, sample 4 *\n PmexPichS_5: Poecilia mexicana Pichucalco, sulfidic, sample 5 *\n PmexPichS_6: Poecilia mexicana Pichucalco, sulfidic, sample 6 *\n PmexPuyS_1: Poecilia mexicana Puyacatengo, sulfidic, sample 1 *\n PmexPuyS_2: Poecilia mexicana Puyacatengo, sulfidic, sample 2 *\n PmexPuyS_3: Poecilia mexicana Puyacatengo, sulfidic, sample 3 *\n PmexPuyS_4: Poecilia mexicana Puyacatengo, sulfidic, sample 4 *\n PmexPuyS_5: Poecilia mexicana Puyacatengo, sulfidic, sample 5 *\n PmexPuyNS_1: Poecilia mexicana Puyacatengo, nonsulfidic, sample 1 *\n PmexPuyNS_2: Poecilia mexicana Puyacatengo, nonsulfidic, sample 2 *\n PmexPuyNS_3: Poecilia mexicana Puyacatengo, nonsulfidic, sample 3 *\n PmexPuyNS_4: Poecilia mexicana Puyacatengo, nonsulfidic, sample 4 *\n PmexPuyNS_5: Poecilia mexicana Puyacatengo, nonsulfidic, sample 5 *\n PmexPuyNS_6: Poecilia mexicana Puyacatengo, nonsulfidic, sample 6 *\n PmexTacNS_1: Poecilia mexicana Tacotalpa, nonsulfidic, sample 1 *\n PmexTacNS_2: Poecilia mexicana Tacotalpa, nonsulfidic, sample 2 *\n PmexTacNS_3: Poecilia mexicana Tacotalpa, nonsulfidic, sample 3 *\n PmexTacNS_4: Poecilia mexicana Tacotalpa, nonsulfidic, sample 4 *\n PmexTacNS_5: Poecilia mexicana Tacotalpa, nonsulfidic, sample 5 *\n PmexTacNS_6: Poecilia mexicana Tacotalpa, nonsulfidic, sample 6 *\n PmexTacS_2: Poecilia mexicana Tacotalpa, sulfidic, sample 2 *\n PmexTacS_3: Poecilia mexicana Tacotalpa, sulfidic, sample 3 *\n PmexTacS_4: Poecilia mexicana Tacotalpa, sulfidic, sample 4 *\n PmexTacS_5: Poecilia mexicana Tacotalpa, sulfidic, sample 5 *\n PmexTacS_6: Poecilia mexicana Tacotalpa, sulfidic, sample 6 *\n PlimNS_1: Poecilia limantouri, nonsulfidic, sample 1 * PlimNS_2: Poecilia\n limantouri, nonsulfidic, sample 2 * PlimNS_3: Poecilia limantouri,\n nonsulfidic, sample 3 * PlimNS_4: Poecilia limantouri, nonsulfidic, sample\n 4 * PlimNS_5: Poecilia limantouri, nonsulfidic, sample 5 *\n PlimNS_6: Poecilia limantouri, nonsulfidic, sample 6 #### File:\n envpredictors.csv **Description:** Measurements of water quality\n parameters for all sites included in this study, which served as\n environmental predictor variables during data analysis. ##### Variables *\n id: Row ID * type: Habitat type (all S for sulfidic) * lineage:\n Evolutionary lineage; corresponds to tip label in the tree * lat: Latitude\n (decimal degrees) * long: Longitude (decimal degrees) * DO: Dissolved\n oxygen (mg/l) * Temperature: Temperature (degrees Celsius) * pH: pH *\n SpCond: Specific conductivity (microS/cm) * H2Smol: Hydrogen sulfide\n concentration (mol/l) #### File: gene_annotations.csv\n **Description:** Gene annotation file for all genes in the FPFM matrix.\n This is based on a Blast search against the human SwissProt database.\n ##### Variables * gene ID: A unique identifier assigned to each gene,\n matching the entries in the FPKM matrix. * gene name: The standard or\n predicted name of the gene. * Subject sequence ID: The identifier of the\n reference (subject) sequence in the database that matched the query gene\n during the sequence alignment search. * % of identical matches: The\n percentage of aligned positions where the query and subject sequences have\n exactly the same nucleotide or amino acid. Higher values indicate greater\n sequence similarity. * Alignment length: The total number of positions\n included in the alignment between the query and subject sequences,\n including matches, mismatches, and gaps. * Mismatches: The number of\n aligned positions where the query and subject sequences differ. * Gap\n openings: The number of gaps introduced into the alignment to maximize\n sequence similarity. Gaps represent insertions or deletions (indels). *\n Start of alignment query: The position in the query gene where the\n alignment begins. * End of alignment query: The position in the query gene\n where the alignment ends. * Start of alignment subject: The position in\n the reference (subject) sequence where the alignment begins. * End of\n alignment subject: The position in the reference (subject) sequence where\n the alignment ends. * E-value: The expected number of matches with a\n similar score that would occur by chance in a database search. Lower\n E-values indicate more statistically significant matches. * Bit score: A\n normalized alignment score that reflects the quality of the sequence\n alignment. Higher bit scores indicate stronger sequence similarity and\n more reliable matches. * Protein annotations: Functional information\n associated with the matched protein in the SwissProt\n database.Code/software All data can be opened with a simple text editor or\n imported into R. ## Software versions Analyses were conducted using R\n version 4.6.0 (2026-04-24) on macOS Tahoe 26.6.1. The following R packages\n were used: * ape 5.8-1 * brms 2.23.0 * evemodel 0.0.0.9008 * geodata 0.6-9\n * ggplot2 4.0.3 * ggrepel 0.9.8 * plyr 1.8.9 * terra 1.9-34 * tibble 3.3.1\n ## License Code in this repository is released under the CC0 License. ##\n Access information Gene expression data was derived from sequencing data,\n which are available at the National Center for Biotechnology Information\n (NCBI) under Bio-Project accession numbers PRJNA473350 and PRJNA608180."}],"geoLocations":[],"fundingReferences":[{"funderIdentifierType":"ROR","funderName":"Division of Integrative Organismal Systems","funderIdentifier":"https://ror.org/01rvays47","awardTitle":"\n        ROL: Collaborative research: extreme environments, physiological\n        adaptation, and the origin of species\n      ","awardNumber":"2311366"},{"funderIdentifierType":"ROR","funderName":"Division of Integrative Organismal Systems","funderIdentifier":"https://ror.org/01rvays47","awardTitle":"\n        ROL: Collaborative research: extreme environments, physiological\n        adaptation, and the origin of species\n      ","awardNumber":"2423844"},{"funderName":"Des Lee Collaborative Vision in Zoological Studies"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.8sf7m0d53","contentUrl":null,"metadataVersion":3,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-07T09:46:15Z","registered":"2026-08-07T09:46:16Z","published":null,"updated":"2026-08-20T14:47:37Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.q573n5v02","type":"dois","attributes":{"doi":"10.5061/dryad.q573n5v02","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["Finnish Environment Institute"],"name":"Toivonen, Marjaana","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0002-9215-8643"}]},{"nameType":"Personal","affiliation":["Kone Foundation"],"name":"Humberg, Paula","nameIdentifiers":[]}],"titles":[{"title":"Data from: Comparing time-lapse photography and traditional human observation in collecting crop pollinator visitation data"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Agricultural sciences","subjectScheme":"fos"},{"subject":"camera surveillance"},{"subject":"crop pollination"},{"subject":"flower visitation"},{"subject":"trail camera photography"},{"subject":"Malus domestica"},{"subject":"Carum carvi"},{"subject":"Pollinator community composition"}],"contributors":[],"dates":[{"date":"2026-07-01T10:25:39Z","dateType":"Created"},{"date":"2026-08-18T02:55:16Z","dateType":"Submitted"},{"date":"2026-08-20T00:00:00Z","dateType":"Issued"},{"date":"2026-08-20T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[],"relatedItems":[],"sizes":["404870 bytes"],"formats":[],"version":"2","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Agricultural pollinator management requires knowledge on pollinator-crop\n interactions, including information on pollinator visits to crop flowers.\n Time-lapse trail cameras may provide an efficient and simple tool to\n collect such information. However, research evaluating their performance\n against traditional direct observation of crop pollinators is scarce.\n Here, we compare trail camera and human observer performance for\n collecting flower visitation data in two crops, apple and caraway, in\n boreal farmland. Five commercially available trail cameras recorded\n time-lapse images of crop flowers throughout one flowering season. Near\n each camera, a human observer monitored flower visits five times per crop.\n We examined how monitoring method affects observed pollinator diversity,\n community composition, the working time required for, and the costs of the\n monitoring. In addition, since CAM was run 24-h, it enabled analysis of\n diel activity patterns of flower visitors. In apple, cameras recorded more\n diverse pollinator assemblages than the human observer. This was due to\n better detection of less conspicuous flower visitors during similar\n observation periods, as well as to larger temporal coverage of the camera\n monitoring. Honeybees dominated flower visits in apple but with a lower\n proportion in the camera monitoring. In caraway, hoverflies\n dominated flower visits in both methods. Their proportion was similar\n between the methods during similar observation periods, but lower in the\n camera monitoring, when using full camera data with larger temporal\n coverage. Detecting pollinators from images was harder for caraway than\n apple. Time-lapse photography with manual image analysis was more time-\n and cost-efficient than human observation. In conclusion, trail camera\n photography appears as a competitive method for monitoring pollinator\n flower visits in apple but not in caraway. Automated pollinator detection\n from images can improve the method’s efficiency in the future."},{"descriptionType":"Methods","description":"We compared trail camera (CAM) and human observer (HUM)\n performance for collecting flower visitation data in two crops, apple\n (\u003cem\u003eMalus domestica\u003c/em\u003e) and caraway (\u003cem\u003eCarum\n carvi\u003c/em\u003e) in Southern Finland. Five close-focus trail cameras\n recorded time-lapse images of crop flowers at 10-second intervals\n continuously 24 hours a day throughout one flowering season in one apple\n orchard and one caraway field. The trail camera images were compiled into\n time-lapse videos, from which a researcher recorded flower visits\n manually. Near each camera, the same researcher monitored flower visits\n five times per crop, 15 min at a time. In the apple orchard, she monitored\n one branch close to each camera, and in the caraway field, a 2 × 2 m\n observation plot close to each camera."},{"descriptionType":"TechnicalInfo","description":"# Data from: Comparing time-lapse photography and traditional human\n observation in collecting crop pollinator visitation data Dataset DOI:\n [10.5061/dryad.q573n5v02](https://doi.org/10.5061/dryad.q573n5v02) ##\n Description of the data and file structure The dataset contains ten .csv\n files. 'Apple_flower_visits_CAM.csv' and\n 'Caraway_flower_visits_CAM.csv' include full camera data for\n apple and caraway, respectively.\\ 'Apple_flower_visits_HUM.csv'\n and 'Caraway_flower_visits_HUM.csv' include full human\n observation data for apple and caraway, respectively. \n 'Apple_flower_visits_similar_observation_periods.csv' and\n 'Caraway_flower_visits_similar_observation_periods.csv'\n include total numbers of flower visits, and the number of flower visits by\n main pollinator groups in CAM and HUM during similar observation periods\n in apple and caraway, respectively.\n 'Apple_pollinator_prop_diversity_similar_observation_periods.csv' and 'Caraway_pollinator_prop_diversity_similar_observation_periods.csv' include proportions of flower visits by main pollinator groups, and pollinator diversity in CAM and HUM during similar observation periods in apple and caraway, respectively. The files include only monitoring time and site combinations where both monitoring methods had produced at least one recorded flower visit.  'Apple_pollinator_composition_for_multivariate_analyses.csv' and 'Caraway_pollinator_composition_for_multivariate_analyses.csv' include total numbers of flower visits by each pollinator species and group over the whole monitoring period for each camera and human observation site in apple and caraway, respectively. Pollinator species and groups with less than ten recorded flower visits were excluded. This data were used for the multivariate analyses. ## Files and variables #### File: Apple_flower_visits_HUM.csv **Description:** Human observation data from the study apple orchard. Each row repsesents one observation period (15 min) at one observation site (apple branch, S1-S5). ##### Variables * Site: ID of the observation site (S1-S5) * Date: Calendar date (day.month.year) * Starting_time: Time of starting the observation period (hh:mm) * Observation_round: Sequence number of the human observation round (1-5) * Temperature_Celcius: Temperature in Celcius degrees * Sun_%: Time the sun shone during HUM as a percentage of the total observation time * Wind_Beaufort_scale: Wind on the Beaufort scale * No_open_flowers: Number of open flowers in the observed apple branch * Honeybees: Number of flower visits by honeybees (*A. mellifera*) * Solitary_bees: Number of flower visits by solitary bees * Bombus_lucorum_group: Number of flower visits by bumblebees of the *Bombus lucorum* -group * Bombus_pascuorum: Number of flower visits by *Bombus pascuorum* * Bombus_lapidarius: Number of flower visits by *Bombus lapidarius* * Bombus_all: Total number of flower visits by *Bombus sp.* * Syrphinae: Number of flower visits by Syrphinae * Eristalinae: Number of flower visits by Eristalinae * Syrphidae_all: Total number of flower visits by hoverflies * Other_Brachycera: Number of flower visits by other Brachycera * Heteroptera: Number of flower visits by Heteroptera * Coccinellidae: Number of flower visits by Coccinellidae * Vespidae: Number of flower visits by Vespidae * Other_Hymenoptera: Number of flower visits by other Hymenoptera * All_flower_visits: Total number of all flower visits #### File: Apple_flower_visits_CAM.csv **Description:** Full trail camera data from the study apple orchard. Each row represents observation(s) at one time point at one observation site (S1-S5). When one pollinator made several consecutive flower visits, the time of the first visit and the total number of consecutive visits was recorded in one row. Due to a technical problem, three out of five cameras were inactive for one day between 30 and 31 May. ##### Variables * Site: ID of the observation site (S1-S5) * Date: Calendar date (day.month.year) * h: Hour of day * min: Minute of hour * Honeybees: Number of flower visits by honeybees (*A. mellifera*) * Solitary_bees: Number of flower visits by solitary bees * Syrphinae: Number of flower visits by Syrphinae * Eristalinae: Number of flower visits by Eristalinae * Syrphidae_unidentified: Number of flower visits by unidentified hoverflies * Syrphidae_all: Total number of hoverfly flower visits * Other_Brachycera: Number of flower visits by other Brachycera * Nematocera: Number of flower visits by Nematocera * Other_Diptera_all: Total number of flower visits by Diptera other than hoverflies * Bombus_pascuorum: Number of flower visits by *Bombus pascuorum* * Bombus_lucorum_group: Number of flower visits by bumblebees of the *Bombus lucorum* -group * Bombus_hortorum: Number of flower visits by *Bombus hortorum* * Bombus_lapidarius: Number of flower visits by *Bombus lapidarius* * Bombus_hypnorum: Number of flower visits by *Bombus hypnorum* * Bombus_pratorum: Number of flower visits by *Bombus pratorum* * Bombus_schrencki: Number of flower visits by *Bombus schrencki* * Bombus_all: Total number of bumblebee flower visits * Cetonia_aurata_Protaetia_cuprea: Number of flower visits by *Cetonia aurata* or *Protaetia cuprea* * Coccinellidae: Number of flower visits by Coccinellidae * Cantharidae: Number of flower visits by Cantharidae * Other_Coleoptera: Number of flower visits by other Coleoptera * Coleoptera_all: Total number of flower visits by all Coleoptera * Lepidoptera_unidentified: Number of flower visits by Lepidoptera * Chrysopidae: Number of flower visits by Chrysopidae * Formicidae: Number of flower visits by Formicidae * Heteroptera: Number of flower visits by Heteroptera * All_identified_flower_visitors: Total number of flower visits by flower visitors identified at least to the order level * Unidentified_flower_visitors: Number of flower visits by unidentified flower visitors #### File: Apple_flower_visits_similar_observation_periods.csv **Description:** Total number of flower visits, and the number of flower visits by main pollinator groups in CAM and HUM during similar observation periods in apple. A subset of CAM data includes flower visits from one hour before to one hour after the 15-min HUM observation at the same site. ##### Variables * Site: ID of the observation site (S1-S5) * Observation_round_HUM: Sequence number of the human observation round (1-5) * Method: CAM = camera monitoring method, HUM = human observation method * All_flower_visits: Total number of flower visits * Honeybees: Number of flower visits by honeybees (*Apis mellifera*) * Hoverflies: Number of flower visits by hoverflies * Other_flies: Number of flower visits by other flies * Bumblebees: Number of flower visits by bumblebees #### File: Apple_pollinator_composition_for_multivariate_analyses.csv **Description:** Data used for the multivariate analyses. Total numbers of apple flower visits by each pollinator species and group over the whole monitoring period for each camera and human observation site. Pollinator species and groups with less than ten recorded flower visits were excluded from this data. ##### Variables * Site: ID of the observation site (S1-S5) * Method: CAM = camera monitoring method, HUM = human observation method * A.melli: Number of flower visits by honeybees (*Apis mellifera*) * Solit.bee: Number of flower visits by solitary bees * Syrphinae: Number of flower visits by Syrphinae * Eristalinae: Number of flower visits by Eristalinae * Brachycera: Number of flower visits by other Brachycera * Nematocera: Number of flower visits by Nematocera * B.luco: Number of flower visits by bumblebees of the *Bombus lucorum* -group * B.pasc: Number of flower visits by *Bombus pascuorum* * B.lapi: Number of flower visits by *Bombus lapidarius* * B.hypn: Number of flower visits by *Bombus hypnorum* * B.prat: Number of flower visits by *Bombus pratorum* * Coccinellidae: Number of flower visits by Coccinellidae * Cetoniini: Number of flower visits by* Cetonia aurata* or *Protaetia cuprea* * Chrysopidae: Number of flower visits by* *Chrysopidae * Formicidae: Number of flower visits by* *Formicidae #### File: Apple_pollinator_prop_diversity_similar_observation_periods.csv **Description:** Proportions of flower visits by main pollinator groups, and pollinator diversity in CAM and HUM during similar observation periods in apple. A subset of CAM data includes flower visits from one hour before to one hour after the 15-min HUM observation at the same site. The data includes only monitoring time and site combinations where both monitoring methods had produced at least one recorded flower visit. ##### Variables * Site: ID of the observation site (S1-S5) * Observation_round_HUM: Sequence number of the human observation round (1-5) * Method: CAM = camera monitoring method, HUM = human observation method * Honeybees_prop: Proportion of flower visits by honeybees (*Apis mellifera*) * Hoverflies_prop: Proportion of flower visits by hoverflies * Other_flies_prop: Proportion of flower visits by other flies * Bumblebees_prop: Proportion of flower visits by bumblebees * Shannon_diversity: Shannon-Wiener diversity index (*H’*) describing the pollinator diversity of flower visits #### File: Caraway_flower_visits_similar_observation_periods.csv **Description:** Total number of flower visits, and the number of flower visits by main pollinator groups in CAM and HUM during similar observation periods in caraway. A subset of CAM data includes flower visits from one hour before to one hour after the 15-min HUM observation at the same site. ##### Variables * Site: ID of the observation site (S1-S5) * Observation_round_HUM: Sequence number of the human observation round (1-5) * Method: CAM = camera monitoring method, HUM = human observation method * All_flower_visits: Total number of all flower visits * Hoverflies: Number of flower visits by hoverflies * Other_flies: Number of flower visits by other flies #### File: Caraway_flower_visits_CAM.csv **Description:** Full trail camera data from the study caraway field. Each row represents observation(s) at one time point at one observation site (S1-S5). When one pollinator made several consecutive flower visits, the time of the first visit and the total number of consecutive visits was recorded in one row. ##### Variables * Site: ID of the observation site (S1-S5) * Date: Calendar date (day.month.year) * h: Hour of day  * min: Minute of hour * Syrphinae: Number of flower visits by Syrphinae * Eristalinae: Number of flower visits by Eristalinae * Syrphidae_unidentified: Number of flower visits by unidentified hoverflies * Syphidae_all: Total number of hoverfly flower visits * Sarcophagidae: Number of flower visits by Sarcophagidae * Lucilia: Number of flower visits by *Lucilia* sp. * Other_Brachycera: Number of flower visits by other Brachycera * Tipulidae: Number of flower visits by Tipulidae * Other_Nematocera: Number of flower visits by other Nematocera * Other_Diptera_all: Total number of flower visits by Diptera other than hoverflies * Coccinellidae: Number of flower visits by Coccinellidae * Cetonia_aurata_Protaetia_cuprea: Number of flower visits by* Cetonia aurata* or *Protaetia cuprea* * Cantharidae: Number of flower visits by* *Cantharidae * Other_Coleoptera: Number of flower visits by other Coleoptera * Coleoptera_all: Total number of Coleoptera flower visits * Heteroptera: Number of flower visits by* *Heteroptera * Honeybees: Number of flower visits by honeybees (*Apis mellifera*) * Solitary_bees: Number of flower visits by solitary bees * Bombus_lucorum: Number of flower visits by bumblebees of the *Bombus lucorum* -group * Vespidae: Number of flower visits by Vespidae * Symphyta: Number of flower visits by Symphyta * Other_Hymenoptera: Number of flower visits by other Hymenoptera * Chrysopidae: Number of flower visits by Chrysopidae * Panorpidae: Number of flower visits by Panorpidae * Araneae: Number of flower visits by Araneae * All_identified_flower_visitors: Total number of flower visits by flower visitors identified at least to the order level * Unidentified_flower_visitors: Number of flower visits by unidentified flower visitors #### File: Caraway_flower_visits_HUM.csv **Description:** Human observation data from the study caraway orchard. Each row repsesents one observation period (15 min) at one observation site (2 × 2 m observation plot, S1-S5). ##### Variables * Site: ID of the observation site (S1-S5) * Date: Calendar date (day.month.year) * Starting_time: Time of starting the observation period (hh:mm) * Observation_round: Sequence number of the human observation round (1-5) * Temperature_Celcius: Temperature in Celcius degrees * Sun_%: Time the sun shone during HUM as a percentage of the total observation time * Wind_Beaufort_scale: Wind on the Beaufort scale * Cover_%_flowering_caraway: Percentage cover of flowering caraway in the observation plot * Syrphinae: Number of flower visits by Syrphinae * Eristalinae: Number of flower visits by Eristalinae * Syrphidae_all: Total number of hoverfly flower visits * Sarcophagidae: Number of flower visits by Sarcophagidae * Lucilia: Number of flower visits by *Lucilia* sp. * Other_Brachycera: Number of flower visits by other Brachycera * Tipulidae: Number of flower visits by Tipulidae * Other_Nematocera: Number of flower visits by other Nematocera * Other_Diptera_all: Total number of flower visits by Diptera other than hoverflies * Heteroptera: Number of flower visits by Heteroptera * Coccinellidae: Number of flower visits by Coccinellidae * Cetonia_aurata_Protaetia_cuprea: Number of flower visits by* Cetonia aurata* or *Protaetia cuprea* * Cantharidae: Number of flower visits by Cantharidae * Other_Coleoptera: Number of flower visits by other Coleoptera * Coleoptera_all: Total number of Coleoptera flower visits * Honeybees: Number of flower visits by honeybees (*Apis mellifera*) * Solitary_bees: Number of flower visits by solitary bees * Vespidae: Number of flower visits by Vespidae * Symphyta: Number of flower visits by Symphyta * Hymenoptera_unidentified: Number of flower visits by other Hymenoptera * Chrysopidae: Number of flower visits by Chrysopidae * Panorpidae: Number of flower visits by Panorpidae * Araneae: Number of flower visits by Araneae * All_flower_visits: Total number of all flower visits #### File: Caraway_pollinator_composition_for_multivatiate_analyses.csv **Description:** Data used for the multivariate analyses. Total numbers of caraway flower visits by each pollinator species and group over the whole monitoring period for each camera and human observation site. Pollinator species and groups with less than ten recorded flower visits were excluded from this data. ##### Variables * Site: ID of the observation site (S1-S5) * Method: CAM = camera monitoring method, HUM = human observation method * Syrphinae: Number of flower visits by Syrphinae * Eristalinae: Number of flower visits by Eristalinae * Sarcophagidae: Number of flower visits by Sarcophagidae * Lucilia: Number of flower visits by *Lucilia* sp. * Brachycera: Number of flower visits by other Brachycera * Tipulidae: Number of flower visits by Tipulidae * Nematocera: Number of flower visits by other Nematocera * Heteroptera: Number of flower visits by Heteroptera * Coccinellidae: Number of flower visits by Coccinellidae * Cetoniini: Number of flower visits by* Cetonia aurata* or *Protaetia cuprea* * Cantharidae: Number of flower visits by Cantharidae * Coleoptera: Number of flower visits by other Coleoptera * A.melli: Number of flower visits by honeybees (*Apis mellifera*) * solit.bee: Number of flower visits by solitary bees * Chrysopidae: Number of flower visits by #### File: Caraway_pollinator_prop_diversity_similar_observation_periods.csv **Description:** Proportions of flower visits by main pollinator groups, and pollinator diversity in CAM and HUM during similar observation periods in caraway. A subset of CAM data includes flower visits from one hour before to one hour after the 15-min HUM observation at the same site. The data includes only monitoring time and site combinations where both monitoring methods had produced at least one recorded flower visit. ##### Variables * Site: ID of the observation site (S1-S5) * Observation_round_HUM: Sequence number of the human observation round (1-5) * Method: CAM = camera monitoring method, HUM = human observation method * Hoverflies_prop: Proportion of flower visits by hoverflies * Other_flies_prop: Proportion of flower visits by other flies * Shannon_diversity: Shannon-Wiener diversity index (*H’*) describing the pollinator diversity of flower visits ## Code/software We used the standard desktop version of Microsoft Excel to digitize our data. For the statistical analyses, the files were saved as CSV and imported into the open-source statistical analysis software RStudio with R version 4.4.2. In RStudio, data were analyzed and visualized using the packages glmmTMB, car, DHARMa, stats, vegan, ggeffects and ggplot2."}],"geoLocations":[],"fundingReferences":[{"funderIdentifierType":"ROR","funderName":"Kone Foundation","funderIdentifier":"https://ror.org/05jwty529"},{"funderIdentifierType":"ROR","funderName":"Arts Promotion Centre Finland","funderIdentifier":"https://ror.org/01vmr1p43"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.q573n5v02","contentUrl":null,"metadataVersion":0,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-20T13:52:40Z","registered":"2026-08-20T13:52:41Z","published":null,"updated":"2026-08-20T13:52:41Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.vdncjszbn","type":"dois","attributes":{"doi":"10.5061/dryad.vdncjszbn","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["University of Wyoming"],"name":"Smith, Austin","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0002-2095-5299"}]},{"nameType":"Personal","affiliation":["Wyoming Game and Fish Department"],"name":"Hall, Embere","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Wyoming Game and Fish Department"],"name":"Clapp, Justin","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0002-8099-1830"}]},{"nameType":"Personal","affiliation":["United States Department of Agriculture"],"name":"Squires, John","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Wyoming"],"name":"Bennett, Drew","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Wyoming Game and Fish Department"],"name":"Maichak, Eric","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Wyoming Game and Fish Department"],"name":"O'Brien, Heather","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Wyoming Game and Fish Department"],"name":"Turnbull, Zach","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Wyoming Game and Fish Department"],"name":"Newkirk, Eric","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Waterloo"],"name":"Fedy, Bradley","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Waterloo"],"name":"Kirol, Christopher","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Agricultural Research Service"],"name":"Porensky, Lauren","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Agricultural Research Service"],"name":"Dufek, Nickolas","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Wyoming"],"name":"Buxbaum, Kelsie","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Arizona"],"name":"Duchardt, Courtney","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Wyoming"],"name":"Paolini, Kelsey","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["University of Wyoming"],"name":"Holbrook, Joseph","nameIdentifiers":[]}],"titles":[{"title":"Data and code from: Mapping for management: Habitat suitability of bobcats across Wyoming, USA"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Wildlife","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Remote sensing","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Ecology","subjectScheme":"PLOS Subject Area Thesaurus"}],"contributors":[],"dates":[{"date":"2026-08-17T18:44:56Z","dateType":"Created"},{"date":"2026-08-17T18:44:58Z","dateType":"Submitted"},{"date":"2026-08-20T00:00:00Z","dateType":"Issued"},{"date":"2026-08-20T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[],"relatedItems":[],"sizes":["66816892 bytes"],"formats":[],"version":"2","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Wildlife conservation requires reliable information on species’\n distribution, demography, and habitat relationships. Carnivores pose\n challenges for understanding their distributions due to their rarity, high\n vagility, and need for large, interconnected habitats. The bobcat (Lynx\n rufus) is a cryptic and elusive carnivore that inhabits diverse habitats\n across much of North America, complicating a synthetic understanding of\n habitat suitability for bobcats. Moreover, bobcats inhabiting the\n mountainous landscapes of western North America have received\n comparatively little attention in terms of scientific investigations, and\n thus our goal was to spatially characterize bobcat habitat suitability\n using remote cameras and direct observations of bobcats across the state\n of Wyoming, USA. We paired remote camera detections (n = 73) and direct\n observations (n = 337) with environmental covariates to develop a\n statewide bobcat habitat suitability model using Random Forest. We\n assessed model performance by estimating the out-of-bag error and\n validated predictive performance using 5 withheld subsets of bobcat\n observations (i.e., a random sample of 50% of the direct observations).\n Our Random Forest model achieved a mean out-of-bag error of 15.93%.\n Independent evaluation indicated a strong correlation (mean rs = 0.93, p\n \u0026lt; 0.001) between predicted habitat suitability and withheld bobcat\n observations. Predicted bobcat habitat suitability was positively related\n to heat load index, vector ruggedness measurement, canopy cover, foliage\n height diversity, herbaceous biomass production, Normalized Difference\n Vegetation Index, and tree occurrence. Contrastingly, predicted\n suitability was negatively associated with snow depth, topographic\n position index, distance to agriculture, and shrub occurrence. Results\n from our model estimated that approximately 50% of Wyoming has moderately\n to highly suitable habitat for bobcats. This work provided essential\n information on bobcat-environment relationships in Wyoming and within the\n Rocky Mountain chain, but additional research (e.g., diet, movement,\n population density, and genetics) is necessary to gain more granular\n insights on bobcats. Our analytical process can be applied in other\n environmental contexts to predict habitat suitability for bobcats and\n other elusive, widespread carnivores."},{"descriptionType":"TechnicalInfo","description":"# Data and code from: Mapping for management: Habitat suitability of\n bobcats across Wyoming, USA ## Description of the data and file structure\n The data were collected from remote cameras and direct observations\n throughout Wyoming, USA, as well as from various remotely sensed products.\n Because the deployment locations of several remote camera stations are\n sensitive and to protect the privacy of many private landowners, we\n excluded coordinates from the datasets. The R script\n (Bobcat_HSM_WY_Analyses.R)and accompanying RDS files contain all the\n material necessary to replicate our analyses. Below, we provide additional\n context for the RDS files, such as column names, column descriptions, etc.\n #### Remote camera detections (1) and pseudo-absence (0) with extracted\n covariate values **CameraDetections_PseudoAbsences.rds** Column names and\n information: * ID = 1 = camera detections; 0 = pseudo-absence * hli = heat\n load index (index) * sdepth = snow depth (m) * tpi = topographic position\n index (index) * vrm = vector ruggedness measurement (index) * cover =\n fractional canopy cover (percentage) * Dist2ag = distance to agriculture\n (m) * fhd = foliage height diversity (na) * herb = herbaceous biomass\n production (kg ha\u003csup\u003e-1\u003c/sup\u003e) * ndvi = normalized difference\n vegetation index (index) * shrub = shrub occurrence (proportion) * tree =\n tree occurrence (proportion) #### Withheld bobcat observations\n **WithheldObservations.rds** Column names and information: * ID = 1 =\n withheld observations * hli = heat load index (index) * sdepth = snow\n depth (m) * tpi = topographic position index (index) * vrm = vector\n ruggedness measurement (index) * cover = fractional canopy cover\n (percentage) * Dist2ag = distance to agriculture (m) * fhd = foliage\n height diversity (na) * herb = herbaceous biomass production (kg\n ha\u003csup\u003e-1\u003c/sup\u003e) * ndvi = normalized difference vegetation\n index (index) * shrub = shrub occurrence (proportion) * tree = tree\n occurrence (proportion) #### Random Forest training data - model 1\n **Model1_FinalTraining_Dataset.rds** Column names and information: * ID =\n 1 = presence; 0 = pseudo-absence * hli = heat load index (index) * sdepth\n = snow depth (m) * tpi = topographic position index (index) * vrm = vector\n ruggedness measurement (index) * cover = fractional canopy cover\n (percentage) * Dist2ag = distance to agriculture (m) * fhd = foliage\n height diversity (na) * herb = herbaceous biomass production (kg\n ha\u003csup\u003e-1\u003c/sup\u003e) * ndvi = normalized difference vegetation\n index (index) * shrub = shrub occurrence (proportion) * tree = tree\n occurrence (proportion) #### Random Forest training data - model 2\n **Model2_FinalTraining_Dataset.rds** Column names and information: * ID =\n 1 = presence; 0 = pseudo-absence * hli = heat load index (index) * sdepth\n = snow depth (m) * tpi = topographic position index (index) * vrm = vector\n ruggedness measurement (index) * cover = fractional canopy cover\n (percentage) * Dist2ag = distance to agriculture (m) * fhd = foliage\n height diversity (na) * herb = herbaceous biomass production (kg\n ha\u003csup\u003e-1\u003c/sup\u003e) * ndvi = normalized difference vegetation\n index (index) * shrub = shrub occurrence (proportion) * tree = tree\n occurrence (proportion) #### Random Forest training data - model 3\n **Model3_FinalTraining_Dataset.rds** Column names and information: * ID =\n 1 = presence; 0 = pseudo-absence * hli = heat load index (index) * sdepth\n = snow depth (m) * tpi = topographic position index (index) * vrm = vector\n ruggedness measurement (index) * cover = fractional canopy cover\n (percentage) * Dist2ag = distance to agriculture (m) * fhd = foliage\n height diversity (na) * herb = herbaceous biomass production (kg\n ha\u003csup\u003e-1\u003c/sup\u003e) * ndvi = normalized difference vegetation\n index (index) * shrub = shrub occurrence (proportion) * tree = tree\n occurrence (proportion) #### Random Forest training data - model 4\n **Model4_FinalTraining_Dataset.rds** Column names and information: * ID =\n 1 = presence; 0 = pseudo-absence * hli = heat load index (index) * sdepth\n = snow depth (m) * tpi = topographic position index (index) * vrm = vector\n ruggedness measurement (index) * cover = fractional canopy cover\n (percentage) * Dist2ag = distance to agriculture (m) * fhd = foliage\n height diversity (na) * herb = herbaceous biomass production (kg\n ha\u003csup\u003e-1\u003c/sup\u003e) * ndvi = normalized difference vegetation\n index (index) * shrub = shrub occurrence (proportion) * tree = tree\n occurrence (proportion) #### Random Forest training data - model 5\n **Model5_FinalTraining_Dataset.rds** Column names and information: * ID =\n 1 = presence; 0 = pseudo-absence * hli = heat load index (index) * sdepth\n = snow depth (m) * tpi = topographic position index (index) * vrm = vector\n ruggedness measurement (index) * cover = fractional canopy cover\n (percentage) * Dist2ag = distance to agriculture (m) * fhd = foliage\n height diversity (na) * herb = herbaceous biomass production (kg\n ha\u003csup\u003e-1\u003c/sup\u003e) * ndvi = normalized difference vegetation\n index (index) * shrub = shrub occurrence (proportion) * tree = tree\n occurrence (proportion) #### Validation: Reserved (withheld) data Contains\n each model's withheld data intersected with the accompanying model\n **Validation_HSM_Values_allModels.rds** * ID = reserved (withheld) bobcat\n observations across Wyoming * Binnum = predicted bobcat habitat\n suitability values (binned 1 to 10) intersected by each ID * model =\n represents which model (1 through 5) the ID and Binnum are associated with\n #### Pseudo-absence locations with extracted covariate values and\n predicted habitat suitability model binned values to construct Figure 4.\n **PseudoAbsence_Covariates.rds** Column names and information: * sdepth =\n snow depth (m) * hli = heat load index (index) * tpi = topographic\n position index (index) * vrm = vector ruggedness measurement (index) *\n cover = fractional canopy cover (percentage) * Dist2ag = distance to\n agriculture (m) * fhd = foliage height diversity (na) * herb = herbaceous\n biomass production (kg ha\u003csup\u003e-1\u003c/sup\u003e) * ndvi = normalized\n difference vegetation index (index) * shrub = shrub occurrence\n (proportion) * tree = tree occurrence (proportion) * HSM_values =\n predicted bobcat habitat suitability values (binned 1 to 10) intersected\n by each pseudo-absence location The R code is divided into seven sections:\n The first section outlines the workflow for randomly sampling half of the\n withheld observations with replacement for each model (n = 5), and joining\n the randomly sampled withheld observations with the remote camera bobcat\n detections, while also reserving the non-sampled withheld observations for\n model validation of the accompanying model. The second section employs\n Random Forest modeling to assess the environmental gradients that\n characterize bobcat habitat and predict suitable habitat across Wyoming.\n The third section provides the workflow for reclassifying each predicted\n model in 10-equal area bins. The fourth section validates the predicted\n habitat suitability maps using reserved (withheld) data. The fifth section\n provides the workflow to create the final habitat suitability maps by\n calculating the mean and standard deviation of the five models. The sixth\n section provides the steps to reclassify the final habitat suitability map\n into 10 equal-area bins. The final section focuses on understanding trends\n in bobcat habitat suitability and its covariates."}],"geoLocations":[],"fundingReferences":[{"funderName":"Wyoming Governor’s Big Game License Coalition"},{"funderIdentifierType":"ROR","funderName":"Wyoming Game and Fish Department","funderIdentifier":"https://ror.org/046em8f15"},{"funderIdentifierType":"ROR","funderName":"University of Wyoming","funderIdentifier":"https://ror.org/01485tq96"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.vdncjszbn","contentUrl":null,"metadataVersion":0,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-20T12:14:29Z","registered":"2026-08-20T12:14:30Z","published":null,"updated":"2026-08-20T12:14:30Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.b2rbnzssq","type":"dois","attributes":{"doi":"10.5061/dryad.b2rbnzssq","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["Linköping University"],"name":"Brorsson, Ann-Christin","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0001-7651-3556"}]},{"nameType":"Personal","affiliation":["Linköping University"],"name":"Elovsson, Greta","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Linköping University"],"name":"Klingstedt, Therese","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Linköping University"],"name":"Nilsson, Peter","nameIdentifiers":[]}],"titles":[{"title":"Data and code from: Diversity of Aβ aggregates produced in a gut-based \u003cem\u003eDrosophila\u003c/em\u003e model of Alzheimer’s disease"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Drosophila melanogaster","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Protein misfolding","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Alzheimer's disease","subjectScheme":"PLOS Subject Area Thesaurus"}],"contributors":[],"dates":[{"date":"2025-06-17T09:49:47Z","dateType":"Created"},{"date":"2025-06-17T09:49:50Z","dateType":"Submitted"},{"date":"2026-08-20T00:00:00Z","dateType":"Issued"},{"date":"2026-08-20T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[{"relationType":"IsCitedBy","relatedIdentifier":"10.1101/2024.11.19.624423","relatedIdentifierType":"DOI"},{"relationType":"IsCitedBy","relatedIdentifier":"10.1371/journal.pone.0314832","relatedIdentifierType":"DOI"}],"relatedItems":[],"sizes":["445899 bytes"],"formats":[],"version":"8","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Alzheimer’s disease (AD) is a neurodegenerative disease manifested by\n memory loss and premature death. One major histopathological\n hallmark of AD is the amyloid plaques formed by aggregates of the\n amyloid-β (Aβ) peptide and the Aβ aggregation process results in\n amyloid fibrils with different structures. Herein, we investigate\n the heterogeneity of Aβ aggregates produced by Drosophila\n melanogaster expressing the Aβ1-42 peptide with the Arctic\n mutation E22G (Arctic flies) or a dimeric construct of Aβ1-42\n (T22Aβ1-42 flies) in the digestive tract. Staining of the gut of the flies\n using luminescent conjugated oligothiophenes (LCOs) revealed that\n the amount of Aβ aggregates increased in both genotypes with age.\n The LCOs also exhibited distinct staining patterns in the flies.\n The expression of T22Aβ1-42 resulted in a heavier Aβ load\n compared to Aβ1-42 with the Arctic mutation. Since the genotypes have\n similar median survival times, the result indicates that the\n toxicity of the combined number of aggregates in the Arctic flies\n is higher compared to the T22Aβ1-42 flies. Stability measurements\n showed that the most accumulated Aβ species in the Arctic and\n the T22Aβ1-42 flies were found in the 4 M and 5 M\n Gua-HCl-fraction, respectively. This indicates that prefibrillar\n Aβ aggregates constitute the toxic species in Arctic flies\n while the cause of death in T22Aβ1-42 flies might be the massive\n load of insoluble aggregates. The study shows that even though\n the different Aβ peptides resulted in an equal reduction of the\n lifespan, they formed an array of different aggregates\n confirming the heterogeneity of this process. Overall, our\n findings support that distinct Aβ aggregates can exhibit\n different pathological effects, and we foresee that\n our Drosophila models can potentially aid in identifying anti-Aβ\n agents targeting different types of aggregated Aβ species."},{"descriptionType":"Methods","description":"We utilized \u003cem\u003eDrosophila melanogaster\u003c/em\u003e to\n overexpress Aβ peptides in the fly gut in order to investigate their\n toxicity and morphological characteristics. Fly lines expressing Aβ were\n generously provided by D. Crowther (AstraZeneca, Floceleris, Oxbridge\n Solutions Ltd., London, United Kingdom). Gut barrier integrity was\n assessed visually by incorporating erioglaucine disodium salt into the fly\n food. Aβ aggregates were detected using aggregate-binding fluorophores and\n analyzed with an inverted Zeiss LSM 780 laser scanning confocal microscope\n (Zeiss, Oberkochen, Germany). Quantification of Aβ levels in the flies was\n performed using the Meso Scale Discovery (MSD) assay."},{"descriptionType":"TechnicalInfo","description":"# Data and code from: Diversity of Aβ aggregates produced in a gut-based\n *Drosophila* model of Alzheimer’s disease Dataset DOI:\n [10.5061/dryad.b2rbnzssq](10.5061/dryad.b2rbnzssq) ## Description of the\n data and file structure * **Smurf assay:** The gut leakage analysis was\n performed when the flies were 8 days old. The smurf phenotype of the flies\n was analysed where spreading of the blue dye in the hemocoel and body\n indicated gut leakage. Data were processed in excel. * **Quantifications\n of the co-staining in T22Aβ1-42- and Arctic** **flies**: Nonbiased scoring\n of Aβ1-42 species per mm2 in Arctic- and T22Aβ1-42 flies at day 8, n=6.\n Only Aβ1-42 species where HS-84 or HS-169 colocalized with the antibody\n were counted. Data were processed in GraphPad Software 9 *\n **Quantification of Aβ species by MSD analysis:** Data were collected\n using A V-PLEX Human Aβ1-42 Peptide (6E10) kit (K151LBE-1, Meso Scale\n Discovery). Data were transformed to excel. * **Statistical analysis:**\n The data were analyzed using GraphPad Software 9. ### Files and variables\n #### File: Data_gut_leakage_v1_3.xlsx **Description:** Data from the gut\n leakage analysis. ##### Variables * Vial#: Refers to the number of the\n vial that was examined * Age (days): Refers to the age of the flies in\n days * Genotype: Refers to the genotype of the fly * no smurf: Refers to\n the amout of flies with no smurf phenotype * Unsure: Refers to the amout\n of flies that have an unsure smurf phenotype * Smurf: Refers to the amount\n of flies with smurf phenotype #### File: Data_MSD_20240821v1_3.xlsx\n **Description:** Data from quantification of Aβ species by MSD analysis at\n 0, 2.5, 4 and 5 M Gua n/a = not applicable ##### Variables **Sheet - Raw\n data**: This sheet contains all raw data * Sample: Refers to the unit\n tested within a serie. *  Assay: Refers to the MSD assay that was used. *\n Well: Refers to the number of the well that was used in the MSD plate. *\n Spot: Refers to the spot in the well that was used in the MSD plate. *\n Dilution: Refers to the dilution factor. * Concentration: Values of\n concentrations that were used in the standard. * Signal: Refers to raw\n electrochemiluminescence (ECL) signal. * Adjusted signal: Refers to raw\n electrochemiluminescence (ECL) signal after correcting for background\n noise, plate effects, or other non‑specific sources of light. * Mean:\n Refers to the mean of raw electrochemiluminescence (ECL) signals for\n repeats. * Adj. Sig. Mean: Refers to the mean of raw\n electrochemiluminescence (ECL) signal after correcting for background\n noise, plate effects, or other non‑specific sources of light for repeats.\n * CV: Refers to coefficient of variation. It is a measure of how\n consistent (or variable) the replicate measurements are. * % Recover:\n Refers to how close the measured value is to the expected value. * %\n Recovery Mean: Refers to the mean of % Recover for repets * Calc.\n Concentration (pg/ml): Refers to calculated concentration of the tested\n sample from the measurement. * Calc. Conc. Mean (pg/ml): Refers to\n calculated mean concentration of tested replicates from the measurement. *\n Calc. Conc. CV: Refers to coefficient of variation for the calculated\n concentrations. #### File: Data_MSD_240221v1_3.xlsx **Description:** Data\n from quantification of Aβ species by MSD analysis ot 0 and 5 M Gua n/a =\n not applicable ##### Variables **Sheet - Raw data**: This sheet contains\n all raw data * Sample: Refers to the unit tested within a serie. *  Assay:\n Refers to the MSD assay that was used. * Well: Refers to the number of the\n well that was used in the MSD plate. * Spot: Refers to the spot in the\n well that was used in the MSD plate. * Dilution: Refers to the dilution\n factor. * Concentration: Values of concentrations that were used in the\n standard. * Signal: Refers to raw electrochemiluminescence (ECL) signal. *\n Adjusted signal: Refers to raw electrochemiluminescence (ECL) signal after\n correcting for background noise, plate effects, or other non‑specific\n sources of light. * Mean: Refers to the mean of raw\n electrochemiluminescence (ECL) signals for repeats. * Adj. Sig. Mean:\n Refers to the mean of raw electrochemiluminescence (ECL) signal after\n correcting for background noise, plate effects, or other non‑specific\n sources of light for repeats. * CV: Refers to coefficient of variation. It\n is a measure of how consistent (or variable) the replicate measurements\n are. * % Recover: Refers to how close the measured value is to the\n expected value. * % Recovery Mean: Refers to the mean of % Recover for\n repets * Calc. Concentration (pg/ml): Refers to calculated concentration\n of the tested sample from the measurement. * Calc. Conc. Mean\n (pg/ml): Refers to calculated mean concentration of tested replicates from\n the measurement. * Calc. Conc. CV: Refers to coefficient of variation for\n the calculated concentrations. #### File:\n Data_Quantification_of_aggregates.prism **Description:** Data from\n quantifications of the co-staining in T22Aβ1-42- and Arctic flies ##\n Code/software * **Statistical analysis:** The data were analyzed using\n GraphPad Software 9. ## #### ####"}],"geoLocations":[],"fundingReferences":[{"funderIdentifierType":"ROR","funderName":"Gun och Bertil Stohnes Stiftelse","funderIdentifier":"https://ror.org/00cv16a87"},{"funderIdentifierType":"ROR","funderName":"Åhlén-Stiftelsen","funderIdentifier":"https://ror.org/00gsykk45"},{"funderIdentifierType":"ROR","funderName":"Alzheimerfonden","funderIdentifier":"https://ror.org/027ey5735"},{"funderIdentifierType":"ROR","funderName":"Torsten Söderbergs Stiftelse","funderIdentifier":"https://ror.org/03p9y1j46"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.b2rbnzssq","contentUrl":null,"metadataVersion":0,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":3,"referenceCount":0,"citationCount":2,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-20T11:28:28Z","registered":"2026-08-20T11:28:29Z","published":null,"updated":"2026-08-20T11:28:29Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}},{"id":"10.5061/dryad.r4xgxd2v7","type":"dois","attributes":{"doi":"10.5061/dryad.r4xgxd2v7","identifiers":[],"creators":[{"nameType":"Personal","affiliation":["University of Minnesota"],"name":"Jones, Joshua","nameIdentifiers":[{"nameIdentifierScheme":"ORCID","schemeUri":"https://orcid.org","nameIdentifier":"https://orcid.org/0000-0003-2010-861X"}]},{"nameType":"Personal","affiliation":["Indiana University"],"name":"Newton, Irene","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Indiana University"],"name":"Moczek, Armin","nameIdentifiers":[]}],"titles":[{"title":"Data from: Microbiome-mediated immunity: Limited support from the gazelle dung beetle"}],"publisher":"Dryad","container":{},"publicationYear":2026,"subjects":[{"schemeUri":"https://web-archive.oecd.org/2012-06-15/138575-38235147.pdf","subject":"FOS: Biological sciences","subjectScheme":"fos"},{"subject":"Onthophagus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Microbiome","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Dung beetles","subjectScheme":"PLOS Subject Area Thesaurus"},{"schemeUri":"https://github.com/PLOS/plos-thesaurus","subject":"Serratia marcescens","subjectScheme":"PLOS Subject Area Thesaurus"},{"subject":"Digitonthophagus"}],"contributors":[{"name":"Indiana University","contributorType":"Sponsor","affiliation":[],"nameIdentifiers":[]}],"dates":[{"date":"2026-05-28T17:28:11Z","dateType":"Created"},{"date":"2026-08-08T09:40:13Z","dateType":"Submitted"},{"date":"2026-08-20T00:00:00Z","dateType":"Issued"},{"date":"2026-08-20T00:00:00Z","dateType":"Available"}],"language":"en","types":{"schemaOrg":"Dataset","resourceTypeGeneral":"Dataset","citeproc":"dataset","bibtex":"misc","ris":"DATA","resourceType":"dataset"},"relatedIdentifiers":[],"relatedItems":[],"sizes":["12537 bytes"],"formats":[],"version":"3","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/publicdomain/zero/1.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Zero v1.0 Universal","rightsIdentifier":"cc0-1.0"}],"descriptions":[{"descriptionType":"Abstract","description":"The data presented here were used to assess the role of a host\n insect's (Digitonthophagus gazella) complex microbiome in supporting\n host immune functions. Data include basal protein concentrations, metrics\n for phenoloxidase and lysozyme-like activity in the larval hemolymph, and\n survival rates for hosts experiencing experimentally induced and\n standardized immune challenges. The accompanying publication includes\n specifics on the beetle population of origin, the developmental time\n points of these different tests, and the subsequent analyses used."},{"descriptionType":"TechnicalInfo","description":"# Data from: Microbiome-mediated immunity: Limited support from the\n gazelle dung beetle Dataset DOI:\n [10.5061/dryad.r4xgxd2v7](https://doi.org/10.5061/dryad.r4xgxd2v7) ##\n Description of the data and file structure 1. Immune_potential_assay.csv:\n Protein concentration (ug/uL), phenoloxidase activity (change in\n absorbance per hour), lysozyme-like activity (change in absorbance across\n five hours), and mass (mcg) of beetles reared with or without their\n maternal microbiome. Immune metrics were collected from the larval\n hemolymph and measured using various methods described in the associated\n paper. 2. Protein_concentration_across_development.csv: Protein\n concentration (ug/uL) and beetle mass (mcg) of larvae at three subsequent\n developmental timepoints. Protein concentration was measured using methods\n described in the associated paper. 3. Serratia_16S_sequence.fasta: 16S\n sequence of the *Serratia* isolate used throughout the paper. 4.\n egg_surface_inoculation_metrics.csv: Survival data of beetles developing\n in environments enriched for *Serratia* and either provided their maternal\n microbiome inoculum or deprived of it. 5.\n hemolymph_inoculation_survival_metrics.csv: Survival data, hemolymph\n protein concentrations (ug/uL), and masses (mcg) of beetles injected with\n *Serratia* partway through larval development with or without their\n maternal microbiome inoculum. ### Files and variables #### File:\n Digitonthophagus_immunity_data.zip **Description:** Contains 5 files:\n **Immune_potential_assay.csv:** * plate/well/batch - Beetles were reared\n in individual wells (well) across multiple 12-well plates (plate). In some\n experiments, beetles had to be collected from several different batches of\n adult egg-laying setups (batch). * microbiome - If the beetle received (+)\n or was deprived of (-) their maternal microbiome. * mass_10 - Their mass\n at day 10 post-hatching (mcg). * PC - Hemolymph protein concentration\n (ug/uL). * PO - Hemolymph phenoloxidase activity (change in absorbance per\n hour). * LYS - Hemolymph lysozyme-like activity (change in absorbance\n across five hours). **Protein_concentration_across_development.csv:** *\n plate/well - Same as above. * microbiome - Same as above. * weight -\n Larval mass at the time of measuring (mcg). * age - Days post-hatching\n when measurement was taken. * pc - Hemolymph protein concentration\n (ug/uL). **Serratia_16S_sequence.fasta** : * Sequence from *Serratia* 16S\n amplicon sequencing. **egg_surface_inoculation_metrics.csv**: *\n batch/plate/well - Same as above. * microbiome - Same as above. * pathogen\n - If the individual was inoculated with *Serratia* (+) or PBS (-). *\n infection - (Not relevant to analysis) If the individual was inoculated\n with a bacterial pathogen or a fungal pathogen. * hatch - Whether the\n individual died (0) as an egg or hatched (1). * pupation_mass -\n Individual's mass after pupation (mcg). \"n/a\" in column\n represents individuals that died prior to pupation and, therefore, could\n not be weighed after pupation. * survival - Whether the individual died\n before pupation (0) or not (1).\n **hemolymph_inoculation_survival_metrics.csv**: * well - Same as above. *\n sample - Individual identifier. * microbiome - Same as before. *\n inoculation - Whether the individual was injected with *Serratia* (+) or\n PBS (-). * weight - Mass prior to inoculation (mcg). \"n/a\"\n represents individuals where the weight was not recorded during\n inoculation. * survive - Whether the individual died before pupation (0)\n or not (1). * PC - protein concentration prior to inoculation (ug/uL).\n \"n/a\" represents individuals that were not used for protein\n concentration."}],"geoLocations":[],"fundingReferences":[{"funderIdentifierType":"ROR","funderName":"Division of Integrative Organismal Systems","funderIdentifier":"https://ror.org/01rvays47","awardTitle":"\n        Cis-regulation and conditional chromatin remodeling in development and\n        evolution of ontogenies in horned beetles\n      ","awardNumber":"2243725"},{"funderIdentifierType":"ROR","funderName":"Division of Integrative Organismal Systems","funderIdentifier":"https://ror.org/01rvays47","awardTitle":"\n        Origin and diversification of evolutionary novelties: insights through\n        the study of beetle horns, insect wings, and bilaterian heads\n      ","awardNumber":"1901680"},{"funderIdentifierType":"ROR","funderName":"Division of Graduate Education","funderIdentifier":"https://ror.org/00whkrf32","awardTitle":"Graduate Research Fellowship Program (GRFP)","awardNumber":"2141416"}],"url":"https://datadryad.org/dataset/doi:10.5061/dryad.r4xgxd2v7","contentUrl":null,"metadataVersion":0,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-08-20T10:32:30Z","registered":"2026-08-20T10:32:31Z","published":null,"updated":"2026-08-20T10:32:31Z"},"relationships":{"client":{"data":{"id":"dryad.dryad","type":"clients"}}}}],"meta":{"total":3770,"totalPages":151,"page":1},"links":{"self":"https://api.datacite.org/dois?client-id=dryad.dryad\u0026registered=2026","next":"https://api.datacite.org/dois?client-id=dryad.dryad\u0026page%5Bnumber%5D=2\u0026page%5Bsize%5D=25\u0026registered=2026"}}