Data Package Design for Special Cases
Version 1.1
published
Community-developed considerations for creating well-designed datasets that include data specialized by type, format, or acquisition method. Examples are images, code, documents, and raw data in other repositories.
Abstract
This book covers considerations for creating well-designed datasets for publishing data types outside the most common cases for research data. Specialized needs for preparing and publishing a dataset may arise because the data are not tabular in structure, are stored in non-standard file formats, are collected with unique or new acquisition methods, or are large in size. Some examples of specialized data types described here are imagery, code, document archives, spatial data files, and raw genomic data stored in other repositories. This is the first version of the book, written and edited in 2019 through 2021 by the Non-tabular Data Working Group of the U.S. LTER Network Information Management Committee. Minor edits and reformatting took place in 2021, 2022 and 2024 to produce this current online version (Version 1.1).
Keywords
data management, EML, dataset, imagery, big data, documents, drone, code, research, environmental science, metadata, publishing, guide, repository
Reuse
Citation
BibTeX citation:
@book{gries2021,
author = {Gries, Corinna and Beaulieu, Stace and F. Brown, Renée and
Elmendorf, Sarah and Garritt, Hap and Gastil-Buhl, Gastil and Hseih,
Hsun-Yi and Kui, Li and Martin, Mary and E. Maurer, Gregory and T.
Nguyen, An and H. Porter, John and Sapp, Adam and Servilla, Mark and
Whiteaker, Tim and Network Information Management Committee, LTER},
title = {Data {Package} {Design} for {Special} {Cases}},
version = {1.1},
date = {2021-02-19},
url = {https://prerelease-edi-docs.netlify.app/guide-special-cases/},
doi = {10.6073/pasta/9d4c803578c3fbcb45fc23f13124d052},
langid = {en},
abstract = {This book covers considerations for creating well-designed
datasets for publishing data types outside the most common cases for
research data. Specialized needs for preparing and publishing a
dataset may arise because the data are not tabular in structure, are
stored in non-standard file formats, are collected with unique or
new acquisition methods, or are large in size. Some examples of
specialized data types described here are imagery, code, document
archives, spatial data files, and raw genomic data stored in other
repositories. This is the first version of the book, written and
edited in 2019 through 2021 by the Non-tabular Data Working Group of
the U.S. LTER Network Information Management Committee. Minor edits
and reformatting took place in 2021, 2022 and 2024 to produce this
current online version (Version 1.1).}
}
For attribution, please cite this work as:
Gries, Corinna, Stace Beaulieu, Renée F. Brown, et al. 2021. Data
Package Design for Special Cases. In Environmental Data
Initiative, version 1.1. https://doi.org/10.6073/pasta/9d4c803578c3fbcb45fc23f13124d052.