Structure
Organize the selected source data into consistent tables and fields. Standardize timestamps, identifiers, units, location data, and missing-value conventions in the finished dataset.
Dataset Construction
For researchers who want the dataset—not the work of constructing it.
Study data is often too fragmented and inconsistent to analyze as collected. Preparing it takes time and specialist judgment about how the sources relate and where they disagree.
I handle that work efficiently and deliver a quality-assessed, documented dataset within an agreed schedule.
Receiver logs vary by format and software version; study metadata varies in completeness and quality across seasons and contributors. Bringing them together takes time and specialist judgment to align fields, identifiers, and relationships within a consistent structure—so you can work without navigating each source’s complexity.
Years of constructing datasets for acoustic telemetry positioning—where small errors in timestamps or station positions can produce much larger errors in estimated animal positions—have given me a methodical approach and the informed judgment this work demands; see my work history.
Original study data stays unaltered; I work from copies and run checks throughout construction. You confirm decisions that affect how the source data are interpreted or whether the finished dataset supports its intended use. Analysis and scientific interpretation are separately scoped.
Organize the selected source data into consistent tables and fields. Standardize timestamps, identifiers, units, location data, and missing-value conventions in the finished dataset.
Join the structured source data using explicit relationships among identifiers, deployment periods, and locations. Link receiver detections to station deployments, animal tagging records, and available device metadata.
Find duplicates, gaps, mismatches, anomalies, and timestamp or identifier conflicts. Compare detections with available receiver diagnostics, make corrections supported by the source data, and mark uncertain cases as unresolved.
Record the source data and versions used, processing steps, check results, assumptions, decisions, unresolved issues, and known limitations. Include that documentation with the delivered dataset.
I build the dataset in the agreed format, schema, and table structure, or recommend open formats suited to R or Python. For hundreds of millions of detections, I can use a storage layout that lets you read selected rows or columns without loading the full dataset into memory. Sample query code is included when agreed.
You receive a versioned dataset in the agreed format and structure, with clear documentation of its contents, construction, checks, and known limitations.
Tell me about the study data you have and the dataset you need. I’ll arrange the transfer, then review the study data you provide.