Leveraging an observed-data likelihood improves the use of machine learning labels in a Bayesian hierarchical model for bioacoustic data

dc.contributor.authorOram, Jacob K.
dc.contributor.authorBanner, Katharine M.
dc.contributor.authorStratton, Christian
dc.contributor.authorHoegh, Andrew
dc.contributor.authorIrvine, Kathryn M.
dc.date.accessioned2026-07-30T21:49:46Z
dc.date.issued2025-12
dc.description.abstractClassification of massive datasets by machine learning (ML) algorithms is promising for many scientific domains, especially wildlife monitoring programs that rely on passive acoustic surveys for detecting species. However, treating ML-predicted class labels (e.g., species identity) as truth biases inferences of focal parameters within common modeling frameworks. One solution is to model the misclassification process explicitly using human-validated true-class labels for a subset of observations. Validation by experts can present a substantial bottleneck in otherwise efficient workflows that use ML predictions. Bioacoustics practitioners seek guidance on both the quantity and process for selecting ML-labeled data to validate by an expert. We derive an alternative model formulation that jointly models human-validated and ML-predicted class labels with an observed-data likelihood (ODL) and use empirically informed simulations motivated by a real-data application to explore different probability designs for selecting class labels for validation. Simulation results suggest that with smaller validation sets the ODL formulation increases computational speed and reduces estimation error compared to a default MCMC data augmentation routine. Our methodology is transferable to applications that treat predictions from classification algorithms as the response variable of interest.
dc.identifier.doi10.1214/25-AOAS2096
dc.identifier.issn1932-6157
dc.identifier.urihttps://scholarworks.montana.edu/handle/1/20074
dc.language.isoen_US
dc.publisherInstitute of Mathematical Statistics
dc.rightsFind the version of record at https://doi.org/10.1214/25-AOAS2096
dc.rights.urihttps://web.archive.org/web/20200526231745/https://www.imstat.org/wp-content/uploads/import/copyrightTA.pdf
dc.subjectmachine learning algorithms
dc.subjectBayesian hierarchical model
dc.subjectbioacoustic data
dc.titleLeveraging an observed-data likelihood improves the use of machine learning labels in a Bayesian hierarchical model for bioacoustic data
dc.typeArticle
mus.citation.extentfirstpage1
mus.citation.extentlastpage24
mus.citation.issue4
mus.citation.journaltitleThe Annals of Applied Statistics
mus.citation.volume19
mus.relation.collegeCollege of Letters & Science
mus.relation.departmentMathematical Sciences
mus.relation.universityMontana State University - Bozeman

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
oram-bayesian-hierarchical-model-bioacoustic-data-2025.pdf
Size:
6.39 MB
Format:
Adobe Portable Document Format