Episode 15: Bioinformatics war stories
In this episode of the microbinfie podcast, hosts share candid 'war stories' from bioinformatics research, revealing the complex and sometimes unpredictable challenges in genomic data analysis and laboratory work.
We share short war stories about bioinformatics from the trenches of research. We share tales about a mystery massive Listeria outbreak, bubbles on flowcells, GAII woes, contaminated databases, the mass murderer, water pouring into a sequencing centre, changing protocols without validation, Excel issues and the ultimate complexity in science - Humans.
Extra notes
-
Contamination Challenges:
- There was a significant outbreak linked to Listeria monocytogenes contamination due to contaminated laboratory culture media, highlighting the need for vigilance beyond raw data acquisition.
- An episode involving contamination of Salmonella detections due to genomic similarities with Shigella emphasizes the complexities in distinguishing closely related bacterial species in bioinformatics analyses.
-
Technological Artifacts:
- Issues with early sequencing technologies such as consistent adapter sequences causing sequencing errors were discussed. This underscores the importance of introducing complexity in samples to avoid such errors.
- The "bubble controversy" with flow cells causing image distortions is an example of technical imperfections impacting data quality, which the manufacturer eventually addressed.
-
Data Verification and Trust:
- Common pitfalls such as over-reliance on database accuracy can lead to misidentifications, as seen in cases where bacterial data matched unrelated organisms due to errors in previous sequencing submissions.
-
Importance of Metadata:
- Highlighted was the catastrophic impact of mixing metadata, where a mis-sorted Excel sheet led to irreparably shuffled metadata, compromising the integrity of the genomic data set.
-
Human Error and Its Effects:
- Human errors, like an alteration in the experimental protocol without proper documentation, were discussed as a significant cause of data corruption or misinterpretation.
- Suggestions were made for instituting positive control mechanisms to detect and mitigate such errors.
-
Tool Usage and Data Handling:
- Examples discussed the crucial role of using tools like Kraken correctly and emphasized thorough examination of genomic regions using integrated genome viewers (IGVs) to accurately diagnose and resolve data inconsistencies.
-
Lessons from Situations:
- Regularly auditing and verifying alignments and annotations during sequence analysis was recommended to avoid misleading conclusions, by ensuring sequence data corresponds correctly with classified organisms in databases.
-
Design Considerations:
- There was a cautionary note about building sequencing facilities on floodplains, learning from past damages to expensive sequencing equipment at the Sanger Institute.