Hello, and thank you for listening to the MicroBinfeed podcast. Here we will be discussing topics in microbial bioinformatics. We hope that we can give you some insights, tips, and tricks along the way. There's so much information we all know from working in the field, but nobody writes it down. There is no manual, and it's assumed you'll pick it up. We hope to fill in a few of these gaps. My co-hosts are Dr. Nabil Ali Khan and Dr. Andrew Page. I am Dr. Lee Katz. Andrew and Nabil work in the Quadram Institute in Norwich, UK, where they work on microbes in food and the impact on human health. I work at Centers for Disease Control and Prevention and am an adjunct member at the University of Georgia in the U.S. Several years ago, when we were doing the Listeria real-time WGS project, I mean, we're past the pilot. We're still going strong on that. In the first few years when we were starting off, we started noticing that there was a widespread outbreak of Listeria, like bigger than anything ever before, and it was just across several states. It turned out to be that we did have contamination from the sheep blood. There was a Listeria outbreak in sheep, and we had no idea. As we were growing Listeria on the blood agar, we were also apparently getting Listeria from the plate itself. Oh, God. Oh, that's terrible. And I need to look that up and actually put that in the show notes with you guys. It's like that mass murderer, I think it was in Germany. Yeah, I totally remember that one. The mass murderer. I just happened to be the person who was making the swabs that the police were using and was licking our fingers. But yeah, when they did the DNA test, they just kept seeing the same signature over and over again for all these murders all over the country, and it didn't make any sense. Those are ridiculous examples. I love those war stories. But I suppose it does highlight the importance of a bit of common sense and not just blindly trusting the sequencing data that you get back. I found the paper. I'm sorry to interrupt you. The paper is called Two Listeria Monocytogenes Pseudo- Outbreaks Caused by Contaminated Laboratory Culture and Media, and that was published in the Journal of Clinical Microbiology. Awesome story if you want to go up and go read on it. Thanks. Do you have any other war stories you want to talk about? Well, I remember in the early days of TRADIS, that's transposon mutagenesis, that on the GA2, so that, you know, it's a long time ago. Oh, wow. Yeah. They would, they'd add the same adapter to the very front of every piece of DNA. But the problem was that when you sequenced it, when you look at the picture from Illumina, every cluster had the same color at the same time. So it was either all black or all white. And then the sequencer just kind of exploded, you know, and just fell over. So one thing to be aware of is that mixed samples and some kind of complexity is a good idea for sequencing. Are you saying that the sequencer just couldn't handle incorporating the same base at the same time for a single cycle? Yep. So even every single base in one cycle was the same? I would freak, yeah, I would freak out, I would freak out. And it's because, I suppose, the sequencer uses the first few bases, 15 bases to do templating to, you know, kind of warm itself up. And then there was the bubble controversy as well. So for ages, this is many years ago now, there's a little bubble would appear on the flow cell and that would cause the camera to go out of focus. And of course, then you get, you know, worse data. And for ages, the manufacturer, you know, well, didn't admit it was there. And eventually they came up with a fix for that. But, you know, it was just someone using a bit of common sense saying, this doesn't look right, you know, it's the same tile having the same problem every time, you know, what's up here. And then someone eyeballing and going, oh yeah, there's a bubble. Yeah, don't forget in your FASTQ files, that information is usually written in the header. That's what all that gobbledygook at the top is. It's got the machine and the tile, the coordinates. And so you can, for this sort of problem, if you ever worried about it, you can go back and reconstruct it. Are you saying that if there is a bubble, then the coordinates would show that the spots are just not centered around one place or another? Yeah, it's sort of like if you average the quality over all of the sector for each sector, it should drop off dramatically for the ones that have the bubble. I have, this isn't really my worst story, but when I was doing the Haiti cholera outbreak, I was looking at previous samples and someone just at that time, like it was so early, they just introduced this idea of using a metagenomics database to see if every single sample matched cholera only. And I did see some previous results that were, that we did report in our manuscript that were contaminated. And it was just really shocking that they were included in previous studies to ours. They were just like 50% contaminated or, you know, even up to like 25% contaminated with just random things that were just not even related to Vibrio. I know I've seen some cases where you sequence some salmonella and it comes back as like a mountain beetle. And what's happened is, you know, someone's mashed up a beetle and sequenced it, forgetting that there's bacteria in the gut of the bug. Good point. And of course, once that gets into a database, it's hard to get it out again. Yeah, I've had something similar with that with pseudomonas and camels. This was a long time ago. What, they mash up the camels? No, no, it was just, there was, I don't know how it got in there. I have one more war story that, again, this isn't really mine, but from our team that just using standard Kraken for a few different runs in a row, we had some samples that kept looking like salmonella, but they were Shigella. And we found out that there were like two salmonella genomes in the standard database that basically shared regions, shared with Shigella. And we looked at IGV later on, we found that there was tremendous disparity between regions that were covered or not covered. And it took a very long time to nail that down. Yeah, well, I remember one time, the big project, you know, a thousand samples, and the first few plates were good. And then suddenly loads and loads of samples are kind of mixed strains and took a while to figure it out. Actually, someone in the lab had made an executive decision to change the fundamental protocol. And of course, that didn't work. So they were, they decided instead of, you know, doing single colony picks and culturing stuff up, they would go, okay, there's one bead, it's probably got one bug on it, I'll just take straight from the beads and sequence and save myself, you know, many, many hours of time. And no. No, no. They didn't tell anyone. And it was only in the data then, you know, when it took, you know, quite a lot of work and quite a lot of sequencing to figure this out. And quite a lot of wasted money, because then people had to go back and do it again from scratch. Yeah, I think, I think human contamination is probably the worst thing for me. And that's where a human being introduces something awful into your data set. This is the worst, the worst one was someone missorting an Excel sheet with all the metadata. And then that's important. And it wasn't like reverse, like front to back, it was just shuffled in a way that could not be reconstructed. So now you have a paper, you have a, you have the genome, the data was fine, but you have all your contextual information is buggered. You have no idea what year this is, what country it is, you just, you can't trust anything on it. And you could, yeah, we wound up reconstructing some of it, but there was a bunch of samples we just had to throw out because we didn't know what they were. And just because someone, someone clicked the wrong button somewhere. Yeah, I suppose there's other things that are out of control as well, you know, when you get samples shipped in, and maybe they're meant to be kept refrigerated, and you know, people don't do that. Or customs decide, oh, I'll take off, you know, the lids or the seal, see what's in there. Ah, just shove it back on again. And then of course, it's all destroyed or contaminated or messed up totally. Yeah. And that's definitely where you want that secret positive control to go through for those for exactly these kind of reasons. Because yeah, if are you getting sort of just as a check, are you getting what you think you're getting on the other side? Because a lot of these, it's not uncommon to have samples sitting out in the heat by accident. I'm really curious about this one. Sanger sequencing room flooding. What were you, what's your note on with that one? The Sanger Institute was built in a floodplain for the River Cam. And of course, the sequencing facility was built, you know, where the water will come in on the ground floor. So, you know, the water is rising and it was at that point that they realized that maybe the builders hadn't built the building as well as they should have and there's some leaks and holes and water was coming in. And so they had all these really expensive, would have been the ABI sequencers back in the day. And so carefully calibrated big heavy machines and they had to, you know, carry them as quickly as possible up the stairs as the water was coming in. And from then on, the sequencing facility was on the upper floor. So don't build a sequencing facility on a floodplain. I think that's the message of the day. I think that these were really great stories, and if any of the listeners here have any good war stories, we'd love to hear them in the comments too. Thank you. Thank you all so much for listening to us at home. If you like this podcast, please subscribe and like us on iTunes, Spotify, SoundCloud, or the platform of your choice. And if you don't like this podcast, please don't do anything. This podcast was recorded by the Microbial Bioinformatics Group and edited by Nick Waters. The opinions expressed here are our own and do not necessarily reflect the views of CDC or the Quadrant Institute.