Hello, and thank you for listening to the Microbinfeed podcast. Here we will be discussing topics in microbial bioinformatics. We hope that we can give you some insights, tips, and tricks along the way. There is so much information we all know from working in the field, but nobody writes it down. There is no manual, and it's assumed you'll pick it up. We hope to fill in a few of these gaps. My co-hosts are Dr. Nabil Ali Khan and Dr. Andrew Page. I am Dr. Lee Katz. Both Andrew and Nabil work in the Quadram Institute in Norwich, UK, where they work on microbes in food and the impact on human health. I work at Centers for Disease Control and Prevention and am an adjunct member at the University of Georgia in the U.S. Hello and welcome to the Microbial Bioinformatics podcast. Today we're talking about getting your head around our favorite enteric microbes, E. coli and Salmonella. Why do they have some of the names that they have? Let's shoot from the hip. First question comes from Andrew. Right. So I'm a computer scientist, so I can feign ignorance of all of this. Lee, you work in the enteric branch, so what on earth is an enteric and why should I be concerned? Okay. Enteric. Enteric means it's in your gut. So just going back to like basic anatomy, you have like your whole person, right? You have places in your body where bacteria don't belong, like inside your muscles, inside your skeleton, inside whatever, but bacteria actually do belong in the area going from inside of you from your mouth all the way down to your butt. So bacteria are enteric if they are somewhere in you from your stomach to your intestines to wherever in between. Does that mean that like Clostroides is enteric? Yeah. So you can accidentally eat Clostroides and then it will be inside of your gut and it will cause enteric disease. So enteric doesn't necessarily mean like it's naturally there. It can also mean it's an enteric pathogen, which means once it's found in your gut, it will cause problems. So now there's a new activity. What can we eat to... So now there's a new activity, you know, what can we eat to make the enteric branch have to look at more pathogens? So SARS-CoV-2, right? Coronavirus. That's found in stool. Is that an enteric? Oh, good one. Yeah, you can't just like eat a thing and call it enteric. Shedding something is a different matter, I think. And there are plenty of animals out there that shed what we would consider enteric pathogens that do nothing to the animal. It's just more about the transmission. You could eat a thing, but your stomach acids might destroy it. So you can't just like eat dirt and call the dirt enteric. Your gut will mash that down and it won't cause any problems. And you can eat a virus that causes a respiratory disease, but it'll get broken down your stomach and it won't cause anything. There are lots of studies about that stuff. So I like to say you can't eat COVID actually, because you can eat SARS-CoV-2, but you can't eat the disease. I mean, that's a deep philosophical question for me. What is a pathogen? What is enteric? Where do you draw the line? Yeah. Okay. So let's move on to, say, E. coli, right? The names with E. coli always confuse the hell out of me now. I think I'm on a paper, an E-Tech paper, right? And Leah's one as well. Nabil, you're left out, unfortunately, sorry. But I always get confused about E-Tech and E-Pec and S-Tech and E-everything. What on earth do they mean? Do they mean anything? And is it what they look like in the lab? Is it a disease they cause? Is it where they're found? What do all these E-things mean? And are they even E's? There's E-Pec, I know, as well, like an S-Tech, and God, I'm confused already. Nabil, can you explain? I'll jump in with some history and explain why this makes no sense. The reason is that these terms existed before we knew anything about Bacillus coli, or what we call E. coli. So people recognized that there were microbes that cause meningitis and neonates. They recognized that there were microbes in form, in terms of, you know, pathogens, and they were microbes that cause enteric disease. And the fun part was Escherich, who said, hang on a second, these actually might just be the same guy. And that's, and called that Bacillus coli, which we call it Escherichia coli, in his honor. So these terminologies are all based around clinical representation that we've then sort of adopted to talk about a microbe, or different types of microbes. Well, it's kind of the same thing, but not really the same thing. And how are they related to genomics, now that we have sequenced probably lots and lots of E. coli? Well, it comes back to your first thing about, you know, is an enteric pathogen, like, you know, because you can take a common E. coli that lives in your gut, and it can cause a urinary tract infection. Is it a, is it a, you're a pathogen then? Yes. But it doesn't have the hallmarks, the genomic hallmarks of what you expect a uropathogenic E. coli to have. So you probably wouldn't say it's like a canonical one. Mostly the difference between, they are, it depends on which type you're looking at. Uropathogenic E. coli, the phylogroup that's commonly associated with that is very different to those that you'd find for enteric E. coli. But then, yeah, you can have enteric E. coli going in the wrong place and causing problems. So it's sort of like, well, yeah. Similar definitions like Shiga-toxigenic E. coli or Verotoxigenic E. coli, pretty much nowadays is just the fact that it expresses Shiga toxin. So and Shiga toxin being mobilized by a prophage, by a phage, as can be just be found in a lot of different E. coli all over the place. So that's driven by sort of something else, mobile genetic elements, which isn't tied to a specific clade or anything like that. So how does that relate to Shigella? Is Shiga toxin in Shigella the same as your STEC? One of them is. And I know Shigella is just like a plasmid, it's a E. coli with a plasmid and it shouldn't really have its own name, but there you go. Yeah. So again, Shigella is, it's again the same story where Shigella was something that was classified because it caused dysentery. So dysentery is like named for that fact. And then people put the dots together and realize, hang on a second, this Shigella thing is more or less a subset of, of what we call E. coli. It is very much this problem that we've sort of gone ass backwards and we've talked a lot about the clinical presentation of these and then realizing later that actually that doesn't necessarily line up with the, with the genealogy of the microbe. So can you tell from a genome, like if I give you an E. coli genome, can you tell which one of these it is? Is it EPEC? Is it ETEC? Or EPEC? Oh yeah. I mean, I tried a lot during my doctorate to try and make a categorical list. It's a little tricky, usually for things like EPEC and terapathogenic E. coli, they have to create attaching and effacing lesions. So that's where the, and that is only really possible by having type 3 secretion. So you can just look for a locus of enterocyte effacement, which encodes type 3 secretion, and then you're probably going to be looking at an EPEC. Shigetoxigenic, you find Shigetoxin, you're done. Other guys are a little bit, so the difference between say enterohemorrhagic E. coli and Shigetoxigenic E. coli is a little bit more tricky because again, enterohemorrhagic E. coli is sort of more about the, the clinical presentation rather than just it having a particular genomic feature. You can find, some of them are a little bit more tricky. So the difference between canonical uropathogens to enteric pathogens is hard, but normally you can find certain genomic islands that encode things that help, help an E. coli live in an environment that's very, very like sort of nutrient deficient. So often they have a lot of things relating to, to acquiring iron, iron acquisition, because that's actually quite difficult for an E. coli to do in, in, in the urinary tract. So often they have a lot of genes to help them out with that. A lot of things with like lysing cell, other cells, host cells to get, get, get access to nutrients and things like that, which the enteric ones probably won't have. So what I've always found very curious is that E. coli has a massive pattern genome. And if you look at say any two E. coli, they're going to be usually quite different if you randomly take two. And is that because of what you say, it is so good at adapting to different environments that it can take up different pieces of DNA maybe that it needs? So for me, I think if you take away the usual suspect that we worry about, so you're 015787, you know, the guy in the news, or you take away your ST131, so you take away. these are quite clonal pathogens that have rapidly expanded reasonably recently. And they're very static. If you take normal commensal E. coli and you compare them amongst each other, they are recombining like it's nobody's business. Maybe not on the same level of like Campy or some of the others out there, but they are shunting stuff all over the place. And that allows them to rapidly sort of disseminate a lot of genetic information and also create a lot of diversity. So really the, are you saying kind of the commensals that are maybe in our guts are the ones that are like the harbingers of doom. They're the ones in the background shuffling around all the dangerous stuff that E. coli or the pathogen might need just in case like it's the backup arsenal. Honestly, I have no idea. That is like a thing that I'm always fascinated with and I'm always trying to get my bearings around. I mean, obviously if you take the, you take the like the E-tech stuff, you take the heat label and heat stable toxins out of it. Like, yeah, okay, that's very clear, cut, cut and dry. But if you think about a problem like the microbe has to adhere to host epithelium, how is those systems passed around? That's not clear to me, actually. It does seem to be stratified, but then it also seems to be just juggled as any old thing. A lot of what we see in that regard, can be pulled in from commensals and sort of remixed. And I mean, I think the classic example everyone remembers in E. coli for like in recent terms for this remixing is obviously the European 104 outbreak from what, like 10 years now, I think. So, I mean, that one was, that wasn't a commensal, but that was an entire aggregate of E. coli that has a completely different pathogenesis that picked up a sugar toxin and became something else. So S-tech plus EA became that. So that remixing thing, and we got to see that firsthand of that. And that's the kind of thing E. coli seems to like to do, but probably quite quietly most of the time, we don't notice it or we haven't dug into it that much. So how does MLST, like if you take 7G and MLST, how does that relate to all of these different names? No, it doesn't for the most part. I'll just be, yeah. I wouldn't take the ST. The STs for things that are very clear cut, like a one, five, seven, yeah, that's fixed. But if you pick a random bug, then yeah, your ST is not going to be necessarily a good predictor of its pathology. We like to shoehorn all of our phylogenetic methods with all of these pathogenic methods, but these pathogen types, but that doesn't work, does it? No, I mean, no, it doesn't. It doesn't work. MLST is not like a good predictor for that. And plus the other bonus, I mean, this is taking the host angle out of it as well. That introduces its own thing. Like you have so many publications upon publication that talk about sugar-toxigenic E. coli, for instance, I'm sure with the E-tech as well, that it's in a person, the person's healthy, nothing happens, and then the next person has got like serious systemic sequelae. Like it's just, and it doesn't make any sense. You're like, why, what happened? So that angle out of it, the genotype doesn't necessarily line up with the pathotype very well. Okay, so this has always confused me, but Salmonella is quite closely related to E. coli. You know, there's what, 150 million years between them, but it seems to operate very differently than because you've got these cerevars and whatnot, and they seem to be much more stable compared to E. coli. Is it really though? I mean, if you look at the literature, they'll always point out that there's over 2,500 cerevars, so sero, like antigenic types of Salmonella are out there. That's a lot of variation. I think if you do like a sort of ANI measurement, there's E. coli might be more diverse than Salmonella, but it's within the same ballpark, I think, if you sort of counted the number pairwise, the number of SNPs between one side of the population to the other. But I guess within a cerevar, you have, it does seem to be quite well conserved, say on a Salmonella typhi, you know, you're talking about 95% at least similarity or vastly greater, actually, probably 97% easily. That for me, I think is the magic part of Salmonella is when you get the top 20 cerevars that show up in the clinic, and you look at those, they all have exactly what you're saying. They have these very, very delineated, very stratified streams. Like you make them on a tree or whatever, at some point in the other, all of these just decided to diversify and then just stay put. And then when you look at it on a clinical perspective, you go, well, Salmonella is quite samey. I mean, there's all these separate, very clear buckets of Dublin and typhi and typhimurium and the paratyphies. And then you sort of go, oh, well, they're all this very like static, boring, boring bug. But then when you step out of those, and then you start talking about the guys who are sort of on the fringes, the ones you find in rivers, the ones you find in reptiles, the ones you find in like all over the place, they start getting very messy and they start looking more like E. coli in terms of this sort of moving, shuffling, all sorts of genomic stuff around. There's a whole thing of research from the people behind ANI on the continuity of species like that. It's kind of interesting to me. Costas Konstatidis from Georgia Tech has this whole thing about ANI and he can find two different kinds of E. coli. And I've seen this slide deck a few different times that he can find like the next closest bacteria and the next closest and keep sequencing and get closer and closer and closer with ANI and where's the species delineation. And he goes really well into depth on that. I'm not gonna go into depth on that because I cannot represent him, but I think it's an interesting line of research. Yeah. The odd thing with salmonella is it doesn't, there seems to have been an event at some point where a lot of these cerevals just sort of rapidly expanded and then stayed put. And even though we, no matter how much we sequence, we can't seem to pin down, we don't see this graduation so much in salmonella. So E. coli, it's more fuzzy, this graduation. Is that because we're so focused on pathogens and not on commensals and we're just kind of missing huge branches of the tree of life of salmonella because 99% is all on like say typhi, typhimurium and virtually nothing is on all these weird and wonderful things that you find in soil and in random animals. Yeah, I think there's a horrible bias of salmonella genomics, obviously the bulk of it being typhimurium, enteritidis and then I think typhi after that. And they're making up between those three cerevals that's making up like maybe two thirds of all of the salmonella genomes out there. So that sample bias is obviously a problem, but I think also we can't see that far back into the past because all of the missing links are gone. They're dead, they're all disappeared. And we just see these sort of branches, these offshoots of these now reasonably stable cerevals and we can no longer impute what was the chain of events that caused them to appear in the first place. Or we need to go and do a lot more sequencing. We can, I'm all for more sequencing. Except stamp collecting in weird and wonderful places. Can we talk about Nabil's viral tweet? Speaking of stamp collecting. Yeah, I know. That was funny. What did you want to specifically mention on that? I don't even know. I just thought it was hilarious. Well, we'll put it in the show notes. Let's want to check out. You had on there though, I think one of them was stamp collecting. Yeah, so this is a knockoff of an XKCD comic which talks about the types of scientific papers. And I made a specialized, just rewrote over it to write a version that says microbial bioinformatics papers. And one of the types of papers you always see is we sequenced a bunch of stuff, but it's not stamp collecting. We made up the biology after we sequenced it. Is that on there too? No, I didn't have any. That's a good caption. No, the closest one I had to that was we ran BEAST for 92 days and then we chose the relaxed clock model because the tree was pretty. Someone who I will not name instant messaged me today and said that it hit too close to home to her. Most of those headings are inspired by stuff I've done. So I'm picking the Mickey out more out of myself than anyone else. It's not specifically targeting anyone. Thank you. I accidentally rewrote bed tools with a pro. I used a pro module that basically does an array in span or something. And it basically does what bed tools does. And I like rewrote bed tools, I think. I don't know which is worse, rewriting bed tools or rewriting more MLSD folders. I've done that too. Yep, yep, yep, yep. They're both bad. Yeah, yeah. I reviewed a few of those papers actually, because of our MLST caller paper, you know, they keep sending them to me and I keep going, Jesus, do I really want to allow the world to have yet another MLST caller? It's all right. It's either that or they'll start writing SARS-CoV-2 pipelines or something. Oh, we all need those. Yeah. We need another one. Yeah. I'm guilty. Okay. I guess we don't have to go too deep into it. It was just really funny. Props to you. Any more questions? You're stuck with Andrew, with E. coli and Salmonella? I don't know. I just give up. I think we should just long read sequence everything, every Salmonella and E. coli we can get our hands on, and that'll solve everything. It's not stamp collecting. It's like high quality premium stamp collecting. Yeah. Yeah. But I think we're never going to get rid of the naming of Salmonella serifase as a thing. No. I suppose we need more informatics methods to be able to replace the lab methods, because people aren't going to give up the lab methods, the phenotyping methods, until they can actually replicate those in silico. It's getting there. I think there's probably, PHE's put off enough publications of it, and then the Canadian public health have also put a fair few that make the case pretty strongly that you can get the same information from the genome rather than doing serotyping. Yeah. There's E.C. Typer, and Seq0, and Sistr. What are the other guys? I mean, those are the main ones. Those are the ones people should be using. From John Nash's group, is it? The Sistr, yeah. And I think it's Deng for the Seq0. Yeah. It's Deng. You know, the disclaimer is that I work with him. I'm an adjunct professor with Deng. But if you're a thing who decides what the serovar names are, it's a cheap buster. Yeah. What happened? For a while, they were giving them like crazy names. It were colons and God knows what. They were like 30 characters long. They got rid of like Dublin and simple things like that. What happened? So what happened was, I can't remember when they decided it, but they decided that serovars, the thing is, is antigenic formulas do not describe salmonella very well outside of subspecies one out of S. enterica enterica. So they decided like, we're not going to, and there's too many of them. There's too many random combinations of these antigenic formulas. And then when you look at them, they don't have anything, the ones that share it don't have anything in common. So the designations are meaningless. So they decided not to name those outside of any of those out of a side of subspecies one. So you're supposed to refer to, I mean, there are people who still refer to them with some of them did get named and people still use those names, but I think you're supposed to use the antigenic formula, which describes the, yeah, the antigens that, that, that, you know, the microbe reacts with. So that's as bad as a computer scientist saying, yeah, please use the MD5 hash as the, the antigen. Please. Basically. Yeah, it is. It's not a very, it's not a very informative thing, but the thing is, is like, you're not really supposed to use that. The serology doesn't help you. It's like, it's like, stop using this. We're going to make it more difficult because we don't want you to use it. It's not going to help you trust us. We all like a good name. It's a lot easier to talk to people and say, listen, now, you know, you, you've got the Salmonella Kentucky, you know, you didn't get into Kentucky fried chicken, but you know, you definitely have that server. Yeah. You can make a game and see if your home city or home country has a name, has a server named after it. It's, it's quite amazing how many different places are in there. Oh yeah. Like Dublin represented. It's not one associated with cows. Yeah. I think, is there an Atlanta? I don't know. I know that all the people around me are going to say, Oh yeah, Oh, there was an Atlanta. It's called Mississippi now. Sorry. Yeah. Yeah. There's a Mississippi. I don't think there's one after my home. No, there's not. Oh yeah. There's one Brisbane. So my hometown gets represented as well. So that's good. Awesome. This is simply to represent where these were isolated. It doesn't mean these originated from there or anything. It's the same problem as the SARS-CoV stuff. Yeah. But then more complicated, you get like typhi, paratyphi, paratyphi, A, B, C, what on earth has gone on there? Are they different? What are they? What's the difference between A, B, C and D? Are they very closely related genomically or are they quite random? This is, this is the classic like software engineering problem of sort of making it up as you're going along. So when they designed the scheme, so this is Kaufman and White back in the turn of middle of last century, they looked at astrology like, yeah, it makes sense. And then they tried to come up with names that were descriptive, that describe the type of disease it caused, you know? So cholera sewage is like, okay, it causes cholera in pigs. It kind of doesn't really, but there you go. And then typhimurium and typhi in mice and so on. And the paratyphi. So the paratyphi is like, it causes typhoid fever-like symptoms, but it's not typhus. It's not really typhoid fever, so paratyphi. So, okay. And then you have A, B and C and so on. But, but then you kind of run out of that pretty quickly. And so then they're like, well, now what do we do? So they picked another system and so they go, oh, we'll just name it after the place that it's from some of the time, or we'll name it after, oh, the abortus sequi is another one that explains what it does, how, what kind of disease it presents, causes abortion in horses. I used to not name them after people because I saw like in a, in anopheles that, and plasmodium often they will name them after like particular people who lived a hundred years ago, like Bill Collins, Collins eye and stuff like that. And it's like, hmm, yeah. Do you really want to be associated with like this deadly pathogen or vector? I might be wrong, but the only reason they probably didn't do that was because that kind of naming is for species. And they don't want to, they, they, they're very careful with this, with that. There's a lot of problems with this because of, of having cerevis considered as a species because they're not. I did notice that people working in salmon are like a very annoyed if you do italics for the server name, they got, no, it's a big, big, no, no, it's a big no, no, because it's not, it isn't a species or subspecies. And then the problem was is that they used to do that. So it used to be as color as Sue is like all italics and it's like, it's not, then they demoted it down. Colossus down to just being a Sarah of our term. Yeah. No, but that's a bugbear for me. So I also get annoyed when the, the nomenclature isn't right because I have a whole document somewhere from people I worked with the CDC on how to name salmonella appropriately. And I lost it. Yeah. There's a CDC paper. There's a CDC white paper that just explains. This is what you're supposed to do, which I, which I come back to sometimes when I'm writing up just to make sure I haven't accidentally gotten mixed up again. It's complicated. It's complicated. You got it. It's, it's, it's, it's horrible. It's the world. It's these really, really long names, salmonella enterica, subspecies enterica, Sarah of our type in Miriam. And it's just like, okay, the S enterica is italicized, the subspecies is not the second enteric is italicized. The Sarah of our type in Miriam is not like you just get, you, you get some, you get like RSI, just, just trying to use your, just trying to use the hotkeys of the italics ground. Well, I'm happy with that. Okay. So do we call it a day there? Yeah. Yeah. And on that bombshell, thank you for listening to the salmonella, zero of our micro bin feed podcast. We've been talking about E. coli and salmonella and all the wonderful naming schemes that have been created as we grapple with these enteric pathogens. Thank you so much for listening to us at home. If you like this podcast, please subscribe and rate us on iTunes, Spotify, SoundCloud, or the platform of your choice. Follow us on Twitter at micro bin feed. And if you don't like this podcast, please don't do anything. This podcast was recorded by the microbial bioinformatics group. The opinions expressed here are our own and do not necessarily reflect the views of CDC or the Quadram Institute.