Transformer-Diffusion model for molecular battery material generation
See how a Transformer-Diffusion model generates novel battery materials autonomously, bypassing DFT calculations for rapid, physically viable designs.
Overview
I built Simer Energy, an end-to-end generative AI pipeline that uses a hybrid Transformer-Diffusion architecture to autonomously design, physically relax and thermodynamically validate novel battery materials from scratch.
For the demo, I will execute a terminal-based run of the by inputting strict material constraints (e.g., elemental bounds for a cobalt-free transition metal oxide, target stoichiometry, and symmetry parameters) into a fine-tuned model. I’ll show how the Transformer maps these constraints into a discretespace groups and passes it as a conditioning vector to a diffusion model. You will see the model denoise the 3D spatial atomic coordinates, followed by the critical “zero-strain” and then passing the raw generated structure directly into a Universal Machine Learning Force Field (CHGNet) to instantly relax the atomic coordinates, bypassing days of expensive DFT calculations and then run the relaxed structure through ALIGNN to validate its Energy Above Hull, resulting in a mathematically viable file generated in under three minutes.
Video
Transcript
Generated about 2 months ago
Summary
Generating a talk summary...
View full transcript
Speaker 0: Efficient Intern of doing tons and tons of, testing in labs. So what we build is an AI model that can generate, technical crystalline, chemical structures that can be used to generate more efficient, batteries specifically. So for, generating molecules, in my first time, I tried to use transformer Model. But, that that while it did generate molecules that looked right, this does look like a molecule. However, when you look at a macroscopic level, it was not as accurate.
Speaker 0: This is because it's like training an image model on on base 64. Like, if we, train a transform model to just just Armaan the base 64 Stack Context of the Places. And then once you try to generate, it might generate something that looks like an image, but it wouldn't understand what the features are. So, then we AI to a diffusion model, which was, which was able to, Event which reliable to learn the structures, of of author schemas of of a water molecule would be. So, for this, BITS used a Matrigen, as the base model, Microsoft's Matrigen.
Speaker 0: And then, we we used Mhate material project dataset. So from that, we tool certain molecules that are related to batteries and then fine tuned AI tuned the model based on allergens, based on these molecules. BITS what we did different is, we we categorized different aspects of the, different aspects of the molecule with, with, different diffusion models. Like, we using, d 3 PM for, atoms because it is able to have a mask structure. So that would be able to, because since Litmus it's it's technical, so you cannot put random noise into it.
Speaker 0: So, through that, we because we use d 3 PM for Custom structures. And then for, AI, lattice coordinates Yeah. For, for for, the coordinates of where the atoms would be, we used Venue. That is a it's a variation exploring, diffusion model. So here, what since the noise increases as, as you add more, as it continues to train, there it would be able to, identify how this it would be able to know the, the distances to be AI/ML within the molecular.
Speaker 0: It's in current. Yeah. And then for, for for the matrix or where the positions of these molecules would be, it is, we using v Voice. So this is because of this, the diffusion model would be able to, know that the size the distances between the lengths are different. So with this, first, on on our first attempt, we were able to get, stats.
Speaker 0: Yeah. We were able to get, e hul. E hul we put conditions on e hul and the stability. So what this does is, if the, molecules are closer to 0, it'll generate a more stable molecule. So, we put we sent this as a conditioning Event, which, resulted in we we, initially got 7, reliable structure stable structures.
Speaker 0: And then, but the issue with that was it it was generalized based on since MacroGen had a a lot of, molecules, we had to, train Mode even with, even with more molecules. So what in our third, attempt, we we we added Transitions, specific molecules only. But then instead of retraining it again, the diffusion Model, what we did was we took the, stored, the checkpoints from the v 2 model and then merged it to another model which we trained on specifics. So it was able to still have the idea of how to generate molecular, at the same time sharing a more train, constrained Day. And with that, it was I can show up out Mode outputs.
Speaker 0: So each time the flow is Presenter a remote. Let's say AI need a molecule that has, like, lithium or tool lithium or oxygen or ruthenium, and then it takes those it takes a text prompt and then turns them into a vector that can be put HINTS a diffusion model. And that once that, that Event has gone through Gemma, which is, which is part of, which which is part of Allergen. What it does, it it Stack the multiple, aspects of a molecular, like Zoom, the where the position is and how stable it needs to be, and then puts them together. And with that and after that, it generates a report like this with, with which molecules are stable.
Speaker 0: What it does to check if a molecule is stable or not, it runs through CHG NET. It's a it's a graph neural network for, for checking the stability of molecules. So once it once it does that, it also checks it with the materials project dataset if a molecule generated is novel or not. And, with that, it generation, Stack on how stable each of the molecule is. And another thing, it also checks using, CSG net if how how much capacity can the molecular generated Folder because enforce, batteries, that's a very important thing.
Speaker 0: And BITS, this this remote it generates are useful for scientists and chemists who are working in, the battery molecule generation field. The its main output is apart from this is to generate CIF files. Those are the files that can be inputted into programs, such as Event where, they can visualize it. So the CIF files are its main output app well as a free Courts of how the molecules will be generated. In our also in our first attempt, what we, did was we pregenerated Artist of molecule, molecules with the diffusion model and then, use whenever a person enters a text prompt, it would BITS would Custom filter through the existing dataset.
Speaker 0: But then that wouldn't like, if there's something completely new, it wouldn't be it would just hallucinate. So, each so now what we Day is each time a prompt is written, it BITS goes back to it runs author, AI, author training loop where it Stack, it takes the, like, specific molecules and then fine tunes it. Zoom no. Okay. Yeah.
Speaker 0: So through this, we were able to the green molecule the green highlighted ones are something that can be that has potential to be, researched. Yet BITS also calculates the volume of also calculates the volume of each molecule. This is also through, Scene. There are sometimes issues where, the molecule is not it like correct, but then the angles are not right. So the it goes through relaxation using AI menus well, tool correct to correct it as well.
Speaker 0: Yeah. That's about it. Thank you. Do we have a word?
Speaker 1: Teaching about, when you're giving the text prompt for the molecules, do you define the percentage of the implement, or is it as per your AI? Or how is it
Speaker 0: You type in the text free, like, generate a scale, generate a stable molecule structure that contains, you AI/ML the molecule l I
Speaker 1: t LLM.
Speaker 0: And all that. And you can say that it needs to be under APIs, Co-Founder, like, 0.01 stability.
Speaker 1: Okay.
Speaker 0: So based on that, it will, like, generate the proportions of how much of each molecule tool.
Speaker 1: So the screening sorry. 1 more question. The screening list is a list of combinations of the
Speaker 0: tool Conversation and the proportions of it as well.
Speaker 1: So there's multiple that can be generated? Yes. Oh, okay. Okay. That's it.
Speaker 0: You can you can set how much you want to be generated as ceo. Okay. Yeah.
Speaker 1: Yeah. That's it. Thank you.
Speaker 0: Yes? I don't know.
Speaker 2: Thank you very MCP, first of all. What what data did you input? Did you use Simpl, protein data-packet set before and when you started?
Speaker 0: Yeah. We used, Mhate material project dataset. So that had billions of molecules. But what lead, Scene this is specific to battery molecular. We just, we just run a search tool just find molecules training to that.
Speaker 2: Yeah. And, and, you mentioned training. How long did each training run on leveraging run? For how long did you train and, what, GPU and compute did you use?
Speaker 0: Okay. We used the Videos 28 10 g GPU. We for each training run, it was around, around 24 hours. Yeah.
Speaker 2: Yeah. Thank you very much. Thanks.
Speaker 1: I think we'll have time for 1 more question. Anyone's app for it? Alright. That's it. Thanks so much,
Speaker 2: Hishaam. Good.