MenuMind: On-Device AI for Menu Translation and Allergen Awareness
Learn to build an on-device AI app for translating menus and identifying allergen risks, focusing on practical food understanding and clear user warnings.
Overview
MenuMind is a Flutter mobile app that helps travelers understand restaurant menus and identify possible allergen risks using on-device multimodal AI. In the demo, I will show how a user can scan a menu image, extract dishes and ingredients, translate unfamiliar items, flag common allergens such as nuts, dairy, gluten, seafood, and eggs, and explain the risk in simple language before ordering.
The app is designed as an awareness assistant, not a medical guarantee: it highlights possible allergen risks and encourages users with severe allergies to confirm with restaurant staff.
Video
Transcript
Generated 2 months ago
Summary
Generating a talk summary...
View full transcript
Speaker 0: Are you good? Whatever you're ready. Are you good? AI? Force good?
Speaker 0: Okay. Good morning, everyone. Train you for AI the AI Tinkerers Dubai. So I will talk about today for on device LLM model. So so you know Gemini?
Speaker 0: Gemini, it's it's a popular LLM model for Google. Gemini also support l l on device model. It's called JIMA. This JIMA, chat can launch in on device. How can it he can run-in mobile application in Android or AI.
Speaker 0: So my idea now or my example, it was for Menu AI. This is for Menu trans translation Zustand for allergy detection. We can call as a breakpoint for this application. I only, using this AI how I can use in the local model local Model Mobile application to a good idea. So as we know, when we travel for many different country with new culture, with new language, the the AI for the problem for us force how I know the language for this, food and if this tool, I have allergy on it or not.
Speaker 0: So I create simple app. This simple app, it was in Flutter, and it was take take for me. We can run the demo at first, and I will explain. Here, this is the application. When BITS run, AI I here, I initialize the model.
Speaker 0: After I initialize the model, here, I scan the menu. After I scan the menu, I can select from image or select free gallery. I correct from gallery some Places, Scene. So I select for APIs Scene, and I need Zustand what this Context, and I have allergy for this food or not. The main idea for this using a AI model without without Internet, when you visit any country without and you you don't have any accessibility for Internet, you have the power of of LLM model in your device.
Speaker 0: And here, as we see, after after he, after he analyze after he analyze the, the image, what he detect he detect 2026, this is the main the main Date, and he give he give me a brief about this lead, what's the what's the content and some flag. BITS, it's it's it's Allergen. OpenAI here, I chat go for profile and select chat allergy I have. This is the common allergy. So I selected here my Allergen.
Speaker 0: It was for milk and eggs. So when I scan any meal, it's contained for egg force Sahil or mail, he will give alert. So this lead, it's have allergy for you. And if I can scan again for another meal, here we can we are here we will scan a plate for, mix of eggs switch the cheese. So he will detect now this Places.
Speaker 0: It's have allergy for me. So what happened here, I Cloud the Mode model with the the the sample prompt for analyze this analyze analyze this analyze this Code analyze this picture with my allergy. So if we go to force our code, the main, the main hero for this story, it was the open source, open source package. It's supported for Google. It's called Scene it's called AI.
Speaker 0: This MediaPipe, what we have, this is the MediaPipe. This MediaPipe BITS created from Google. It's open source. It's only for Date AI, LLM model on device. We can Simpl, we can Simpl.
Speaker 0: It was like Gemma. We can run LLM model in our machine. So training the same, we can run LLM model in device reducing this managing, this force media bio. So what we hear, Google, when Google created this package, it's supported for it's created mini mini version for this. Flutter created the Gemini AI create Menu bio and created also what we have here.
Speaker 0: It's supported BITS have a media pipe for Android, media pipe for iOS, and 2026 media pipe for gen AI for generated for generated Day, and also Menu Gemini AI for Flutter. So by using this package, we can, we can run. How we can achieve force what we can using by this pipeline, we can, it's train, it's a can run-in Android. It's a can run-in iOS. And also chat we can Day, we can we can do IDENTIFICATION.
Speaker 0: We can do building, and we can we can do language detection and infrastructure. All of this we can Day in Mobile device. So if we pack for our implementation, we, we, what I what I implement Intern, I using the Flutter Community. It's created some package. It's Code called Flutter Gemma.
Speaker 0: So when we pack for this, this is a Flutter gma. The Flutter gma, it is it is a new version for from Google to to run a little model in a AI, and they have mini variant. We have here 4 variant. We Mhate, Gemma with 3 n with 4,000,000,000, and we have also for 2,000,000,000, and we have some specification for this. So we can using 1 of APIs model and install in our app.
Speaker 0: So what AI you using? Are you using for 4,000,000,000? What we need this is the open source. This is the open source. So what we what we have now, we can go for this mode data model and this.
Speaker 0: This is for us for this is the first thing. Okay. And is is our Code, it was you share my code? Okay. For this, we can we the first step, we download the model on on the AI.
Speaker 0: Where we can download this device, we have 2 Remotion. We can download the Zoom model from Honey Force or from Kaggle. And if we need, if we are downloading from Honey Force, we lead the access key. So this is turn Use from Honey Zustand free create access token from Free Face to download it. What happened after that?
Speaker 0: I have here the data source remote tool contact with LLM Model. And here, the first step, we initialize the LLM model with the token. And here, what app, AI as Awareness Date, I download the model. After I download the model, I store it in the device storage. But after that, when we LangChain the application, we scale we call this model bus to put it in Flutter Gmail.
Speaker 0: This is a this is a package for it will be helping me to run a Model model in the device. And here we Mhate, I I I created, some common method. The first method, it was for generates text response. This, we can we will create session. This session will contact switch LLM Model, and after that, it was, gets Sponsors based on the text.
Speaker 0: So this function only, the main goal for this, I will give it a text Zustand as a response oppose the text. Why free using using this mac using AI, this, this, this Remotion? AI use it for local. Because Gemma, Litmus supported around free, 15 language. So what I do, I make main, main language is English.
Speaker 0: But if you need another language to detect what you need to display the allergy and the description force, for everything, I use Inggijima also for translation. And also what here, this is the analysis with image. The first thing we, pick up the image from gallery and this image AI will convert it force, AI convert it to list binary Zustand I will give it for the LLM model with the Mode prompt. So the LLM model, I will give it the binary of image and the prompt. The remote, it was it will it was the main task for LLM.
Speaker 0: So here, how what we see, this is our prompt. The old prompt, it was for explaining menu item as a JSON. And for each item, I need the main concept for Zustand the force Mode IDENTIFICATION for each for each dish. The original AI/ML, the transaction name because the the original Date, it was the the chat the name is already in the image. And, also, what we have here, after we, after we got the response, what we do, we extract Scene because what app now, after I call this prompt, the for an for normal LEM model, the LEM model is returning JSON, but the but this JSON has a string.
Speaker 0: So I need extract the real text from this sharing. So I create some, some function to LLM me to extract this extract APIs, extract this train dish. So and here, if I will pack for the same for the initialize, here in the model, what we can do AI the session here force we can I what free if I tool align with my task, we can specify it for the ODeX Agents for Intern? And also here, when you create the session, what you have, you have temperature. This is the temperature to to let the Remote model to be AI for the chat result the design.
Speaker 0: And also for a top k. This is a top k also. So based on all APIs param, we can exchange for give it Mode, accurate data. So you have now a AI LLM model on device, and this is LLM. With your specification here, what you do switch, you can create a Engine, and mid-session, we will create we will contact with LLM model.
Speaker 0: So this is this is also the, the big point for how the Lead model. Use. Thank you.
Speaker 1: Okay. You
Speaker 0: can just stable wrong.
Speaker 1: Okay. AI had a question. So, if I take a picture, the LLM model will analyze it and say, okay. This food item has, nuts. Don't eat that.
Speaker 1: Is it all like, do you also have the capability to give it context? So this menu item is, I don't know, butter chicken Yeah. Instead of paneer butter I don't know.
Speaker 0: Yes. Yes. We can define all of this inside the prompt bit. So inside the prompt bit, we chat, here, here, my prompt bit, it's force symbol. So to, give it more context, you can define all of all of all of what you need inside this context.
Speaker 0: Okay. And, also, it will be clarified what is the image it's contained.
Speaker 1: Okay. Cool. I have 1 more question unless Yeah. Anyone else that's Ajay. I'll just keep going.
Speaker 1: Can it also store, like, context or history? So if you keep going to the same restaurant, you don't wanna keep taking the same picture for different AI. Is it a way tool learn?
Speaker 0: Yeah. For me for me for my idea, I I store every every every message for the AI, I store it already. Because if you have if you scan before this menu, I don't need to tool the LLM model again to scan it. So I store this I store, I store, the response in the local storage to
Speaker 1: Same.
Speaker 0: Let the let to limit the request for LLM model.
Speaker 1: That's cool. Thank you.
Speaker 0: Yeah. You're welcome. Anybody else? Sorry.
Speaker 2: 1 question on the LLM model switch we are putting in the device. So let's say, there are new restaurant in town or something who are having a new disease. So how this LLM is getting refined to have the latest of Day data across the Architecture, AI, for new this is new Student. How this will get AI author because we are putting that model into device. So are we having a capability where we will connect to some Model online or something like that or how is it?
Speaker 0: Yes. All of this Model, it was already in, in a hungry Mhate. And also what we what we have here, we have the extension for enforce Artist, for task. If you Senior model here, what's the Engine? Mid-session City was for, dot task.
Speaker 0: The normal extension it was for dot bin. The force task, it's, it's AI all metadata. So this model BITS Mhate for tokenizer. It's all it's also abandoned for this. So any model, what you can using for us, the the hand build, what we can, you can using the the task extension because this is the Engine.
Speaker 0: It's a Context tool also all, AI. How we how user we can AI the user input to the LLM model. If you have the normal, the normal LLM model, it's look like for, extension force build. So at this moment, you need you have Menu, create your tokenize. So for this, pretty sure you can access any LLM model and use see the specification what the histories because all of APIs, it was be in the RAM.
Speaker 0: So based on specification in your mobile, you can check the capability for each model. Okay. Thanks. Yeah.
Speaker 2: Got a question.
Speaker 0: Yes. The meeting's the last 1 and then we'll head to the next 1.
Speaker 3: Yeah. Okay. Thank you. Again, a question on the model itself. So Google has already released its Artist release of, local models switch this Gemma 4 n.
Speaker 0: Ximena.
Speaker 3: So, like, is there a particular reason where wherein why you chose 3 n itself instead of the,
Speaker 0: like Yes. Because Birla this because the fine using. The the the size for this model when you download it in your app because chat the first time, what we haven't well, as the first time you download the whole model in your device. So what we have now in if use if we pack for for this screen for, GMS 3 n, it's also offered scale, small Model. So we can Date what you what you know to check also on the Model.
Speaker 0: But for GMS 4 n, also, if use if if the AI, it was matched with Courts, mobile IDENTIFICATION, you can using. No, no problem. Because, and also for, the media AI, this is enforce for all local LLL model. And this package is not only specific for Gemma. You can using, any any bubbles, another Mode Model, and then download it in your app, like deep Scene or something else.
Speaker 0: Yeah. Okay. Thank you. Much, Arthur. Thank you.
Speaker 0: Next app, we have a. Hold
Speaker 3: on. Thank you. Where's
Speaker 0: your mic? Okay. Like, focus and focus on this, but right now, all 4 day. Yeah. Okay.
Speaker 0: This is like a AI model I can show start. Yeah.
Speaker 3: I mean, this is that. Just show really what you need. How are you filling this server?