Automatic cooking method of cooking robot based on artificial intelligence
By combining multimodal sensors and machine learning, the robot can perceive and dynamically adjust the state inside the pot in real time, solving the problems of single perception and rigid decision-making in traditional automatic cooking equipment, and realizing a personalized automatic cooking method for stir-frying robots.
Patent Information
- Application Number
- CN202511908741.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-17
- Publication Date
- 2026-01-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional automatic cooking equipment lacks the ability to perceive and dynamically adjust the real-time status of ingredients in the pot, making it unable to adapt to complex changes in different ingredients, initial conditions, and user tastes. Its level of intelligence is limited, and its multimodal data fusion method is simple, affecting the stability and personalization of cooking results.
Multimodal sensors are used to simultaneously collect images, temperature and humidity data inside the pot. A multimodal fusion matrix is generated through dynamic weighted fusion. Combined with a machine learning decision model and a self-learning module, the status of the ingredients is analyzed in real time and personalized cooking control instructions are generated. The model is optimized based on feedback information.
It achieves refined and adaptive state perception of complex cooking environments, improves the stability and personalization of cooking results, and ensures scientific decision-making regarding ingredient ripeness and heat transfer efficiency.
Smart Images

Figure CN121348779A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence technology, specifically relating to an automatic cooking method using an artificial intelligence-based stir-fry robot. Background Technology
[0002] Traditional automated cooking equipment or robots typically rely on preset programs and time-temperature curves to operate, lacking the ability to perceive and dynamically adjust the real-time state of the ingredients in the pot. Furthermore, most devices rely solely on temperature sensors for control, lacking synchronous and comprehensive analysis of the visual state of the ingredients (such as color, shape, and texture) and the humidity environment. This results in unstable cooking effects. Control strategies are often based on fixed logic or simple conditional rules, unable to adapt to the complex changes in different ingredients, initial conditions, and user tastes. Their level of intelligence is limited, and the systems are usually closed-loop, making it difficult to receive and integrate real-time user feedback or long-term preferences during the cooking process, thus failing to achieve personalized flavor customization. When attempting to use multiple sensors, the fusion methods for different modal data (such as images, temperature, and humidity) are simple (such as fixed weights or simple splicing), failing to adaptively weight according to the dynamic changes in the cooking process, affecting the accuracy of state judgment. Therefore, there is an urgent need for an automated cooking method that can perceive the cooking status in real time and in multiple dimensions, and make dynamic and personalized decisions based on intelligent algorithms. Summary of the Invention
[0003] (1) Technical problems to be solved This invention discloses an automatic cooking method for a stir-fry robot based on artificial intelligence, which aims to solve the problems of traditional automatic cooking equipment having limited perception capabilities, rigid decision-making, and a lack of intelligence and self-learning ability in cooking control strategies.
[0004] (2) Technical solution This invention discloses an automated cooking method using an artificial intelligence-based stir-fry robot, comprising the following steps: Step 1: Simultaneously acquire image, temperature, and humidity data inside the pot using a multimodal sensor, and preprocess and dynamically weighted fuse the data to generate a multimodal fusion matrix; Step 2: Based on the multimodal fusion matrix, analyze the visual state of the food and the thermodynamic state inside the pot, and fuse them to generate a comprehensive state vector for decision-making; Step 3: Input the comprehensive state vector into the machine learning-based decision model to generate a control instruction sequence that includes the heat adjustment amount and the stir-frying action; Step 4: Execute the control command sequence, monitor the execution process and collect data after execution, and generate a feedback result vector containing the execution deviation and the post-execution state; Step 5: Based on the feedback result vector, perform stability updates on the parameters of the decision model and its internal optimization target state; Step 6: Receive user preference input and evaluation feedback, and update the decision model and optimization target state in a personalized manner based on the preference and feedback information.
[0005] Further, step 1 includes the following steps: Step 101: Acquire cooking data inside the pot using a multimodal sensor, including image data, temperature data, and humidity data inside the pot; Step 102: Preprocess the acquired image data, temperature data, and humidity data inside the pot respectively; Step 103: Calculate dynamic fusion weights for the preprocessed in-pot image data, in-pot temperature data, and in-pot humidity data respectively; Step 104: Using the calculated dynamic fusion weights, the image feature data, temperature feature data, and humidity feature data are weighted and fused to generate a multimodal fusion matrix.
[0006] Furthermore, the dynamic fusion weights The calculation formula is:
[0007] in, For modality Dynamic fusion weights; Modal index; Modal index; Image modality Temperature modes; Modality The energy gradient; Modality The energy gradient; For modality Biological thermal sensitivity coefficient; For modality Biological thermal sensitivity coefficient; For the current moment The humidity value inside the pot.
[0008] Furthermore, step 2 includes the following steps: Step 201: Extract and decode visually relevant information from the multimodal fusion matrix, input it into the food ingredient state analysis model, identify and output the intermediate state feature vector of the food ingredient; Step 202: Extract and decode temperature and humidity related information from the multimodal fusion matrix, and calculate the heat conduction index based on the heat conduction-humidity coupling model. :
[0009] in, The thermal conductivity index; Thermal conductivity; The rate of temperature change; Humidity correction factor; Initial humidity; : for the current moment The humidity value inside the pot; : for the current moment The temperature value inside the pot; Step 203: The intermediate state feature vector of the food ingredient is weighted and fused with the thermal conductivity index to generate a comprehensive state vector for decision-making.
[0010] Furthermore, step 3 includes the following steps: Step 301: Input the comprehensive state vector into the decision model based on offline reinforcement learning to predict and output the optimal cooking action in the current state; Step 302: Based on the predicted optimal cooking action, compare the comprehensive state vector with the optimized state vector provided by the self-learning module to obtain the state deviation; Step 303: Input the state deviation into a Sigmoid function and map it into a smoothing adjustment coefficient. Multiply the difference between the current pot temperature and the system's preset maximum safe temperature by the smoothing adjustment coefficient to obtain the heat adjustment amount. Step 304: Based on the calculated heat adjustment amount and the preset stirring angle, generate the control command sequence for the current moment.
[0011] Furthermore, step 4 includes the following steps: Step 401: According to the control command sequence, drive the actuator to perform the corresponding heat adjustment and stirring actions; and use sensors to monitor and obtain the actual temperature value and actual stirring angle during the execution process in real time; Step 402: Based on the actual temperature value and actual stir-frying angle obtained from monitoring, compare them with the preset target temperature and target stir-frying angle in the control command sequence, and calculate a comprehensive execution deviation based on the difference between the two. Step 403: After the control command is executed, the image, temperature and humidity data inside the pot are collected again by the multimodal sensor, and the pot state vector after execution is generated based on this data; the execution deviation is combined with the pot state vector after execution to construct a feedback result vector.
[0012] Furthermore, step 5 includes the following steps: Step 501: Calculate the loss of the decision model based on the difference between the in-pot state vector and the optimized state vector in the feedback result vector; Step 502: Update the internal parameters of the decision model in step 301 based on the calculated loss; Step 503: Based on the executed in-pot state vector and the user's historical preference data, perform iterative stability updates on the optimized state vector; Step 504: Apply the updated decision model parameters and the optimized state vector to the next round of cooking decision-making process.
[0013] Furthermore, step 6 includes the following steps: Step 601: Receive cooking preference information input by the user and convert the preference information into optimization parameters that the system can recognize, so as to adjust the optimization state vector; Step 602: During the cooking process, based on the multimodal fusion matrix and the integrated state vector, provide the user with real-time feedback information including cooking progress and ingredient status; Step 603: After cooking is completed, receive sensory evaluation data of the dish made by the user based on real-time feedback information, and preprocess the evaluation data into standardized user feedback signals; Step 604: Input the optimization parameters and user feedback signals into the self-learning module to update the parameters of the decision model and the optimization state vector in a personalized manner.
[0014] Furthermore, based on step 103, an initial ecological entropy balance weight for generating information entropy is introduced:
[0015] in, Modality The initial ecological entropy balance weight; Image modality; Temperature modes; Humidity mode; Modality The energy gradient; : for the current moment The humidity value inside the pot; Modality Ecological sensitivity coefficient Modality Information entropy; Entropy balance adjustment parameter.
[0016] Furthermore, a spatial attention mechanism and a temporal recursive network are introduced into the initial ecological entropy balance weights to generate spatiotemporally adjusted weights:
[0017] in, Modality In spatial location and time step Spatiotemporal adjustment weights; Spatial attention function, used to calculate temperature field through convolution operations. Local importance; Temporal recursive networks capture the dynamic changes of weights over time; : Normalization function.
[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. A dynamic fusion weight calculation formula based on energy gradient and humidity inside the pot is proposed, which enables the fusion weight of image, temperature and humidity data to be adjusted in real time with the cooking process. Furthermore, the concept of "initial ecological entropy balance weight" based on information entropy is introduced, and "spatiotemporal adjustment weight" is generated by combining spatial attention mechanism and temporal recursive network, realizing the fine-grained and adaptive adjustment of data fusion in spatial and temporal dimensions. This greatly improves the system's perception accuracy and robustness to the comprehensive state under complex cooking environment, and makes multimodal data fusion closer to the real physical and chemical change process.
[0019] 2. In addition to analyzing the visual state of the ingredients, the system also calculates the heat transfer index using a "heat conduction-humidity coupling model" and merges the two to generate a "comprehensive state vector" for decision-making. This vector simultaneously encodes the biochemical changes of the ingredients and the physical thermodynamic state inside the pot, providing a more comprehensive and scientific state description basis for subsequent intelligent decision-making, enabling the decision-making model to consider both the ripeness of the ingredients and the efficiency of heat transfer. Attached Figure Description
[0020] Figure 1 This is a schematic diagram of the overall process flow of the present invention.
[0021] Figure 2 This is a schematic diagram of step 1 of the present invention.
[0022] Figure 3 This is a schematic diagram of step 2 of the present invention. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] like Figure 1 As shown, this invention discloses an automated cooking method using an artificial intelligence-based stir-fry robot, comprising the following steps: Step 1: Simultaneously acquire image, temperature, and humidity data inside the pot using a multimodal sensor, and preprocess and dynamically weighted fuse the data to generate a multimodal fusion matrix; Step 2: Based on the multimodal fusion matrix, analyze the visual state of the food and the thermodynamic state inside the pot, and fuse them to generate a comprehensive state vector for decision-making; Step 3: Input the comprehensive state vector into the machine learning-based decision model to generate a control instruction sequence that includes the heat adjustment amount and the stir-frying action; Step 4: Execute the control command sequence, monitor the execution process and collect data after execution, and generate a feedback result vector containing the execution deviation and the post-execution state; Step 5: Based on the feedback result vector, update the parameters of the decision model and the internal optimization target state; Step 6: Receive user preference input and evaluation feedback, and use the preference and feedback information to drive the decision model and optimize the personalized adjustment of the target state.
[0025] Specifically, step 1 includes the following steps: Step 101: Acquire cooking data inside the pot using a multimodal sensor, including image data of the inside of the pot. Boiler temperature data Humidity data inside the pot Simultaneously, the camera, temperature sensor, and humidity sensor are activated to capture real-time images of the inside of the pot, measure the temperature at various points, and detect the steam humidity, obtaining a complete picture of the cooking process in one go without leaving any blind spots. All data is completely synchronized in time to ensure the accuracy of the analysis. For example, it can strictly correspond the "color of the ingredients at a certain moment" with the "accurate temperature at the same moment".
[0026] Multimodal sensor data: Image data: Images of food in a pot were captured using a high-resolution camera (1920×1080 resolution, 30fps). The dataset contains 5,000 images of the cooking process, covering 10 common Chinese dishes (such as Kung Pao Chicken, Shredded Pork with Garlic Sauce, and Stir-fried Pork with Green Peppers) and different stages of cooking. The data comes from a simulated kitchen environment built in the laboratory and is supplemented by the public dataset "Fd-101".
[0027] Temperature data: The temperature field distribution inside the pot is collected by an infrared temperature sensor array (16×16 resolution, measurement range -20°C to 500°C, accuracy ±0.5°C). The data is sampled once per second, and the dataset contains about 100,000 sets of temperature field data, covering different cookware (iron pot, non-stick pot) and heat sources (induction cooker, gas stove).
[0028] Humidity data: High-precision humidity sensors are used to collect humidity changes inside the pot, and temperature data is recorded synchronously. The size of the dataset is consistent with that of the temperature data.
[0029] The system simultaneously triggers visual sensors (such as RGB-D cameras) deployed above the cookware, temperature sensor arrays integrated into the bottom or wall of the pot, and non-contact humidity sensors to synchronously capture images, temperature field distribution, and humidity data inside the cooking pot at a fixed time frequency (such as 10 times per second). Strict time synchronization ensures that the images, temperature, and humidity data are completely aligned on the time axis. This allows subsequent analysis to accurately correlate "color change in an image of a certain area" with "the actual measured temperature of that area" and "the current steam humidity," avoiding misjudgments caused by time misalignment.
[0030] Step 102: The image is normalized in size and corrected in color, encoding semantic information such as the color, texture, and shape of the ingredients. The temperature field data is interpolated to generate a regular two-dimensional temperature distribution map. At the same time, the humidity reading is filtered and smoothed to obtain a stable current humidity value. The preprocessing steps (such as noise reduction and filtering) effectively filter sensor noise and instantaneous interference (such as image occlusion caused by oil droplet splashes and humidity reading spikes caused by instantaneous steam), improving the reliability of the data. Feature extraction condenses the high-dimensional raw data (such as millions of pixels in the image) into low-dimensional but more information-density feature vectors, which greatly reduces the complexity of subsequent calculations and highlights the information most relevant to cooking.
[0031] Step 103: Based on the real-time physical state inside the pot, intelligently determine which data to trust more, for each mode. (in, Calculate the dynamic fusion weights for image mode, temperature mode, and humidity mode respectively; The dynamic fusion weight The calculation formula is:
[0032] in, For modality Dynamic fusion weights; Modal index; Modal index; Image modality; Temperature modes; Humidity index; Modality The energy gradient; Modality The energy gradient; The biological thermosensitivity coefficient; For modality Biological thermal sensitivity coefficient; For the current moment The humidity value inside the pot.
[0033] When there are drastic local changes inside the pot (such as a sudden temperature rise at a certain point) The energy gradient of the temperature mode is large, or oil splashes cause this. (The energy gradient of the image mode is large), if the environment is dry at this time ( (Small), will automatically and significantly reduce the weight of the corresponding mode. If the data indicates an "outlier" or "unstable state," its impact on the overall judgment should be reduced.
[0034] When the humidity inside the pot is very high ( (Larger), the denominator increases, and the modal weights are affected by the energy gradient. The resulting differences were smoothed out, which simulates the physical fact that cooking in a steam environment is gentler and the overall data reliability is higher.
[0035] This design achieves context-aware adaptive fusion, much like an experienced chef who can dynamically decide whether to trust what the eyes see or what the thermometer displays based on "how dry, hot, and unstable the pot is." This mechanism can automatically reduce the contribution of abnormal data and improve the overall robustness of the system.
[0036] Step 104: Using the calculated dynamic fusion weights, process the in-pot image data. Boiler temperature data Humidity data inside the pot Perform weighted fusion to generate a multimodal fusion matrix. .
[0037] Since the above weight calculation formula does not take into account the lack of adaptability, ignores spatiotemporal correlation, and does not model sensor uncertainty, a multimodal fusion weight calculation method based on "ecological entropy balance and dynamic self-learning" is proposed.
[0038] By introducing an entropy optimization mechanism, the distribution of fused weights is ensured to avoid "overfitting" to a single mode and to maintain information diversity. At the same time, by introducing a temporal recurrent network (RNN) and a spatial attention mechanism, the two-dimensional distribution and temporal changes of the temperature field inside the pot are captured. Furthermore, learnable parameters are introduced to enable end-to-end training capability for weight calculation. Based on a Bayesian probability framework, sensor noise and uncertainty are explicitly modeled to improve the robustness of weight calculation.
[0039] First, based on step 103, an initial ecological entropy balance weight is generated by introducing an ecological entropy balance model. The formula for calculating the initial ecological entropy balance weight is as follows:
[0040] Symbol explanation: Modal The initial ecological entropy balance weight; Image modality; Temperature modes; Humidity mode; Modality The energy gradient; Modality The energy gradient, analogous to the activation energy in the Arrhenius equation, for temperature modes... For visual modalities (Image gradient intensity), for humidity mode (Variance of humidity variation); Humidity value, as an environmental regulating factor Modality The ecological sensitivity coefficient, analogous to the enzyme-catalyzed reaction rate, has an initial value of 0.1 and can be adaptively adjusted through training. Modality Information entropy, calculated as ,That It is the probability distribution of modal features, used to measure the amount of modal information; Entropy balance adjustment parameter (typical value 0.5), used to control the contribution of the entropy term to the weights; : Summing over all modes for normalization.
[0041] This formula introduces information entropy. To avoid a single modality (such as temperature) from becoming overly dominant in certain scenarios, we must ensure the diversity of fusion weights and achieve dynamic balance.
[0042] Furthermore, considering the spatial distribution (two-dimensional / three-dimensional) and temporal changes of the temperature field and food state within the pot, this application also introduces a spatial attention mechanism and a temporal recurrent network (RNN) to perform secondary optimization on the optimized dynamic fusion weights to obtain spatiotemporal adjustment weights. The calculation formula for the secondary optimized dynamic fusion weights is as follows:
[0043] Symbol explanation: Modality In spatial location and time step Spatiotemporal adjustment weights; Spatial attention function, used to calculate temperature field through convolution operations. Local importance; Temporal recursive networks capture the dynamic changes of weights over time; : Normalization function, ensuring that the sum of all modal weights is 1.
[0044] Through spatial attention mechanism, the weights are no longer scalars, but are related to the position inside the pot, capturing the uneven distribution of the temperature field; through RNN, the weights evolve dynamically over time, adapting to the continuous changes in the state of the ingredients during the cooking process.
[0045] To model sensor noise and uncertainties (such as grease reflection in visual images and temperature sensor errors), a Bayesian framework is introduced to correct the dynamic fusion weights after secondary optimization. The correction formula is as follows:
[0046] in: :mold The final fusion weight; Modality The reliability probability is determined by historical data. The estimation is based on Bayesian posterior probability calculation. : Summing over all modes for normalization.
[0047] By explicitly modeling sensor uncertainties using Bayesian probabilistics, the system can automatically reduce the weights of relevant modes and improve fusion robustness in cases of visual bias (such as complex lighting) or large temperature measurement errors.
[0048] The final fused multimodal matrix Calculated by weighted summation:
[0049] Symbol explanation: Current moment ,Location The multimodal matrix at the location; Modality The characteristics are represented.
[0050] To improve adaptability, key parameters in the formula (such as...) The parameters of the attention network and RNN are all set as learnable parameters and optimized through end-to-end training. The loss function is designed by combining the cooking results (such as the color and doneness of the dish) with the intermediate fusion quality (such as the information entropy balance) to ensure that the system can adapt to different ingredients, cookware and user preferences.
[0051] Through entropy term The introduction of this approach avoids the dominance of a single modality, ensures information diversity, and combines spatial attention with temporal RNNs to enable weight calculation to have spatial resolution and temporal continuity, breaking through the limitations of scalarization. By explicitly modeling sensor uncertainties, it improves the stability of the system in complex cooking environments (such as oil reflection and light changes).
[0052] Specifically, step 2 includes the following steps: Step 201: Extract and decode visual information from the multimodal fusion matrix, input it into the food state analysis model, identify and output the intermediate state feature vector of the food; this step inputs the visual information in the fusion matrix into an AI model trained on cooking images. The model will identify and quantify the current state of the food like an experienced chef, such as: "The chicken is golden yellow (about 80% cooked)" and "The vegetables are starting to soften".
[0053] Step 202: Extract and decode temperature and humidity related information from the multimodal fusion matrix, and calculate the heat conduction index based on the heat conduction-humidity coupling model. :
[0054] in, The thermal conductivity index; Thermal conductivity; The rate of temperature change; Humidity correction factor; This design considers the initial humidity level; it goes beyond just temperature readings to assess how well the heat is being used. For example, it can identify situations where, despite high heat output, significant moisture evaporation is taking away heat, resulting in low actual heating efficiency.
[0055] Step 203: The intermediate state feature vector of the ingredients and the thermal conductivity index are weighted and fused. The system will dynamically adjust the importance of the two according to the cooking stage, such as focusing more on physical temperature rise in the early stage of aroma-making and more on color change in the late stage of sauce reduction, to generate a comprehensive state vector for decision-making. The comprehensive state vector The calculation formula is:
[0056] in, For perceived weights; This design simultaneously knows how the food looks and the heat conditions of the pot, avoiding the blind spots of one-sided judgment. It can intelligently focus on different key points at different cooking stages, simulating the decision-making logic of a human chef.
[0057] Specifically, step 3 includes the following steps: Step 301: Input the comprehensive state vector into a decision model based on offline reinforcement learning (based on OfflineRL, such as the Overcooked-AI variant) to predict and output the best macroscopic cooking action in the current state; input the comprehensive state vector representing the current cooking state into an AI decision-making brain (offline reinforcement learning model) trained on massive amounts of cooking data. This model does not only determine the current action, but predicts the optimal action sequence over a period of time in the future, such as: "stir-fry over high heat for 30 seconds, then turn to medium heat and stir-fry, adding the sauce in two batches," enabling the robot to have the planning ability to "think three steps ahead," ensuring that the entire cooking process is coherent and efficient, and avoiding step conflicts or flavor loss caused by short-sighted decisions.
[0058] Step 302: Based on the predicted optimal cooking action, compare the comprehensive state vector with the optimized state vector provided by the self-learning module to obtain the state deviation; Step 303: Input the aforementioned state deviation into a Sigmoid function, map it to a smoothing adjustment coefficient, and multiply the difference between the current pot temperature and the system's preset maximum safe temperature by the smoothing adjustment coefficient to obtain the heat adjustment amount. :
[0059] in, The amount of heat adjustment (unit: kW); This represents the sigmoid sensitivity (typical value 0.2). To optimize the state vector (obtained from self-learning); Maximum safe temperature; First, calculate the current overall cooking state vector. With optimization of state vector The difference between the ideal state and the desired state, obtained through vector operations, reflects the degree of deviation of the current food's doneness, color, thermal efficiency, and other aspects from the ideal state. This state difference is then processed by a sigmoid function. This function uses a sensitivity coefficient... The system adjusts its response strength to state discrepancies, outputting a smoothing coefficient between 0 and 1. When the state is severely unsatisfactory, the coefficient approaches 1, indicating that strong intervention is needed; when the state is close to ideal, the coefficient approaches 0.5, indicating that only fine-tuning is required to maintain the desired state. The system then obtains the current temperature inside the pot. With the preset maximum safe temperature Calculate the difference between the two. This difference represents the theoretically permissible temperature rise while ensuring cooking safety and preventing burning. Multiplying the intelligent decision coefficient obtained in the second step by the safe adjustment range obtained in the third step yields the final heat adjustment amount. , A positive value indicates that firepower needs to be increased, a negative value indicates that firepower needs to be reduced, and a zero value indicates that the current firepower needs to be maintained.
[0060] Based on the difference between the current integrated state vector and the optimized state vector, an innovative "dynamic fire control law" is used to calculate precise fire adjustment. The core of this control law is an S-shaped response function that smoothly maps state differences to fire changes and automatically "applies the brakes" when approaching safety limits. It can achieve fine control of kilowatt-level firepower at the ±0.1kW level, much like an experienced chef adjusting a gas valve by feel. Step 304: Based on the calculated heat adjustment and preset stirring angle, generate the current stirring action sequence. This transforms complex cooking strategies into strictly executable digital commands, eliminating ambiguity and ensuring precise timing coordination between heat changes and stirring actions. For example, "start stirring while increasing heat" to maximize heat transfer efficiency. The system uses the current comprehensive cooking state vector... Based on factors such as ingredient distribution and viscosity, as well as cooking experience, we can plan the optimal stirring angle and rhythm to ensure that the ingredients are heated evenly and do not spill out of the pan.
[0061] To improve the overall state vector To enhance the expressive power of the system, a multimodal sensor data fusion and physical-data hybrid modeling approach is adopted, combined with innovative ideas of computer vision and thermodynamic modeling, to improve the perception of food status and temperature field inside the pot.
[0062] Traditional state vector The current model only includes scalar data such as temperature and humidity, lacking modeling of the spatial distribution of the temperature field inside the pot. This paper proposes a physical-information neural network that combines thermodynamic partial differential equations (PDEs) with deep learning to model the two-dimensional temperature field inside the pot. Explicit modeling is performed, and the state representation is dynamically updated by combining the food distribution information captured by visual sensors. The formula for modeling the temperature field is as follows:
[0063] Position inside the pot Time spent Temperature; Thermal diffusivity, which characterizes the thermal conductivity of cookware materials; The Laplace operator for the temperature field, representing the space-time thermal conduction effect; The heating source is located at... Heat input at the location; Heat input conversion efficiency; Heat loss due to the evaporation of moisture from food ingredients; : Coefficient of heat loss due to water evaporation; By training with PINN and combining discretized temperature sensor data with food distribution information in visual images, the above PDE is approximately solved to generate a high-resolution temperature field state. and its visual characteristics (e.g., changes in the color and shape of ingredients) and humidity characteristics The states are merged to form an optimized integrated state vector. .
[0064] The distribution of ingredients inside the pot is extracted from the visual image and mapped onto a temperature field grid. and By defining the boundary conditions, this visual-physical coupling method not only improves the accuracy of state representation but also enhances the system's ability to adapt to new ingredients and cookware.
[0065] Specifically, step 4 includes the following steps: Step 401: Based on the stir-frying action sequence, drive the actuator to perform corresponding heat adjustment and stir-frying actions; simultaneously, monitor and acquire the actual temperature value during the execution process in real time through sensors. Compared to the actual angle of stir-frying The system drives the robotic arm and stove to precisely execute the "stir-frying angle" and "heat adjustment" instructions issued in step 3. At the same time, the sensors immediately start up, recording the actual temperature, actual stir-frying angle and other data in real time, just like a supervisor. This ensures that the intelligent decision is faithfully and accurately translated into physical actions. This is the "execution starting point" of the intelligent closed loop. Synchronous monitoring creates a complete data chain from "instruction" to "result". Any execution deviation will be recorded immediately.
[0066] Step 402: Compare the actual temperature value and actual stirring angle obtained from monitoring with the expected target temperature and target angle in the control command sequence, and calculate a comprehensive execution deviation based on the difference between the two; this provides a measurable and comparable "report card" for the execution effect of each execution. This method replaces the vague subjective judgment of approximate completion, and can clearly distinguish whether it is due to improper decision-making (problems with the command itself) or "poor execution" (insufficient mechanical precision), making the direction of subsequent optimization clear.
[0067] The formula for calculating the execution deviation is:
[0068] Symbol Explanation For deviation; Angle weight (typical value 0.1); The expected angle for stirring as instructed; This refers to the actual angle at which the food is stir-fried. Monitor actual temperature values; : The expected target temperature value.
[0069] First, for all temperature measurement points inside the pot (e.g., a 4x4 sensor array with 16 points), the difference between the actual temperature and the target temperature is calculated one by one. Then, each difference is squared, and finally, all squared values are summed. This ensures that the deviation at each point is non-negative, while amplifying the impact of larger deviations (e.g., a 5°C deviation contributes 25, while a 2°C deviation only contributes 4), making it more sensitive to significant temperature runaways. This allows for a comprehensive assessment of the temperature control uniformity across the entire pot surface, rather than just the quality of a single point. Next, the absolute value of the difference between the expected and actual stirring angles is calculated. Then multiply by an angle weighting coefficient. Next, the "sum of squared temperature deviations" obtained in the first stage is added to the "weighted angle deviations" obtained in the second stage, and then the square root of the sum is taken to obtain a single, intuitive, and stable comprehensive deviation. This deviation accurately reflects the overall degree of deviation between the execution of this action and the expected goal.
[0070] Step 403: After the control command is executed, the image, temperature and humidity data inside the pot are collected again by the multimodal sensor, and the pot state vector after execution is generated based on this data; the execution deviation is combined with the pot state vector after execution to construct a feedback result vector. Specifically, step 5 includes the following steps: Step 501: The feedback result vector includes the execution deviation degree and the in-pot state vector after execution. Based on the difference between the in-pot state vector after execution and the optimized state vector, the loss of the decision model is calculated. A detailed comparison is made between the "actual dish state" in the report and the "ideal dish state" recorded internally. This can clearly indicate whether the problem is due to inaccurate heat, inadequate stirring, or a deviation in the decision itself, providing a clear direction for optimization.
[0071] Step 502: Based on the difference results obtained from the analysis, update the internal parameters of the decision model in step 301, and calculate the adjustment amount of the internal parameters of the decision model through a self-learning optimization algorithm. Based on the gaps found in the review, the system's self-learning algorithm will accurately calculate how much and in what direction the internal parameters of the "AI decision brain" (the model in step 3) need to be adjusted. Instead of crudely saying "make it hotter next time," it will refine and optimize the decision logic in a data-driven manner, just like adjusting a precision instrument, so that the system's "cooking intuition" becomes more and more accurate and reliable with the increase of usage.
[0072] Step 503: Based on the executed in-pot state vector and user historical preference data, iteratively update the optimized state vector; apply the calculated optimization scheme to the AI decision-making model to complete the upgrade of the "brain". Simultaneously, based on long-term feedback, dynamically update the "optimized state vector state" it pursues, i.e., the "ideal dish" standard that better suits user tastes. Step 504: Apply the updated AI decision-making model and the new optimized target state to the next round of cooking decision-making cycle. Immediately incorporate the updated "brain" and "standards" into the next cooking decision, forming a complete "learning-application" closed loop. This completes the fully automated closed loop of "perception → decision → execution → evaluation → learning," making the system a continuously self-improving intelligent agent. Specifically, step 601: Receive cooking preference information input by the user and convert the preference information into optimization parameters that the system can recognize, in order to adjust the optimization state vector. Users can conveniently input preference commands such as "less oil," "more spicy," and "soft and tender" via a mobile app, voice, or touchscreen. These natural language commands are accurately translated into machine-understandable optimization parameters. Without complex settings, users can customize the flavor and health standards of dishes in their most familiar way, ensuring that the user's personalized requirements are accurately translated into optimization goals.
[0073] Step 602: During the cooking process, based on the multimodal fusion matrix and the comprehensive state vector, provide the user with real-time feedback information including cooking progress and ingredient status; during the cooking process, display to the user in real time through the interface: current heat level, estimated remaining time, ingredient status identified by AI (such as "reducing sauce", "chicken has been colored") and the logic of system decision-making (such as "increasing heat because the temperature is detected to be too low").
[0074] Step 603: After cooking is complete, the system receives sensory evaluation data from the user regarding the dish and preprocesses this data into standardized user feedback signals. Once the dish is finished, the system invites the user to provide a simple rating (e.g., a five-star system) or select a tag (e.g., "just right," "too salty," "could be more tender next time"). This one-click feedback from the user is a valuable "teaching signal" for the system. The system transforms the user's subjective taste perception ("delicious" or "too salty") into objective data that can be processed and learned by the system.
[0075] Step 604: Input the standardized user feedback signal into the self-learning module to drive the parameters of the decision model and the optimized state vector to be updated in a personalized manner to adapt to the user's long-term taste preferences. Deeply integrate all the user's current and historical preference instructions and evaluation feedback to drive the self-learning module in step 5 to perform targeted optimization. This means that the system's "brain" and "gold standard" will continue to evolve towards a taste database unique to the user.
[0076] The innovation of this invention lies in proposing a dynamic fusion weight calculation formula based on energy gradient and humidity inside the pot. This allows the fusion weights of image, temperature, and humidity data to be adjusted in real time during the cooking process. Furthermore, it introduces the concept of "initial ecological entropy balance weight" based on information entropy, and combines spatial attention mechanisms and temporal recursive networks to generate "spatiotemporal adjustment weights," achieving refined and adaptive adjustment of data fusion in both spatial and temporal dimensions. This significantly improves the system's perception accuracy and robustness in complex cooking environments, making multimodal data fusion more closely resemble real physicochemical changes. In addition to analyzing the visual state of the ingredients, it also calculates the heat conduction index through a "heat conduction-humidity coupling model," fusing the two to generate a "comprehensive state vector" for decision-making. This vector simultaneously encodes the biochemical changes of the ingredients and the physical thermodynamic state inside the pot, providing a more comprehensive and scientific state description foundation for subsequent intelligent decision-making, enabling the decision model to simultaneously consider the ripeness of the ingredients and the efficiency of heat transfer.
[0077] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style of the specification is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other implementations that can be understood by those skilled in the art.
[0078] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. An artificial intelligence-based automatic cooking method of a cooking robot, characterized by, The method comprises the following steps: Step 1: synchronously collecting image, temperature and humidity data in the pot by a multi-modal sensor, and pre-processing and dynamically weighting the data to generate a multi-modal fusion matrix; Step 2: analyzing the visual state of the food and the thermodynamic state in the pot based on the multi-modal fusion matrix, and fusing to generate a comprehensive state vector for decision-making; Step 3: inputting the comprehensive state vector into a decision-making model based on machine learning to generate a control instruction sequence containing a fire adjustment amount and a stir-frying action; Step 4: executing the control instruction sequence, monitoring the execution process and collecting post-execution data to generate a feedback result vector containing execution deviation and post-execution state; Step 5: based on the feedback result vector, stably updating the parameters of the decision-making model and the internal optimization target state; Step 6: receiving user preference input and evaluation feedback, and based on the preference and feedback information, individualizing updating the decision-making model and the optimization target state.
2. The automatic cooking method of the cooking robot based on artificial intelligence according to claim 1, characterized in that, The step 1 comprises the following steps: Step 101: acquiring cooking data in the pot by a multi-modal sensor, including pot image data, pot temperature data and pot humidity data; Step 102: pre-processing the acquired pot image data, pot temperature data and pot humidity data respectively; Step 103: calculating dynamic fusion weights for the pre-processed pot image data, pot temperature data and pot humidity data respectively; Step 104: weighting and fusing the image feature data, temperature feature data and humidity feature data using the calculated dynamic fusion weights to generate a multi-modal fusion matrix.
3. The automatic cooking method of the cooking robot based on artificial intelligence according to claim 2, characterized in that, The calculation formula of the dynamic fusion weight is: wherein, is a dynamic fusion weight of the modal ; : modal index; : modal index; : image modal; : temperature modal; : humidity modal; is an energy gradient of the modal ; is an energy gradient of the modal ; is a bio-thermal sensitivity coefficient of the modal ; is a bio-thermal sensitivity coefficient of the modal ; is a current time in-pot humidity value.
4. The automatic cooking method of the cooking robot based on artificial intelligence according to claim 3, characterized in that, The step 2 comprises the following steps: Step 201: extracting and decoding visual related information from the multi-modal fusion matrix, inputting into a food state analysis model, identifying and outputting an intermediate state feature vector of the food; Step 202: Extract and decode temperature and humidity related information from the multi-modal fusion matrix, calculate the heat conduction index based on the heat conduction-humidity coupling model : wherein, is a thermal conduction index; is a thermal conductivity; is a temperature change rate; is a humidity correction factor; is an initial humidity; is a current time in-pot humidity value; is a current time in-pot temperature value; Step 203: weighting and fusing the intermediate state feature vector of the food and the heat conduction index to generate a comprehensive state vector for decision-making.
5. The automatic cooking method of the cooking robot based on artificial intelligence according to claim 4, characterized in that, The step 3 comprises the following steps: Step 301: inputting the comprehensive state vector into a decision-making model based on offline reinforcement learning to predict and output the optimal cooking action under the current state; Step 302: based on the predicted optimal cooking action, comparing the comprehensive state vector with an optimization state vector provided by a self-learning module to obtain a state deviation; Step 303: inputting the state deviation into a Sigmoid function to map it into a smooth adjustment coefficient, multiplying the difference between the current pot temperature and the maximum safe temperature preset by the system by the smooth adjustment coefficient to obtain a fire adjustment amount; Step 304: based on the calculated fire adjustment amount and a preset stir-frying angle, generating a control instruction sequence at the current time.
6. The automatic cooking method of a cooking robot based on artificial intelligence according to claim 5, wherein, The step 4 comprises the following steps: Step 401: driving an execution mechanism to execute corresponding fire adjustment and stir-frying action according to the control instruction sequence, and monitoring and acquiring actual temperature value and actual stir-frying angle in the execution process in real time through a sensor; Step 402: Based on the obtained actual temperature value and actual stirring angle, the preset target temperature and target stirring angle in the control instruction sequence are compared, and a comprehensive execution deviation degree is calculated based on the difference between the two; Step 403: After the control instruction is executed, the image, temperature and humidity data in the pot are collected again through the multi-modal sensor, and the post-execution pot state vector is generated based on this; the execution deviation degree and the post-execution pot state vector are combined to construct a feedback result vector.
7. The automatic cooking method of a cooking robot based on artificial intelligence according to claim 6, characterized in that, The step 5 includes the following steps: Step 501: Based on the difference between the pot state vector in the feedback result vector and the optimized state vector, the loss of the decision model is calculated; Step 502: According to the calculated loss, the internal parameters of the decision model in step 301 are updated; Step 503: Based on the post-execution pot state vector and user historical preference data, the optimized state vector is iteratively updated for stability; Step 504: The updated decision model parameters and the optimized state vector are applied to the next round of cooking decision process.
8. The automatic cooking method of the cooking robot based on artificial intelligence according to claim 7, characterized in that, The step 6 includes the following steps: Step 601: Receive user input cooking preference information, and convert the preference information into system recognizable optimization parameters to adjust the optimized state vector; Step 602: During the cooking process, based on the multi-modal fusion matrix and the comprehensive state vector, real-time feedback information including cooking progress and food material state is provided to the user; Step 603: After cooking is completed, the user's sensory evaluation data of the dish based on real-time feedback information is received, and the evaluation data is preprocessed into standardized user feedback signals; Step 604: The optimization parameters and user feedback signals are input into the self-learning module to update the parameters of the decision model and the optimized state vector. 9.The automatic cooking method of a cooking robot based on artificial intelligence according to claim 3, wherein, Based on the step 103, information entropy is introduced to generate an initial ecological entropy balance weight: wherein, : modalities of initial eco-entropy balance weight; : image modalities; : temperature modalities; : humidity modalities; : energy gradient of modalities ; : the pot humidity value at the current time ; : eco-sensitive coefficient of modalities ; : information entropy of modalities ; : entropy balance adjustment parameter.
10. The automatic cooking method of a cooking robot based on artificial intelligence according to claim 9, wherein, The initial ecological entropy balance weight is introduced into a spatial attention mechanism and a time sequence recurrent network to generate a space-time adjustment weight: wherein, : modalities in spatial locations and current time instances spatio-temporal adjustment weights; : spatial attention function, computed by a convolution operation on the temperature field local importance of the features : temporal recurrent network, capturing dynamic changes of weights over time; : Normalization function.