Multi-task cooking navigation method and device and electronic equipment
By acquiring visual and temporal features combined with historical behavioral features and using a multi-task cooking network for fusion processing, the problem of insufficient perception of food status in traditional cooking navigation devices is solved, achieving intelligent cooking navigation with higher precision and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NINGBO FOTILE KITCHEN WARE CO LTD
- Filing Date
- 2026-01-07
- Publication Date
- 2026-05-12
AI Technical Summary
Traditional cooking navigation devices lack the ability to sense the state of ingredients, resulting in a large discrepancy between cooking instructions and actual operation, which affects the user experience.
By acquiring visual modal features and temporal modal features, and combining them with target task identifiers and historical behavior features, a multi-task cooking network is used for fusion processing to predict cooking results and provide navigation.
It improves the precision and accuracy of intelligent cooking navigation in various cooking scenarios, enhancing the accuracy and safety of the cooking process.
Smart Images

Figure CN122018349A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cooking navigation technology, and in particular to a multi-tasking cooking navigation method, device, and electronic device. Background Technology
[0002] As the demand for smart kitchen appliances continues to grow, the application of smart cooking devices that can assist users in cooking has developed rapidly. Traditional cooking navigation relies heavily on fixed recipe steps and has a weak ability to perceive the state of ingredients (such as changes in temperature, doneness, and weight), resulting in a large discrepancy between guidance and actual operation, poor smart cooking effects, and a negative impact on user experience. Summary of the Invention
[0003] To address the aforementioned problems in the prior art, this invention discloses a multi-task cooking navigation method, device, and electronic device, which can improve the precision and accuracy of intelligent cooking navigation in various cooking scenarios. The technical solution disclosed in this invention is as follows: According to one aspect of the embodiments disclosed in this invention, a multi-tasking cooking navigation method is provided, comprising: The system acquires a target task identifier, visual modal features and temporal modal features during the cooking process, and historical behavioral features of the target user. The visual modal features include food image features and temperature image features of the cooking environment. The temporal modal features include gas concentration temporal features, cookware weight temporal features, temperature temporal features, and stove firepower temporal features. The target task identifier is used to indicate the target cooking task in at least two cooking tasks. Based on the target task identifier, the visual modal features and the temporal modal features are fused to obtain a first fused feature; The first fused feature and the historical behavior feature are input into the cooking multi-task network to obtain the predicted cooking result of the target cooking task; Cooking navigation is performed based on the predicted cooking results.
[0004] Optionally, the step of fusing the visual modal features and the temporal modal features based on the target task identifier to obtain the first fused feature includes: When the target task identifier indicates that the target cooking task is a cooking step reminder task, the weight information corresponding to the visual modality feature and the temporal modality feature is determined, and the weight information represents the importance of the visual modality feature and the temporal modality feature respectively; Based on the weight information corresponding to the visual modality feature and the temporal modality feature respectively, the visual modality feature and the temporal modality feature are weighted and summed to obtain the first fused feature.
[0005] Optionally, the step of fusing the visual modal features and the temporal modal features based on the target task identifier to obtain the first fused feature includes: When the target task identifier indicates that the target cooking task is a cooking warning task, the visual modal feature and the temporal modal feature are spliced together to obtain the first fused feature.
[0006] Optionally, the step of fusing the visual modal features and the temporal modal features based on the target task identifier to obtain the first fused feature includes: When the target task identifier indicates that the target cooking task is a safety alarm task, the target image features corresponding to the safety alarm task are obtained by filtering from the food image features and the temperature image features, and the target time sequence features corresponding to the safety alarm task are obtained by filtering from the gas concentration time sequence features, the weight time sequence features, the temperature time sequence features and the firepower time sequence features; The target image features and the target temporal features are fused to obtain the first fused feature.
[0007] Optionally, the cooking multitasking network includes at least two gate networks and at least two cooking result prediction networks, wherein the gate networks and the cooking result prediction networks correspond one-to-one, and each gate network corresponds to one cooking task; The step of inputting the first fused feature and the historical behavior feature into the cooking multi-task network to obtain the predicted cooking result of the target cooking task includes: The first fusion feature and the historical behavior feature are input into the gate network corresponding to the target task identifier for feature fusion processing to obtain the weight information corresponding to the first fusion feature and the historical behavior feature respectively. Based on the weight information corresponding to the first fusion feature and the historical behavior feature respectively, the first fusion feature and the historical behavior feature are weighted and summed to obtain the second fusion feature. The second fusion feature is input into the cooking result prediction network corresponding to the gate network for cooking result prediction processing to obtain the predicted cooking result of the target cooking task.
[0008] Optionally, acquiring the visual modal features and temporal modal features during the cooking process includes: During the cooking process, images of the target ingredients and the target temperature of the cooking environment are acquired, along with sequences of gas concentration information, weight information, temperature information, and firepower information of the stove. The visual modal features are obtained by stitching together the target food image and the target temperature image and extracting spatial features. The time-series modal features are obtained by splicing together and extracting time-series features from the gas concentration information sequence, the weight information sequence, the temperature information sequence, and the firepower information sequence.
[0009] Optionally, the step of concatenating and extracting time-series features from the gas concentration information sequence, the weight information sequence, the temperature information sequence, and the firepower information sequence to obtain the time-series modal features includes: Acquire human sensor signal sequence, touch screen signal sequence, and vibration information sequence and pot inspection information sequence of the cookware; The time-series modal features are obtained by splicing and extracting time-series features from the gas concentration information sequence, the weight information sequence, the temperature information sequence, the firepower information sequence, the human sensor signal sequence, the touch screen signal sequence, the vibration information sequence, and the boiler inspection information sequence.
[0010] Optionally, acquiring the target ingredient image and the target temperature image of the cooking environment, as well as the gas concentration information sequence, the cookware weight information sequence, the temperature information sequence, and the stove firepower information sequence includes: Acquire initial food images and initial temperature images of the cooking environment at multiple preset times, as well as initial gas concentration information, initial weight information of the cookware, initial temperature information, and initial firepower information of the stove. Based on each preset time, the initial food image, the initial temperature image, the gas concentration information, the weight information, the temperature information, and the firepower information are aligned to obtain the target food image and the target temperature image, as well as the target gas concentration information, the target weight information, the target temperature information, and the target firepower information. The target gas concentration information, target weight information, target temperature information, and target firepower information at the multiple preset times are sorted by time to obtain the gas concentration information sequence, the weight information sequence, the temperature information sequence, and the firepower information sequence.
[0011] According to another aspect of the disclosed embodiments of the present invention, a multi-tasking cooking navigation device is provided, comprising: The acquisition module is used to acquire the target task identifier, visual modal features and temporal modal features during the cooking process, and the historical behavior features of the target user; the visual modal features include food image features and temperature image features of the cooking environment, the temporal modal features include gas concentration temporal features, cookware weight temporal features, temperature temporal features, and stove firepower temporal features, and the target task identifier is used to indicate the target cooking task in at least two cooking tasks; The fusion module is used to fuse the visual modal features and the temporal modal features based on the target task identifier to obtain a first fused feature; The prediction module is used to input the first fused feature and the historical behavior feature into the cooking multi-task network to obtain the predicted cooking result of the target cooking task; A navigation module is used for cooking navigation based on the predicted cooking results.
[0012] According to another aspect of the disclosed embodiments of the present invention, an electronic device for multitasking cooking navigation is provided, including a processor and a memory, wherein the memory stores at least one instruction, the at least one instruction being loaded and executed by the processor to implement the multitasking cooking navigation method described in any of the preceding claims.
[0013] According to another aspect of the disclosed embodiments of the present invention, a computer-readable storage medium is provided, wherein at least one instruction is stored therein, the at least one instruction being loaded and executed by a processor to implement the multitasking cooking navigation method described in any of the preceding claims.
[0014] According to another aspect of the disclosed embodiments of the present invention, a computer program product containing instructions is provided that, when run on a computer, causes the computer to perform the multitasking cooking navigation method described in any of the above embodiments of the present invention.
[0015] The multi-task cooking navigation method provided by this invention has the following technical effects: This invention acquires a target task identifier, visual modal features and temporal modal features during the cooking process, and historical behavioral features of the target user. The visual modal features include food image features and temperature image features of the cooking environment; the temporal modal features include gas concentration temporal features, cookware weight temporal features, temperature temporal features, and stove firepower temporal features. The target task identifier is used to indicate the target cooking task among at least two cooking tasks. Based on the target task identifier, the visual modal features and temporal modal features are fused to obtain a first fused feature. The first fused feature and historical behavioral features are input into a multi-task cooking network to obtain a predicted cooking result for the target cooking task. Cooking navigation is performed based on the predicted cooking result, thereby improving the accuracy and precision of intelligent cooking navigation in various cooking scenarios.
[0016] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating a multi-tasking cooking navigation method according to an exemplary embodiment; Figure 2 This is a schematic diagram illustrating a process for acquiring visual modal features and temporal modal features during the cooking process, according to an exemplary embodiment. Figure 3 This is a schematic diagram illustrating a process for fusing the visual modal features and the temporal modal features according to an exemplary embodiment; Figure 4 This is a schematic diagram illustrating another process for fusing the visual modal features and the temporal modal features according to an exemplary embodiment; Figure 5 This is a schematic diagram illustrating a multi-tasking cooking navigation method according to an exemplary embodiment; Figure 6 This is a block diagram illustrating a multitasking cooking navigation device according to an exemplary embodiment; Figure 7 This is a block diagram illustrating a terminal electronic device for multitasking cooking navigation according to an exemplary embodiment; Figure 8 This is a block diagram illustrating a server electronic device for multitasking cooking navigation according to an exemplary embodiment. Detailed Implementation
[0019] To enable those skilled in the art to better understand the technical solutions disclosed in this invention, the technical solutions in the disclosed embodiments will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0020] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention disclosed herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0021] The following describes a multi-tasking cooking navigation method according to this application. Please refer to [link / reference]. Figure 1 , Figure 1 This is a flowchart illustrating a multi-tasking cooking navigation method according to an exemplary embodiment. This specification provides the operational steps of the method as described in the embodiments or flowcharts, but based on conventional or non-inventive labor, more or fewer operational steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only execution order. In actual system or server products, the methods shown in the embodiments or drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment). Specifically, as... Figure 1 As shown, the above method may include: S101: Obtain the target task identifier, visual modal features and temporal modal features during the cooking process, and historical behavioral features of the target user.
[0022] In one specific embodiment, visual modal features may include food image features and temperature image features of the cooking environment, while temporal modal features may include gas concentration temporal features, cookware weight temporal features, temperature temporal features, and stove heat temporal features. The target user's historical behavioral features can be the user's cooking behavior characteristics during past cooking processes, which can be used to characterize the user's cooking habits and personalize cooking prompts during cooking navigation, such as a preference for high-heat stir-frying at specific times, and personalized cooking prompt times, frequencies, and content.
[0023] In one specific embodiment, a target task identifier can be used to indicate a target cooking task among at least two cooking tasks. One task identifier can correspond to one cooking task, and the corresponding cooking task can be determined based on the task identifier. Cooking tasks can include cooking step reminder tasks, cooking warning tasks, and safety alarm tasks, etc. Specifically, cooking step reminder tasks can be tasks that remind users of cooking steps, such as "heating the pan is complete" or "heating the oil is complete" (for frying, the oil temperature has reached the appropriate level), "adding vegetables" reminder, and "flipping" reminder for frying, etc. Cooking warning tasks are tasks that warn users of abnormalities during the cooking process, such as "dry burning" reminder and "high temperature cooking" reminder. Safety alarm tasks are tasks that provide safety alarms for abnormalities in the cooking environment, such as "gas leak alarm" and "fire alarm".
[0024] In an optional embodiment, such as Figure 2 As shown, obtaining the visual modal features and temporal modal features during the cooking process can include: S201: During the cooking process, acquire images of the target ingredients and the target temperature of the cooking environment, as well as gas concentration information sequences, cookware weight information sequences, temperature information sequences, and stove firepower information sequences.
[0025] Optionally, the acquisition of the target ingredient image and the target temperature image of the cooking environment, as well as the gas concentration information sequence, the cookware weight information sequence, the temperature information sequence, and the stove firepower information sequence may include: Acquire initial food images and initial temperature images of the cooking environment at multiple preset times, as well as initial gas concentration information, initial weight information of the cookware, initial temperature information, and initial firepower information of the stove. Based on each preset time, the initial food image, initial temperature image, gas concentration information, weight information, temperature information and firepower information are aligned to obtain the target food image and target temperature image, as well as the target gas concentration information, target weight information, target temperature information and target firepower information; The target gas concentration information, target weight information, target temperature information, and target firepower information at multiple preset times are sorted by time to obtain the gas concentration information sequence, weight information sequence, temperature information sequence, and firepower information sequence.
[0026] In one specific embodiment, the aforementioned information can be collected by setting up sensors. Specifically, the initial food image can be an RGB image of the food, which can be collected by a camera installed on the range hood; the initial temperature image of the cooking environment can be an infrared temperature image, which can be collected by an infrared temperature sensor installed on the range hood; the initial gas concentration information can be collected by a gas concentration sensor; the initial weight information of the cookware can be collected by a pressure sensor installed on the cookware; the initial temperature information can be the temperature information of the bottom of the pot, which can be collected by a temperature sensor installed on the bottom of the pot; and the initial firepower information of the stove can be determined by an angle sensor or an infrared sensor installed on the stove knob to determine the firepower level.
[0027] In multimodal data, different modalities may have different temporal resolutions, sampling rates, or time points. To fuse them together, time alignment is required. In practical applications, after acquiring the aforementioned information, heterogeneous data synchronization processing can be performed. Timestamp alignment techniques can be used to integrate multiple modal data such as temperature, image, weight, and gas data, aligning them and eliminating single-sensor errors. Specifically, after acquiring the aforementioned information, filtering (such as Kalman filtering) can be performed to remove noise and normalize the data. Based on the data's timestamp or sampling rate, time information is labeled for each modality. Appropriate strategies can be selected to achieve time alignment, such as interpolation, alignment to the nearest timestamp, or alignment to a unified time grid. After time alignment, features can be extracted from each modality's data for subsequent fusion or analysis.
[0028] S203: The target food image and the target temperature image are stitched together and spatial features are extracted to obtain visual modal features.
[0029] Specifically, convolutional neural networks (CNNs) can be used to extract spatial features from the target food image and the target temperature image to obtain visual modal features of a preset dimension, such as 1024 dimensions.
[0030] S205: The gas concentration information sequence, weight information sequence, temperature information sequence and firepower information sequence are spliced together and time series features are extracted to obtain time series modal features.
[0031] Specifically, the temporal dependencies of the above information sequences can be captured by long short time series (LSTM) to obtain temporal modal features of a preset dimension, such as 64 dimensions.
[0032] Optionally, the splicing and time-series feature extraction of the above-mentioned gas concentration information sequence, weight information sequence, temperature information sequence, and firepower information sequence to obtain the time-series modal features may include: Acquire human sensor signal sequences, touch screen signal sequences, as well as cookware vibration information sequences and cookware inspection information sequences; The gas concentration information sequence, weight information sequence, temperature information sequence, firepower information sequence, human sensor signal sequence, touch screen signal sequence, vibration information sequence, and boiler inspection information sequence are spliced together and time-series features are extracted to obtain time-series modal features.
[0033] In one specific embodiment, the aforementioned human sensing signal sequence, touch screen signal sequence, vibration information sequence, and pot inspection information sequence can also be obtained by temporally sorting the human sensing signals, touch screen signals, vibration information, and pot inspection information at multiple acquisition times after time alignment. Specifically, the human sensing signal can be used to sense human body information to detect whether there is a person in the current cooking environment, and the pot inspection information can be used to determine whether a pot is placed on the stove. Similarly, the temporal dependence of the above information sequences can be captured by Long Short-Time Series (LSTM) to obtain temporal modal features of a preset dimension.
[0034] S103: Based on the target task identifier, the visual modal features and temporal modal features are fused to obtain the first fused feature.
[0035] Specifically, depending on the cooking task, the appropriate feature fusion method can be selected to fuse visual modal features and temporal modal features to obtain the first fused feature.
[0036] In some embodiments, such as Figure 3 As shown, based on the target task identifier, visual modal features and temporal modal features are fused to obtain the first fused feature, which includes: S301: When the target task identifier indicates that the target cooking task is a cooking step reminder task, determine the weight information corresponding to the visual modal features and the temporal modal features respectively.
[0037] In one specific embodiment, weight information can characterize the relative importance of visual modal features and temporal modal features.
[0038] S303: Based on the weight information corresponding to the visual modal features and the temporal modal features respectively, the visual modal features and the temporal modal features are weighted and summed to obtain the first fusion feature.
[0039] The cooking step reminder task requires combining temporal and spatial features to detect cooking steps and predict reminder content. Therefore, it needs to focus on the cooking scenario and perform weighted fusion of visual modality features and temporal modality features. Optionally, the first fusion feature can be determined based on an attention network. The visual modality features and temporal modality features are input into the attention network to obtain the weight information corresponding to each of the visual modality features and temporal modality features, and the visual modality features and temporal modality features are weighted and summed accordingly to obtain the first fusion feature.
[0040] In some embodiments, the above-mentioned fusion processing of visual modal features and temporal modal features based on the target task identifier to obtain the first fused feature may include: When the target task identifier indicates that the target cooking task is a cooking warning task, the visual modal features and temporal modal features are spliced together to obtain the first fused feature.
[0041] In one specific embodiment, the cooking early warning task needs to detect cooking abnormalities in advance and issue an early warning. Therefore, feature concatenation and fusion are used to obtain more and more comprehensive information.
[0042] In some embodiments, such as Figure 4 As shown, the above-mentioned fusion processing of visual modal features and temporal modal features based on the target task identifier to obtain the first fused feature may include: S401: When the target task identifier indicates that the target cooking task is a safety alarm task, the target image features corresponding to the safety alarm task are selected from the food image features and temperature image features, and the target time sequence features corresponding to the safety alarm task are selected from the gas concentration time sequence features, weight time sequence features, temperature time sequence features and firepower time sequence features.
[0043] In one specific embodiment, the safety alarm task needs to focus on key information; therefore, key feature filtering is employed to extract key features from visual modal features and / or temporal modal features. For example, for alarms such as gas leaks and fires, the focus can be on monitoring infrared temperature images and gas concentration information.
[0044] S403: The target image features and target temporal features are fused to obtain the first fused feature.
[0045] In one specific embodiment, the key features extracted through screening can be subjected to fusion processing such as splicing and weighted fusion to obtain fused features.
[0046] S105: Input the first fusion feature and historical behavior feature into the cooking multi-task network to obtain the predicted cooking result of the target cooking task.
[0047] In one specific embodiment, the cooking multi-task network may include at least two gate networks and at least two cooking result prediction networks, with a one-to-one correspondence between the gate networks and the cooking result prediction networks, and each gate network corresponding to one cooking task. Specifically, the cooking result prediction network may be a tower network, with different tower networks set up based on different cooking task requirements to obtain corresponding predicted cooking results.
[0048] In an optional embodiment, inputting the first fused features and historical behavioral features into the cooking multi-task network to obtain the predicted cooking result for the target cooking task may include: The first fusion feature and the historical behavior feature are input into the gate network corresponding to the target task identifier for feature fusion processing to obtain the weight information corresponding to the first fusion feature and the historical behavior feature respectively. Based on the weight information corresponding to the first fusion feature and the historical behavior feature respectively, the first fusion feature and the historical behavior feature are weighted and summed to obtain the second fusion feature. The second fusion feature is input into the cooking result prediction network corresponding to the gate network for cooking result prediction processing to obtain the predicted cooking result of the target cooking task.
[0049] In one specific embodiment, predicting cooking results can include the content, frequency, and method of the cooking reminder. Inputting user historical behavior features into the gate network allows for adjusting the model's adaptive strategy based on user cooking habits. Through feedback from user historical behavior, it is possible to achieve adaptive reminder timing (earlier / delayed), personalized reminder content, dynamic threshold adjustment (reducing invalid reminders), and long-term habit learning (such as a preference for high-heat stir-frying on weekends).
[0050] Optionally, the cooking multi-task network can also include multiple expert networks. Each expert network can extract features from the input data. Each gate network is used to assign different weights to each expert network for different tasks. Based on the output features of each expert network and the output weights of the gate networks, a weighted sum is performed to obtain the feature fusion data (i.e., the second fusion feature) corresponding to each gate network, which is the input of the corresponding tower network. The task prediction result is obtained through the corresponding tower network. Optionally, the first fusion feature, historical behavior features, and the collected multimodal information can also be input into each gate network. Each gate network can extract features from the input data to obtain the probability of multiple expert networks being selected by each gate network. The probability of multiple expert networks being selected by each gate network is then normalized to obtain the weights of the multiple expert networks corresponding to each gate network.
[0051] S107: Cooking navigation based on predicted cooking results.
[0052] In one specific implementation, such as Figure 5 As shown, after collecting and aligning the aforementioned multimodal information, feature extraction and fusion are performed on visual modality information and temporal modality information respectively to obtain visual modality features and temporal modality features. Different feature fusion methods are selected according to the different cooking tasks. For cooking step reminder tasks, visual modality features and temporal modality features are weighted and fused based on an attention network. For cooking warning tasks, visual modality features and temporal modality features are concatenated. For safety alarm tasks, key features from visual modality features and / or temporal modality features are selected and fused to obtain fused features corresponding to different cooking tasks. The fused features, user historical behavior features, and multimodal information are weighted and fused based on a gate network, and then input into the corresponding tower network for result prediction to obtain the predicted cooking result. Cooking navigation is then performed based on the predicted cooking result.
[0053] As can be seen from the technical solutions provided in the embodiments of this specification above, this specification obtains a target task identifier, visual modal features and temporal modal features during the cooking process, and historical behavioral features of the target user; the visual modal features include food image features and temperature image features of the cooking environment, and the temporal modal features include gas concentration temporal features, cookware weight temporal features, temperature temporal features, and stove firepower temporal features; the target task identifier is used to indicate the target cooking task among at least two cooking tasks; based on the target task identifier, the visual modal features and temporal modal features are fused to obtain a first fused feature; the first fused feature and historical behavioral features are input into a cooking multi-task network to obtain a predicted cooking result for the target cooking task; cooking navigation is performed based on the predicted cooking result, thereby improving the accuracy and precision of intelligent cooking navigation in various cooking scenarios.
[0054] This invention also provides a multi-tasking cooking navigation device, such as... Figure 6 As shown, the device includes: The acquisition module 610 is used to acquire the target task identifier, visual modal features and temporal modal features during the cooking process, and the historical behavior features of the target user; the visual modal features include food image features and temperature image features of the cooking environment, the temporal modal features include gas concentration temporal features, cookware weight temporal features, temperature temporal features, and stove firepower temporal features, and the target task identifier is used to indicate the target cooking task in at least two cooking tasks; The fusion module 620 is used to fuse the visual modal features and the temporal modal features based on the target task identifier to obtain a first fused feature; The prediction module 630 is used to input the first fused feature and the historical behavior feature into the cooking multi-task network to obtain the predicted cooking result of the target cooking task; Navigation module 640 is used for cooking navigation based on the predicted cooking results.
[0055] Optionally, the fusion module 620 includes: The first weight determination unit is used to determine the weight information corresponding to the visual modal feature and the temporal modal feature respectively when the target task identifier indicates that the target cooking task is a cooking step reminder task. The weight information represents the importance of the visual modal feature and the temporal modal feature respectively. The first feature fusion unit is used to perform a weighted summation of the visual modal features and the temporal modal features based on the weight information corresponding to the visual modal features and the temporal modal features respectively, to obtain the first fused feature.
[0056] Optionally, the fusion module 620 includes: The second feature fusion unit is used to concatenate the visual modal features and the temporal modal features to obtain the first fused feature when the target task identifier indicates that the target cooking task is a cooking warning task.
[0057] Optionally, the fusion module 620 includes: The filtering unit is configured to, when the target task identifier indicates that the target cooking task is a safety alarm task, filter out the target image features corresponding to the safety alarm task from the food image features and the temperature image features, and filter out the target time sequence features corresponding to the safety alarm task from the gas concentration time sequence features, the weight time sequence features, the temperature time sequence features and the firepower time sequence features; The third feature fusion unit is used to fuse the target image features and the target temporal features to obtain the first fused feature.
[0058] Optionally, the cooking multi-task network includes at least two gate networks and at least two cooking result prediction networks, wherein the gate networks and the cooking result prediction networks correspond one-to-one, and each gate network corresponds to one cooking task; the prediction module 630 includes: The fourth feature fusion unit is used to input the first fused feature and the historical behavior feature into the gate network corresponding to the target task identifier for feature fusion processing to obtain the weight information corresponding to the first fused feature and the historical behavior feature respectively, and to perform weighted summation processing on the first fused feature and the historical behavior feature based on the weight information corresponding to the first fused feature and the historical behavior feature respectively to obtain the second fused feature. The cooking result prediction unit is used to input the second fused feature into the cooking result prediction network corresponding to the gate network for cooking result prediction processing, so as to obtain the predicted cooking result of the target cooking task.
[0059] Optionally, the acquisition module 610 includes: The first acquisition unit is used to acquire, during the cooking process, an image of the target ingredient and an image of the target temperature of the cooking environment, as well as a sequence of gas concentration information, a sequence of weight information of the cookware, a sequence of temperature information, and a sequence of firepower information of the stove. A visual modality feature determination unit is used to stitch together the target food image and the target temperature image and extract spatial features to obtain the visual modality features; The first temporal modal feature determination unit is used to splice and extract temporal features from the gas concentration information sequence, the weight information sequence, the temperature information sequence and the firepower information sequence to obtain the temporal modal features.
[0060] Optionally, the time-series modal feature determination unit includes: The second acquisition unit is used to acquire the human sensor signal sequence, the touch screen signal sequence, and the vibration information sequence and pot inspection information sequence of the cookware; The second temporal modal feature determination unit is used to splice and extract temporal features from the gas concentration information sequence, the weight information sequence, the temperature information sequence, the firepower information sequence, the human sensor signal sequence, the touch screen signal sequence, the vibration information sequence, and the boiler inspection information sequence to obtain the temporal modal features.
[0061] Optionally, the first acquisition unit includes: The third acquisition unit is used to acquire initial food images and initial temperature images of the cooking environment at multiple preset times, as well as initial gas concentration information, initial weight information of the cookware, initial temperature information, and initial firepower information of the stove. An information alignment unit is used to align the initial food image, the initial temperature image, the gas concentration information, the weight information, the temperature information, and the firepower information based on each preset time, so as to obtain the target food image and the target temperature image, as well as the target gas concentration information, the target weight information, the target temperature information, and the target firepower information. The sorting unit is used to sort the target gas concentration information, target weight information, target temperature information and target firepower information at the multiple preset times by time, so as to obtain the gas concentration information sequence, the weight information sequence, the temperature information sequence and the firepower information sequence.
[0062] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0063] Figure 7 This is a block diagram illustrating a terminal electronic device for multitasking cooking navigation according to an exemplary embodiment. The electronic device can be a terminal, and its internal structure diagram can be as follows: Figure 7 As shown, the electronic device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a multi-tasking cooking navigation method. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse.
[0064] Figure 8 This is a block diagram of a server electronic device for multi-tasking cooking navigation according to an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as follows: Figure 8As shown, the electronic device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When executed by the processor, the computer program implements a multi-tasking cooking navigation method.
[0065] Those skilled in the art will understand that Figure 7 or Figure 8 The structures shown are merely block diagrams of some structures related to the disclosed solutions of this invention, and do not constitute a limitation on the electronic devices to which the disclosed solutions of this invention are applied. Specific electronic devices may include more or fewer components than those shown in the figures, or combine certain components, or have different component arrangements.
[0066] In an exemplary embodiment, an electronic device for multitasking cooking navigation is also provided, including a processor and a memory storing at least one instruction, which is loaded and executed by the processor to implement the multitasking cooking navigation method as disclosed in the present invention.
[0067] In an exemplary embodiment, a computer-readable storage medium is also provided, which stores at least one instruction, which is loaded and executed by a processor to implement the multitasking cooking navigation method disclosed in this invention.
[0068] In an exemplary embodiment, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform the multitasking cooking navigation method disclosed in this invention.
[0069] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0070] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles disclosed herein and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the following claims.
[0071] It should be understood that the present invention is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is limited only by the appended claims.
Claims
1. A multi-task cooking navigation method, characterized in that, The method includes: The system acquires a target task identifier, visual modal features and temporal modal features during the cooking process, and historical behavioral features of the target user. The visual modal features include food image features and temperature image features of the cooking environment. The temporal modal features include gas concentration temporal features, cookware weight temporal features, temperature temporal features, and stove firepower temporal features. The target task identifier is used to indicate the target cooking task in at least two cooking tasks. Based on the target task identifier, the visual modal features and the temporal modal features are fused to obtain a first fused feature; The first fused feature and the historical behavior feature are input into the cooking multi-task network to obtain the predicted cooking result of the target cooking task; Cooking navigation is performed based on the predicted cooking results.
2. The method according to claim 1, characterized in that, The step of fusing the visual modal features and the temporal modal features based on the target task identifier to obtain the first fused feature includes: When the target task identifier indicates that the target cooking task is a cooking step reminder task, the weight information corresponding to the visual modality feature and the temporal modality feature is determined, and the weight information represents the importance of the visual modality feature and the temporal modality feature respectively; Based on the weight information corresponding to the visual modality feature and the temporal modality feature respectively, the visual modality feature and the temporal modality feature are weighted and summed to obtain the first fused feature.
3. The method according to claim 1, characterized in that, The step of fusing the visual modal features and the temporal modal features based on the target task identifier to obtain the first fused feature includes: When the target task identifier indicates that the target cooking task is a cooking warning task, the visual modal feature and the temporal modal feature are spliced together to obtain the first fused feature.
4. The method according to claim 1, characterized in that, The step of fusing the visual modal features and the temporal modal features based on the target task identifier to obtain the first fused feature includes: When the target task identifier indicates that the target cooking task is a safety alarm task, the target image features corresponding to the safety alarm task are obtained by filtering from the food image features and the temperature image features, and the target time sequence features corresponding to the safety alarm task are obtained by filtering from the gas concentration time sequence features, the weight time sequence features, the temperature time sequence features and the firepower time sequence features; The target image features and the target temporal features are fused to obtain the first fused feature.
5. The method according to claim 1, characterized in that, The cooking multitasking network includes at least two gate networks and at least two cooking result prediction networks. The gate networks and the cooking result prediction networks correspond one-to-one, and each gate network corresponds to a cooking task. The step of inputting the first fused feature and the historical behavior feature into the cooking multi-task network to obtain the predicted cooking result of the target cooking task includes: The first fusion feature and the historical behavior feature are input into the gate network corresponding to the target task identifier for feature fusion processing to obtain the weight information corresponding to the first fusion feature and the historical behavior feature respectively. Based on the weight information corresponding to the first fusion feature and the historical behavior feature respectively, the first fusion feature and the historical behavior feature are weighted and summed to obtain the second fusion feature. The second fusion feature is input into the cooking result prediction network corresponding to the gate network for cooking result prediction processing to obtain the predicted cooking result of the target cooking task.
6. The method according to any one of claims 1 to 5, characterized in that, The acquisition of visual modal features and temporal modal features during the cooking process includes: During the cooking process, images of the target ingredients and the target temperature of the cooking environment are acquired, along with sequences of gas concentration information, weight information, temperature information, and firepower information of the stove. The visual modal features are obtained by stitching together the target food image and the target temperature image and extracting spatial features. The time-series modal features are obtained by splicing together and extracting time-series features from the gas concentration information sequence, the weight information sequence, the temperature information sequence, and the firepower information sequence.
7. The method according to claim 6, characterized in that, The process of concatenating and extracting time-series features from the gas concentration information sequence, the weight information sequence, the temperature information sequence, and the firepower information sequence to obtain the time-series modal features includes: Acquire human sensor signal sequence, touch screen signal sequence, and vibration information sequence and pot inspection information sequence of the cookware; The time-series modal features are obtained by splicing and extracting time-series features from the gas concentration information sequence, the weight information sequence, the temperature information sequence, the firepower information sequence, the human sensor signal sequence, the touch screen signal sequence, the vibration information sequence, and the boiler inspection information sequence.
8. The method according to claim 6, characterized in that, The acquisition of the target ingredient image and the target temperature image of the cooking environment, as well as the gas concentration information sequence, the cookware weight information sequence, the temperature information sequence, and the stove firepower information sequence, includes: Acquire initial food images and initial temperature images of the cooking environment at multiple preset times, as well as initial gas concentration information, initial weight information of the cookware, initial temperature information, and initial firepower information of the stove. Based on each preset time, the initial food image, the initial temperature image, the gas concentration information, the weight information, the temperature information, and the firepower information are aligned to obtain the target food image and the target temperature image, as well as the target gas concentration information, the target weight information, the target temperature information, and the target firepower information. The target gas concentration information, target weight information, target temperature information, and target firepower information at the multiple preset times are sorted by time to obtain the gas concentration information sequence, the weight information sequence, the temperature information sequence, and the firepower information sequence.
9. A multi-tasking cooking navigation device, characterized in that, The device includes: The acquisition module is used to acquire the target task identifier, visual modal features and temporal modal features during the cooking process, and the historical behavior features of the target user; the visual modal features include food image features and temperature image features of the cooking environment, the temporal modal features include gas concentration temporal features, cookware weight temporal features, temperature temporal features, and stove firepower temporal features, and the target task identifier is used to indicate the target cooking task in at least two cooking tasks; The fusion module is used to fuse the visual modal features and the temporal modal features based on the target task identifier to obtain a first fused feature; The prediction module is used to input the first fused feature and the historical behavior feature into the cooking multi-task network to obtain the predicted cooking result of the target cooking task; A navigation module is used for cooking navigation based on the predicted cooking results.
10. An electronic device for multi-tasking cooking navigation, characterized in that, The electronic device includes a processor and a memory, the memory storing at least one instruction, which is loaded and executed by the processor to implement the multitasking cooking navigation method as described in any one of claims 1 to 8.