Food model training method and device and electronic equipment
Through multiple iterative training combined with image data and spectral data, the problem of insufficient spectral data accuracy in food nutritional component testing of infrared spectrometers is solved, and accurate prediction of food nutritional components is achieved.
Patent Information
- Application Number
- CN202510502994.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-22
AI Technical Summary
In the prior art, when using infrared spectrometers to test food nutritional components, there is a problem of low accuracy of spectral data, resulting in inaccurate analysis of nutrient components.
A food model training method is adopted, through multiple iterative training, using image data and spectral data as training samples, the network parameters of the initial food model are updated until the loss value meets the convergence conditions, thereby obtaining the target food model.
By compensating for the deficiency of spectral data by image data, the target food model can accurately predict the nutritional content of food, for example, in pork, accurately identify the proportion of lean meat and fat meat.
Smart Images

Figure CN120014631A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of food testing, and in particular to a food model training method, device and electronic equipment. Background Art
[0002] Food nutritional testing is of great significance and value. It can help users understand the nutrients such as calories, protein, fat, carbohydrates, vitamins and minerals contained in the food, thereby assisting users in evaluating the potential impact of food on health.
[0003] In the related art, an infrared spectrometer is used to perform spectral scanning on food to obtain spectral data of the food, and the spectral data is used to analyze the nutritional components of the food. However, the nutritional components of the food obtained may be inaccurate. Summary of the invention
[0004] The purpose of the present invention is to provide a food model training method, device and electronic equipment to solve the above technical problems.
[0005] In order to achieve the above-mentioned object, a first aspect of an embodiment of the present invention provides a food model training method, the method comprising: using multiple groups of training samples to perform multiple rounds of iterative training on an initial food model to obtain a target food model; each group of training samples in the multiple groups of training samples includes image data and spectral data of the food, and each round of iterative training in the multiple rounds of iterative training includes the following steps: Inputting the image data and the spectral data into an initial food model to obtain predicted nutritional components of the food; Determining a loss value between the predicted nutritional content and the actual nutritional content of the food; updating the network parameters of the initial food model according to the loss value until the loss value satisfies a convergence condition; The initial food model when the loss value satisfies the convergence condition is used as the target food model.
[0006] Optionally, updating the network parameters of the initial food model according to the loss value includes: When the loss value does not satisfy the convergence condition, the network parameters of the initial food model are updated according to the loss value until the loss value satisfies the convergence condition.
[0007] Optionally, the training sample further includes text data, which is a description text of the food; and the step of inputting the image data and the spectral data into an initial food model to obtain predicted nutritional components of the food includes: The image data, the text data and the spectral data are input into the initial food model to obtain the predicted nutritional components of the food.
[0008] Optionally, inputting the image data, the text data and the spectral data into the initial food model to obtain the predicted nutritional components of the food includes: fusing the image data, the spectral data and the text data to obtain a first feature; The first feature is respectively fused with the image data, the spectrum data and the text data to obtain a first fused feature, a second fused feature and a third fused feature; The first fusion feature, the second fusion feature and the third fusion feature are input into the initial food model to obtain the predicted nutritional components of the food.
[0009] Optionally, the training samples include historical training samples and real-time training samples, the historical training samples include historical image data and spectral data, and the real-time training samples include real-time image data and spectral data; the step of inputting the image data and the spectral data into an initial food model to obtain predicted nutritional components of the food includes: The real-time image data and spectral data, as well as the historical image data and spectral data are input into the initial food model to obtain predicted nutritional components of the food.
[0010] Optionally, the method further comprises: Collecting single-angle images of the food at various angles; A surround image of the food is obtained according to the single-angle images of the food at various angles; the surround image contains the image data.
[0011] Optionally, inputting the image data and the spectral data into an initial food model to obtain predicted nutritional components of the food includes: A target area having a change value greater than a preset value in the surround image is extracted; the change value includes a change value in color and brightness.
[0012] The preset image features in the target area are used as image data in the surround view image.
[0013] Optionally, the preset image feature includes at least one of a color feature and a texture feature in the target area.
[0014] Through the above technical solution, when the accuracy of spectral data is low, the spectral data can be compensated by image data, and the target food model can also accurately predict the nutritional components of the food based on the image data. For example, taking pork as an example, the image data shows that the pork has more lean meat and less fat, so the corresponding nutritional components of the pork have more protein content and less fat content. When the spectral data cannot accurately identify the proportion of protein and fat, the image data can be used for compensation, and the target food model can accurately predict the nutritional components of the food.
[0015] Other features and advantages of the present invention will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the present invention but do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 It is a flowchart of the steps of a food model training method according to an exemplary embodiment.
[0017] Figure 2 It is a flowchart of the steps of a food model training method according to an exemplary embodiment.
[0018] Figure 3 It is a flowchart of the steps of a food model training method according to an exemplary embodiment.
[0019] Figure 4 It is a flowchart of the steps of a food model training method according to an exemplary embodiment.
[0020] Figure 5 It is a flowchart of the steps of a food model training method according to an exemplary embodiment.
[0021] Figure 6 The figure is a schematic diagram of the structure of an infrared spectrometer according to an exemplary embodiment.
[0022] Figure 7 It is a flowchart of the steps of a food model training method according to an exemplary embodiment.
[0023] Figure 8 is a block diagram of a food model training device according to an exemplary embodiment.
[0024] Fig. 9 It is a block diagram of a first electronic device according to an exemplary embodiment.
[0025] Fig.10is a block diagram of a second electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0026] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the present invention.
[0027] Figure 1 It is a flowchart of the steps of a food model training method proposed according to an exemplary embodiment. The food model training method can be applied to electronic devices such as infrared spectrometers.
[0028] The food model training method comprises: using multiple groups of training samples to perform multiple rounds of iterative training on the initial food model to obtain the target food model; each group of training samples in the multiple groups of training samples includes image data and spectral data of the food, and each round of iterative training in the multiple rounds of iterative training comprises the following steps: In step S10, the image data and the spectral data are input into an initial food model to obtain predicted nutritional components of the food.
[0029] The image data is image data related to the food to be tested. The image of the food to be tested can be collected by an image acquisition device. The image data can intuitively reflect the proportion of some nutrients in the food.
[0030] For example, the image data collected by the image acquisition device is image data related to pork. If the lean meat accounts for a larger proportion of pork, the protein content of the pork is higher; if the fat meat accounts for a larger proportion of pork, the fat content of the pork is higher. Therefore, the image data can indirectly reflect the proportion of nutrients, and the nutritional components of food can be obtained to a certain extent based on the image data.
[0031] Spectral data is the spectral data of food collected by an infrared spectrometer. The spectral data can reflect the physical structure and chemical composition of the food. Therefore, the nutritional components of the food can be analyzed based on the physical structure and chemical composition of the food.
[0032] The process of obtaining spectral data of food includes: collecting food, grinding, homogenizing and drying the food, etc., wherein grinding is to grind solid food into powder to improve the uniformity of food, homogenizing is to homogenize liquid food to eliminate the influence of particulate matter, and drying is to dry the food to remove the influence of moisture; then the processed food is filled into containers such as cuvettes and quartz plates; then the food filled in the container is placed in an infrared spectrometer for testing; before the test, the parameters of the infrared spectrometer are set, such as scanning range, scanning speed, spectral resolution, etc., and then the infrared spectrometer is run to perform spectral scanning on the food to obtain the spectral data of the food.
[0033] The initial food model is a model to be trained. The initial food model can be trained using multiple groups of training samples and labels corresponding to the multiple groups of training samples. Each group of training samples in the multiple groups of training samples contains image data and spectral data, and the labels corresponding to the multiple groups of training samples contain the actual nutritional components of the food.
[0034] The image data and the spectral data can be input into the initial food model to obtain the predicted nutritional components of the food, which are the nutritional components of the food that are predicted by the initial food model with low accuracy.
[0035] Among them, the nutritional components of food include carbohydrates, proteins, fats, vitamins, minerals, etc.
[0036] In step S20, the loss value between the predicted nutritional components and the actual nutritional components of the food is determined.
[0037] Optionally, the difference between the predicted nutrients and the actual nutrients can be used as the loss value between the predicted nutrients and the actual nutrients of the food. The larger the loss value, the greater the gap between the predicted nutrients predicted by the initial food model and the actual nutrients, and the lower the accuracy of the predicted nutrients predicted by the initial food model, and vice versa.
[0038] Among them, the predicted nutrients and actual nutrients can be substituted into loss functions such as the cross entropy loss function and the root mean square error to obtain the loss value, which will not be repeated here.
[0039] Among them, the actual nutritional content is the accurate nutritional content of the food obtained by multiple measurements of the food.
[0040] In step S30, the network parameters of the initial food model are updated according to the loss value until the loss value meets the convergence condition.
[0041] The training of the initial food model includes a forward propagation stage and a back-propagation stage. In the forward propagation stage, the initial food model includes multiple network layers such as an input layer, a hidden layer, and an output layer. Image data and spectral data can be input into the input layer. After calculation by the hidden layer, the predicted nutritional components are output through the output layer, and then the loss value between the predicted nutritional components and the actual nutritional components is calculated; in the back-propagation stage, starting from the output layer, the gradient of the loss value relative to the network parameters of each network layer is calculated, that is, the partial derivative of the loss value relative to each network layer is calculated, and then the gradient of the network parameters calculated is superimposed on the network parameters to obtain the updated network parameters of the initial food model to reduce the loss value.
[0042] For example, the image data and spectral data of the previous round of training samples can be input into the initial food model to obtain the predicted nutritional components of the food, and then the loss value between the predicted nutritional components and the actual nutritional components can be calculated. The loss value is used to reversely update the network parameters of each network layer in the initial food model to achieve the update of the initial food model. If the loss value of the previous round does not meet the convergence condition, the image data and spectral data of the next round of training samples are input into the updated initial food model to obtain the predicted nutritional components of the food again, and then the loss value between the predicted nutritional components and the actual nutritional components is recalculated. This iteration is repeated until the final loss value meets the convergence condition.
[0043] Among them, the loss value satisfies the convergence conditions including that the loss value is less than a preset value, the loss value no longer decreases, or the training of the initial food model obtains a preset number of iterations.
[0044] In step S40, the initial food model when the loss value satisfies the convergence condition is used as the target food model.
[0045] When the loss value meets the convergence condition, it means that the difference between the predicted nutrients output by the initial food model and the actual nutrients is small, the predicted nutrients output by the initial food model are close to the actual nutrients, or equal to the actual nutrients, and the accuracy of the predicted nutrients output by the initial food model is high. Therefore, the initial food model when the loss value meets the convergence condition can be used as the trained target food model.
[0046] When the target food model is subsequently applied, the collected image data related to the food and the spectral data of the food can be input into the target food model to obtain the accurate nutritional components of the food.
[0047] In the related art, an infrared spectrometer is used to scan the nutritional components of food. However, the nutritional components of food scanned by the infrared spectrometer are inaccurate to a certain extent for the following reasons: First, the nutritional composition of food is complex, and the infrared absorption spectra of different ingredients will overlap with each other, making the spectral data difficult to interpret and affecting the accuracy of the calculated nutritional composition.
[0048] Secondly, there are differences between different samples of the same type of food, such as differences in moisture content, fat content and protein content, which will affect the spectral signal and the accuracy of the calculated nutritional components.
[0049] Furthermore, there are errors in the infrared spectrometer itself. For example, the calibration of the infrared spectrometer and environmental factors will affect the accuracy of the obtained spectral data.
[0050] Through the above technical solution, image data and spectral data can be used as training samples to train the initial food model. Then, the trained target food model will learn the potential impact of image data and spectral data on the nutritional components of food. When the subsequent target food model obtains similar image data and spectral data, it will predict the accurate nutritional components of the food.
[0051] First, when the accuracy of spectral data is low, the spectral data can be compensated by image data, and the target food model can also accurately predict the nutritional composition of the food based on the image data. For example, if the food is pork, the image data shows that the pork has more lean meat and less fat, then the corresponding nutritional composition of the pork has more protein and less fat. When the spectral data cannot accurately identify the ratio of protein to fat, the image data can be used for compensation, and the target food model can accurately predict the nutritional composition ratio of the food.
[0052] Secondly, since the concept of the initial food model is introduced, although the initial training of the initial food model will consume some training samples and training time, after the initial food model is trained into a target food model, the target food model can quickly divide the nutritional components of the food based on image data and spectral data.
[0053] Figure 2 This is an exemplary embodiment of the above step S10, in which the training sample used for interpreting and training the initial food model also includes text data, then the above step S10 also includes the following steps: In step S11, the image data, the text data and the spectrum data are input into the initial food model to obtain the predicted nutritional components of the food.
[0054] The text data is information describing the food in the image data.
[0055] For example, if the image data includes pork, then the text data includes that the food shown in the image is pork, and the pork has more fat and less lean meat.
[0056] When training the initial food model, the image data, spectral data and text data of the previous round of training samples can be input into the initial food model to obtain the predicted nutritional components of the food, and then the loss value between the predicted nutritional components and the actual nutritional components can be calculated. The loss value is used to reversely update the network parameters of each network layer in the initial food model to achieve the update of the initial food model. If the loss value of the previous round does not meet the convergence condition, the image data, spectral data and text data of the next round of training samples are input into the updated initial food model to obtain the predicted nutritional components of the food again, and then the loss value between the predicted nutritional components and the actual nutritional components is recalculated. This iteration is repeated until the final loss value meets the convergence condition.
[0057] Image data is limited by the shooting environment. Factors such as shooting lighting, angle and background will affect the image quality, resulting in poor display quality of the food indicated by the image data. In this case, the nutritional components of the food are predicted based on image data with relatively poor image quality, and the predicted nutritional components will also have errors.
[0058] Through the above technical solution, in addition to using image data and spectral data as training samples for the initial food model, text data of the text description of the food can also be added. The text data describes the freshness of the food, the distribution ratio of the food, etc. from the side, thereby assisting the initial food model to understand the influence of factors such as the freshness of the food and the distribution ratio of the food on the nutritional components of the food, that is, learning the potential relationship between the freshness of the food, the distribution ratio of the food and the nutritional components of the food, so as to obtain more accurate nutritional components.
[0059] For example, as time goes by, the nutritional content of fresh food will gradually decrease. For example, after fresh fruits and vegetables are picked, the content of vitamins and other antioxidants in the fruits will decrease as the storage time increases. By adding text descriptions of the fruits, for example, the freshness of the fruit is low, then when training the initial food model, the image data, spectral data and text data of the fruit will be used to train the initial food model, and the actual nutritional content of the fruit will be used as the training label of the initial food model. After the initial food model is trained with these training samples, it will accurately predict the nutritional content of similar fruits with low freshness in the future.
[0060] Figure 3An exemplary embodiment involved in the above step S10 is shown, which is used to explain a scheme for obtaining predicted nutritional components of food by fusing image data, spectral data and text data, and includes the following steps: In step S12, the image data, the spectrum data and the text data are fused to obtain a first feature.
[0061] Optionally, the image data, spectral data and text data may be directly concatenated to obtain the first feature.
[0062] Optionally, respective weights may be set for the image data, the spectral data, and the text data, and the sum of the three products of the product of the first weight of the image data and the image data, the product of the second weight of the spectral data and the spectral data, and the product of the third weight of the text data and the text data is calculated to obtain the first feature. In this process, the first weight, the second weight, and the third weight may be determined according to actual needs. For example, if the user believes that the spectral data is more accurate, the second weight may be increased to increase the importance of the spectral data.
[0063] Optionally, the image data, spectral data and text data may be mapped from features of different modalities into the same feature space and then fused to obtain the first feature.
[0064] It can be seen that the first feature is the feature after the image data, spectral data and text data are initially fused.
[0065] In step S13, the first feature is fused with the image data, the spectrum data and the text data respectively to obtain a first fused feature, a second fused feature and a third fused feature.
[0066] Optionally, the first feature can be fused with the image data to obtain a first fused feature, so that the first fused feature not only has the image data, but also considers the influence of the first feature after the fusion of image data, spectral data and text data on the image data, so that the subsequent initial food model can better consider features of different dimensions to capture information of different dimensions and improve the robustness of the initial food model.
[0067] Optionally, the first feature can be fused with the spectral data to obtain a second fused feature, so that the second fused feature not only has the spectral data, but also considers the influence of the first feature after the fusion of image data, spectral data and text data on the spectral data, so that the subsequent initial food model can better consider the features of different dimensions to capture information of different dimensions and improve the robustness of the initial food model.
[0068] Optionally, the first feature can be fused with the text data to obtain a third fused feature, so that the third fused feature not only has the text data, but also considers the influence of the first feature after the fusion of image data, spectral data and text data on the text data, so that the subsequent initial food model can better consider features of different dimensions to capture information of different dimensions and improve the robustness of the initial food model.
[0069] In step S14, the first fusion feature, the second fusion feature and the third fusion feature are input into the initial food model to obtain the predicted nutritional components of the food.
[0070] After the first fusion feature, the second fusion feature and the third fusion feature are input into the initial food model, when the initial food model judges the potential relationship between the first fusion feature and the nutritional components of the food, it will not only judge the potential relationship between the image data in the first fusion feature and the nutritional components of the food, but also judge the potential relationship between the information fused from the image data, spectral data and text data and the nutritional components of the food, thereby making the nutritional components of the food obtained based on the first fusion feature more accurate.
[0071] When judging the potential relationship between the second fusion feature and the nutritional components of the food, the initial food model not only judges the potential relationship between the spectral data in the second fusion feature and the nutritional components of the food, but also judges the potential relationship between the information fused from the image data, spectral data and text data and the nutritional components of the food, thereby making the nutritional components of the food obtained based on the second fusion feature more accurate.
[0072] When judging the potential relationship between the third fusion feature and the nutritional components of the food, the initial food model will not only judge the potential relationship between the text data in the third fusion feature and the nutritional components of the food, but also judge the potential relationship between the information fused from the image data, spectral data and text data and the nutritional components of the food, thereby making the nutritional components of the food obtained based on the third fusion feature more accurate.
[0073] The final initial food model will obtain the nutritional components of the food in three cases based on the first fusion feature, the second fusion feature and the third fusion feature. Finally, the nutritional components of these three cases will be subjected to weighted average and other fusion operations to obtain the final nutritional components of the food.
[0074] Through the above technical scheme, in the process of using image data, spectral data and text data as training samples to train the initial food model, it is not simply to use image data, spectral data and text data separately as training samples to train the initial food model, but the first feature after the fusion of image data, spectral data and text data is fused with these three data for a second time, and the obtained first fusion feature, second fusion feature and third fusion feature are input into the initial food model to train the initial food model. It takes into account the mutual influence relationship among the three data, and also enables the initial food model to learn the influence relationship among image data, spectral data and text data, and consider the influence of these three data on the nutritional components of food from multiple dimensions and angles to obtain more accurate nutritional components.
[0075] Figure 4 An exemplary embodiment involved in the above step S10 is shown, which is used to explain that the predicted nutritional components of food will not only take into account historical training samples but also real-time training samples, including the following steps: In step S15, the real-time image data and spectral data, and the historical image data and spectral data are input into the initial food model to obtain the predicted nutritional components of the food.
[0076] Historical training samples include historical image data, spectral data and text data; real-time training samples include real-time image data, spectral data and text data.
[0077] After each use of the initial food model, the current real-time image data, spectral data and text data will be added to the historical image data, spectral data and text data, and the initial food model will be calibrated and updated so that the initial food model can be updated in real time according to the image data, spectral data and text data of each new food tested.
[0078] In the process of testing the nutritional components of food, the types of food are different, and the nutritional components of the same type of food also vary. If fixed historical image data, spectral data and text data are used to train the initial food model, the generalization of the trained initial food model will be poor. The trained initial food model can only be applied to the prediction of the nutritional components of a few foods.
[0079] Through the above technical scheme, each time a new food to be tested is added, the image data, spectral data and text data of the food will be added to the historical image data, spectral data and text data, that is, the real-time training samples are added to the historical training samples, and the network parameters of the initial food model are calibrated and updated in real time, so that the initial food model can be updated in real time each time a new food to be tested is added, thereby improving the generalization ability of the initial food model, so that the initial food model can accurately predict the nutritional components of the food when faced with different types of food and different morphologies of food under the same type.
[0080] Figure 5 An exemplary embodiment of the present invention is shown, which is used to explain an exemplary scheme for collecting image data, and includes the following steps: In step S50, single-angle images of the food at various angles are collected.
[0081] The single-angle image may be an image of the food at a single angle.
[0082] Optionally, see Figure 6 As shown, a tray and multiple image acquisition devices located outside the tray can be configured in the infrared spectrometer, such as multiple cameras A~D. The multiple image acquisition devices are distributed around the tray. After the food is placed on the tray, the multiple image acquisition devices located around the tray can respectively capture single-angle images of the food at various angles. The single-angle images at various angles are images of the food at various viewing angles.
[0083] Optionally, a tray can be configured in the infrared spectrometer, and the tray is rotatable. An image acquisition device can be configured in the infrared spectrometer. When the image acquisition device acquires images of food, the tray can be controlled to rotate so that the food on the tray rotates. In this way, the image acquisition device can acquire images of the food at various viewing angles.
[0084] In step S60, a surround image of the food is obtained based on the single-angle images of the food at various angles, and the surround image contains image data.
[0085] The single-angle images of the food at various angles can be spliced together to obtain a surround-view image of the food, which contains the above-mentioned image data.
[0086] Through the above technical solution, by collecting single-angle images of food at various angles and splicing the single-angle images at various angles into a panoramic image of the food, the panoramic image can be made able to display the food distribution of the food from various viewing angles. For example, the panoramic image can display the fat and lean distribution, fascia distribution, etc. of pork from various viewing angles. After using a more comprehensive panoramic image as a training sample for the initial food model, the nutritional components output by the initial food model will also be more comprehensive and accurate.
[0087] Figure 7 An exemplary embodiment involved in the above step S10 is shown, and the exemplary embodiment is used to explain an exemplary scheme for quickly obtaining predicted nutritional components, which includes the following steps: In step S16, a target area in the surround view image having a change value greater than a preset value is extracted.
[0088] The change value includes the change value of color and brightness. The target area with the change value greater than the preset value in the surround image includes: if the change value between two adjacent pixels in the surround image is greater than the preset value, the two adjacent pixels are used as pixels in the target area, and multiple groups of two adjacent pixels constitute the target area.
[0089] Optionally, an edge detection algorithm may be used to calculate grayscale changes of pixels in the surround image in the horizontal and vertical directions, thereby finding edge lines and dividing lines in the image and locating target areas in the surround image where the change value is greater than a preset value.
[0090] Optionally, a differential algorithm (such as forward differential, backward differential and intermediate differential, etc.) can be used to detect changes in the surround image by calculating the differences between adjacent pixels, thereby locating a target area in the surround image where the change value is greater than a preset value.
[0091] Optionally, a clustering operation may be performed on the pixels in the surround view image, so as to identify a target area with a more obvious change in the surround view image.
[0092] In step S17, the preset image features in the target area are used as image data in the surround view image.
[0093] The preset image features include at least one of color features and texture features.
[0094] For example, if the food is pork, and there is a large distribution of fat and lean meat on the pork, then the target area in the surround image with a change value greater than a preset value, such as the critical area between fat and lean meat, can be used as the target area. The target area includes both the color characteristics and texture characteristics of the lean meat part and the color characteristics and texture characteristics of the fat meat part.
[0095] Through the above scheme, there is no need to use the entire surround image as the input parameter of the initial food model. Instead, the color features and texture features in the target area with more obvious changes in the surround image can be used as the input parameters of the initial food model. In this way, the initial food model does not need to recognize the entire surround image, but can process the color features and texture features in the target area with more obvious changes in the surround image to predict the nutritional components of the food. This reduces the calculation and processing amount of the initial food model, so that the initial food model can predict the nutritional components of the food more quickly.
[0096] Figure 8 is a block diagram of a food model training device according to an exemplary embodiment, the food model training device 800 includes: a training module 810; The training module 810 is configured to perform multiple rounds of iterative training on the initial food model using multiple groups of training samples to obtain a target food model; each group of training samples in the multiple groups of training samples includes image data and spectral data of the food, and each round of iterative training in the multiple rounds of iterative training includes the following steps: Inputting the image data and the spectral data into an initial food model to obtain predicted nutritional components of the food; Determining a loss value between the predicted nutritional content and the actual nutritional content of the food; updating the network parameters of the initial food model according to the loss value until the loss value satisfies a convergence condition; The initial food model when the loss value satisfies the convergence condition is used as the target food model.
[0097] Optionally, the training module 810 is further configured to update the network parameters of the initial food model according to the loss value when the loss value does not meet the convergence condition until the loss value meets the convergence condition.
[0098] Optionally, the training sample also includes text data, which is a description text of the food; the training module 810 is also configured to input the image data, the text data and the spectral data into the initial food model to obtain the predicted nutritional components of the food.
[0099] Optionally, the training module 810 includes: A first fusion submodule is configured to fuse the image data, the spectral data and the text data to obtain a first feature; A second fusion submodule is configured to fuse the first feature with the image data, the spectral data and the text data respectively to obtain a first fusion feature, a second fusion feature and a third fusion feature; The input submodule is configured to input the first fusion feature, the second fusion feature and the third fusion feature into the initial food model to obtain the predicted nutritional components of the food.
[0100] Optionally, the training samples include historical training samples and real-time training samples, the historical training samples include historical image data and spectral data, and the real-time training samples include real-time image data and spectral data; the training module 810 is also configured to input the real-time image data and spectral data, as well as the historical image data and spectral data into the initial food model to obtain the predicted nutritional components of the food.
[0101] Optionally, the food model training device 800 includes: A collection module, configured to collect single-angle images of the food at various angles; The splicing module is configured to obtain a panoramic image of the food according to the single-angle images of the food at various angles; the panoramic image contains the image data.
[0102] Optionally, the training module 810 includes: The extraction submodule is configured to extract a target area in the surround image whose change value is greater than a preset value; the change value includes a change value in color and brightness.
[0103] The determination submodule is configured to use the preset image features in the target area as image data in the surround view image.
[0104] Optionally, the preset image feature includes at least one of a color feature and a texture feature in the target area.
[0105] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0106] Fig. 9 FIG. 9 is a block diagram of a first electronic device 900 according to an exemplary embodiment. Fig. 9 As shown, the first electronic device 900 may include: a first processor 901 and a first memory 902. The first electronic device 900 may also include one or more of a multimedia component 903, a first input / output (I / O) interface 904, and a first communication component 905.
[0107] The first processor 901 is used to control the overall operation of the first electronic device 900 to complete all or part of the steps in the above-mentioned food model training method. The first memory 902 is used to store various types of data to support the operation of the first electronic device 900, and these data may include, for example, instructions for any application or method used to operate on the first electronic device 900, and application-related data, such as contact data, sent and received messages, pictures, audio, video, etc. The first memory 902 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read-Only Memory, referred to as EPROM), programmable read-only memory (Programmable Read-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic memory, flash memory, magnetic disk or optical disk. The multimedia component 903 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone, which is used to receive external audio signals. The received audio signal may be further stored in the first memory 902 or sent through the first communication component 905. The audio component also includes at least one speaker for outputting audio signals. The first input / output interface 904 provides an interface between the first processor 901 and other interface modules, and the above-mentioned other interface modules may be keyboards, mice, buttons, etc. These buttons may be virtual buttons or physical buttons. The first communication component 905 is used for wired or wireless communication between the first electronic device 900 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G or 4G, or a combination of one or more of them, so the corresponding first communication component 905 may include: Wi-Fi module, Bluetooth module, NFC module.
[0108] In an exemplary embodiment, the first electronic device 900 can be implemented by one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), controllers, microcontrollers, microprocessors or other electronic components to execute the above-mentioned food model training method.
[0109] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, and when the program instructions are executed by a processor, the steps of the above-mentioned food model training method are implemented. For example, the computer-readable storage medium can be the above-mentioned first memory 902 including program instructions, and the above-mentioned program instructions can be executed by the first processor 901 of the first electronic device 900 to complete the above-mentioned food model training method.
[0110] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program that can be executed by a processor. When the computer program is executed by the processor, the steps of the above-mentioned food model training method are implemented.
[0111] Fig.10 1 is a block diagram of a second electronic device 1000 according to an exemplary embodiment. For example, the second electronic device 1000 may be provided as a server. Fig.10 The second electronic device 1000 includes a second processor 1022, which may be one or more, and a second memory 1032 for storing a computer program executable by the second processor 1022. The computer program stored in the second memory 1032 may include one or more modules each corresponding to a set of instructions. In addition, the second processor 1022 may be configured to execute the computer program to perform the above-mentioned food model training method.
[0112] In addition, the second electronic device 1000 may further include a power supply component 1026 and a second communication component 1050, the power supply component 1026 may be configured to perform power management of the second electronic device 1000, and the second communication component 1050 may be configured to implement communication of the second electronic device 1000, for example, wired or wireless communication. In addition, the second electronic device 1000 may further include a second input / output (I / O) interface 1058. The second electronic device 1000 may operate based on an operating system stored in the second memory 1032.
[0113] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, and when the program instructions are executed by a processor, the steps of the above-mentioned food model training method are implemented. For example, the computer-readable storage medium can be the above-mentioned second memory 1032 including program instructions, and the above-mentioned program instructions can be executed by the second processor 1022 of the second electronic device 1000 to complete the above-mentioned food model training method.
[0114] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program that can be executed by a processor. When the computer program is executed by the processor, the steps of the above-mentioned food model training method are implemented.
[0115] The preferred embodiments of the present invention are described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, a variety of simple modifications can be made to the technical solution of the present invention, and these simple modifications all belong to the protection scope of the present invention.
[0116] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, the present invention will not further describe various possible combinations.
[0117] In addition, various embodiments of the present invention may be arbitrarily combined, and as long as they do not violate the concept of the present invention, they should also be regarded as the contents disclosed by the present invention.
Claims
1. A food model training method, characterized in that: The method comprises: using multiple groups of training samples to perform multiple rounds of iterative training on the initial food model to obtain the target food model; each group of training samples in the multiple groups of training samples comprises image data and spectral data of the food, and each round of iterative training in the multiple rounds of iterative training comprises the following steps: Inputting the image data and the spectral data into an initial food model to obtain predicted nutritional components of the food; Determining a loss value between the predicted nutritional content and the actual nutritional content of the food; updating the network parameters of the initial food model according to the loss value until the loss value satisfies a convergence condition; The initial food model when the loss value satisfies the convergence condition is used as the target food model.
2. The food model training method according to claim 1, characterized in that: The updating of the network parameters of the initial food model according to the loss value comprises: When the loss value does not satisfy the convergence condition, the network parameters of the initial food model are updated according to the loss value until the loss value satisfies the convergence condition.
3. The food model training method according to claim 1, characterized in that: The training sample also includes text data, and the text data is a description text of the food; The step of inputting the image data and the spectral data into an initial food model to obtain predicted nutritional components of the food comprises: The image data, the text data and the spectral data are input into the initial food model to obtain the predicted nutritional components of the food.
4. The food model training method according to claim 3, characterized in that: The step of inputting the image data, the text data and the spectral data into the initial food model to obtain the predicted nutritional components of the food includes: fusing the image data, the spectral data and the text data to obtain a first feature; The first feature is respectively fused with the image data, the spectrum data and the text data to obtain a first fused feature, a second fused feature and a third fused feature; The first fusion feature, the second fusion feature and the third fusion feature are input into the initial food model to obtain the predicted nutritional components of the food.
5. The food model training method according to claim 1, characterized in that: The training samples include historical training samples and real-time training samples, the historical training samples include historical image data and spectral data, and the real-time training samples include real-time image data and spectral data; the image data and the spectral data are input into the initial food model to obtain the predicted nutritional components of the food, including: The real-time image data and spectral data, as well as the historical image data and spectral data are input into the initial food model to obtain predicted nutritional components of the food.
6. The food model training method according to claim 1, characterized in that: The method further comprises: Collecting single-angle images of the food at various angles; A surround image of the food is obtained according to the single-angle images of the food at various angles; the surround image contains the image data.
7. The food model training method according to claim 6, characterized in that: The step of inputting the image data and the spectral data into an initial food model to obtain predicted nutritional components of the food comprises: Extracting a target area in the surround image whose change value is greater than a preset value; the change value includes a change value in color and brightness; The preset image features in the target area are used as image data in the surround view image.
8. The food model training method according to claim 7, characterized in that: The preset image feature includes at least one of a color feature and a texture feature in the target area.
9. A food model training device, characterized in that: The device comprises: The training module is configured to use multiple groups of training samples to perform multiple rounds of iterative training on the initial food model to obtain a target food model; each group of training samples in the multiple groups of training samples includes image data and spectral data of the food, and each round of iterative training in the multiple rounds of iterative training includes the following steps: Inputting the image data and the spectral data into an initial food model to obtain predicted nutritional components of the food; Determining a loss value between the predicted nutritional content and the actual nutritional content of the food; updating the network parameters of the initial food model according to the loss value until the loss value satisfies a convergence condition; The initial food model when the loss value satisfies the convergence condition is used as the target food model.
10. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the food model training method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Generating hyperspectral image database by machine learning and mapping of color images to hyperspectral domain
US20190311230A1
Systems and methods for food analysis, feedback, and recommendation using a sensor system
US20230153972A1
System and method for nutrition analysis using food image recognition
US9349297B1
Hyperspectral imaging sensor
WO2018165605A1
Nutritional management method and system using deep learning-based food image recognition model
WO2023159909A1