Food model training method, device and electronic device

By conducting multiple iterative training on the initial food model, combining image data and spectral data, the problem of inaccurate infrared spectrometer in food nutritional composition testing is solved, and accurate prediction of food nutritional composition is achieved.

CN120014631BActive Publication Date: 2025-06-27SICHUAN INST OF PROD QUALITY SUPERVISION INSPECTION & TESTING (SICHUAN QUALITY & TECH REVIEW & EVALUATION CENT) +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510502994.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-06-27
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

In the prior art, in the test of food nutritional ingredients, the food nutritional ingredients obtained by infrared spectrometer scanning are inaccurate, mainly due to the complex food ingredients, different samples and infrared spectrometer errors.

Method used

Multiple training samples are used to conduct multiple iterative training on the initial food model, combining image data and spectral data, and the model parameters are updated through loss values ​​until the convergence conditions are met, and the target food model is obtained.

Benefits of technology

By compensating for the insufficient spectral data by image data, the target food model can accurately predict the nutritional content of the food, improving the accuracy of the test.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014631B_ABST
    Figure CN120014631B_ABST
Patent Text Reader

Abstract

The present invention relates to a food model training method, apparatus and electronic device, and relates to the technical field of food testing. The method includes: performing multiple rounds of iterative training on an initial food model using multiple sets of training samples to obtain a target food model; each round of iterative training in the multiple rounds of iterative training includes the following steps: inputting the image data and the spectral data into the initial food model to obtain the predicted nutritional components of the food; determining the loss value between the predicted nutritional components and the actual nutritional components of the food; updating the network parameters of the initial food model according to the loss value until the loss value meets the convergence condition; using the initial food model when the loss value meets the convergence condition as the target food model. Using the food model training method proposed by the present invention, accurate nutritional components can be quickly predicted based on the target food model after the target food model is trained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of food testing, and specifically, to a food model training method, apparatus, and electronic device. Background Art

[0002] Testing the nutritional components of food is of great significance and value, which can help users understand the nutritional components such as calories, proteins, fats, carbohydrates, vitamins, and minerals contained in food, thereby assisting users in evaluating the potential impact of food on health.

[0003] In the related art, an infrared spectrometer is used to perform spectral scanning on food to obtain spectral data of the food, and the spectral data is used to analyze the nutritional components of the food. However, the obtained nutritional components of the food may be inaccurate. Summary of the Invention

[0004] The purpose of the present invention is to provide a food model training method, apparatus, and electronic device to solve the above technical problems.

[0005] To achieve the above purpose, a first aspect of an embodiment of the present invention provides a food model training method, the method including: using multiple groups of training samples to perform multiple rounds of iterative training on an initial food model respectively to obtain a target food model; each group of training samples in the multiple groups of training samples includes image data and spectral data of food, and each round of iterative training in the multiple rounds of iterative training includes the following steps:

[0006] Inputting the image data and the spectral data into the initial food model to obtain the predicted nutritional components of the food;

[0007] Determining the loss value between the predicted nutritional components and the actual nutritional components of the food;

[0008] Updating the network parameters of the initial food model according to the loss value until the loss value meets the convergence condition;

[0009] Taking the initial food model when the loss value meets the convergence condition as the target food model.

[0010] Optionally, the updating the network parameters of the initial food model according to the loss value includes:

[0011] In the case where the loss value does not meet the convergence condition, updating the network parameters of the initial food model according to the loss value until the loss value meets the convergence condition.

[0012] Optionally, the training sample further includes text data, which is a descriptive text of the food; the step of inputting the image data and the spectral data into the initial food model to obtain the predicted nutritional components of the food includes:

[0013] Input the image data, the text data, and the spectral data into the initial food model to obtain the predicted nutritional components of the food.

[0014] Optionally, the step of inputting the image data, the text data, and the spectral data into the initial food model to obtain the predicted nutritional components of the food includes:

[0015] Fuse the image data, the spectral data, and the text data to obtain a first feature;

[0016] Fuse the first feature with the image data, the spectral data, and the text data respectively to obtain a first fused feature, a second fused feature, and a third fused feature;

[0017] Input the first fused feature, the second fused feature, and the third fused feature into the initial food model to obtain the predicted nutritional components of the food.

[0018] Optionally, the training sample includes historical training samples and real-time training samples. The historical training samples include historical image data and spectral data, and the real-time training samples include real-time image data and spectral data; the step of inputting the image data and the spectral data into the initial food model to obtain the predicted nutritional components of the food includes:

[0019] Input the real-time image data and spectral data, and the historical image data and spectral data into the initial food model to obtain the predicted nutritional components of the food.

[0020] Optionally, the method further includes:

[0021] Collect single-angle images of the food from various angles;

[0022] Obtain a panoramic image of the food based on the single-angle images of the food from various angles; the panoramic image contains the image data.

[0023] Optionally, the step of inputting the image data and the spectral data into the initial food model to obtain the predicted nutritional components of the food includes:

[0024] Extract the target regions in the panoramic image where the change values are greater than a preset value; the change values include change values in color and brightness.

[0025] Use the preset image features in the target area as the image data in the surround-view image.

[0026] Optionally, the preset image features include at least one of the color feature and the texture feature in the target area.

[0027] Through the above technical solution, when the accuracy of the spectral data is low, the spectral data can be compensated by the image data, and the target food model can also predict the accurate nutritional components of the food according to the image data. For example, taking the food as pork, if the image data shows that the pork has more lean meat and less fat, then the corresponding nutritional components of the pork have more protein and less fat. When the spectral data cannot accurately identify the proportion of protein and fat, the image data can be used for compensation, and the target food model can accurately predict the proportion of the nutritional components of the food.

[0028] Other features and advantages of the present invention will be described in detail in the subsequent specific implementation part. Brief Description of the Drawings

[0029] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the following specific implementation, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:

[0030] Figure 1 is a flowchart of the steps of a food model training method shown according to an exemplary embodiment.

[0031] Figure 2 is a flowchart of the steps of a food model training method shown according to an exemplary embodiment.

[0032] Figure 3 is a flowchart of the steps of a food model training method shown according to an exemplary embodiment.

[0033] Figure 4 is a flowchart of the steps of a food model training method shown according to an exemplary embodiment.

[0034] Figure 5 is a flowchart of the steps of a food model training method shown according to an exemplary embodiment.

[0035] Figure 6 is a schematic structural diagram of an infrared spectrometer shown according to an exemplary embodiment.

[0036] Figure 7 is a flowchart of the steps of a food model training method shown according to an exemplary embodiment.

[0037] Figure 8It is a block diagram of a food model training device shown according to an exemplary embodiment.

[0038] Figure 9 It is a block diagram of a first electronic device shown according to an exemplary embodiment.

[0039] Figure 10 It is a block diagram of a second electronic device shown according to an exemplary embodiment. Detailed implementation manners

[0040] The following details the specific implementation manners of the present invention with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only for explaining and illustrating the present invention, and are not used to limit the present invention.

[0041] Figure 1 It is a step flowchart of a food model training method proposed according to an exemplary embodiment. This food model training method can be applied to electronic devices such as infrared spectrometers.

[0042] This food model training method includes: using multiple groups of training samples to perform multiple rounds of iterative training on an initial food model respectively to obtain a target food model; each group of training samples in the multiple groups of training samples includes image data and spectral data of food, and each round of iterative training in the multiple rounds of iterative training includes the following steps:

[0043] In step S10, input the image data and the spectral data into the initial food model to obtain the predicted nutritional components of the food.

[0044] The image data is image data related to the food to be measured. The image of the food to be measured can be collected by an image acquisition device, and the image data can intuitively reflect the proportion of some nutritional components in the food.

[0045] For example, the image data collected by the image acquisition device is image data related to pork. If the lean part of the pork accounts for a large proportion, then the protein content of the pork is higher; if the fat part of the pork accounts for a large proportion, then the fat content of the pork is higher. Therefore, the image data can indirectly reflect the proportion of nutritional components, and the nutritional components of the food can be obtained to a certain extent based on the image data.

[0046] The spectral data is the spectral data of the food collected by an infrared spectrometer. This spectral data can reflect the physical structure and chemical composition of the food. Therefore, the nutritional components of the food can be analyzed based on the physical structure and chemical composition of the food.

[0047] The process of obtaining the spectral data of food includes: collecting food, performing at least one treatment such as grinding, homogenizing, and drying on the food. Grinding is to grind solid food into powder to improve the uniformity of the food. Homogenizing is to homogenize liquid food to eliminate the influence of particulate matter. Drying is to dry the food to remove the influence of moisture. Then, the processed food is filled into containers such as cuvettes and quartz plates. Then, the food filled in the container is placed in an infrared spectrometer for testing. Before testing, first set the parameters of the infrared spectrometer, such as the scanning range, scanning speed, spectral resolution, etc., and then run the infrared spectrometer to perform spectral scanning on the food to obtain the spectral data of the food.

[0048] The initial food model is a model to be trained. The initial food model can be trained with multiple groups of training samples and the labels corresponding to the multiple groups of training samples. Each group of training samples in the multiple groups of training samples includes image data and spectral data, and the labels corresponding to the multiple groups of training samples all include the actual nutritional components of the food.

[0049] The image data and spectral data can be input into the initial food model to obtain the predicted nutritional components of the food. The predicted nutritional components are the nutritional components of the food predicted by the initial food model with low accuracy.

[0050] Among them, the nutritional components of food include carbohydrates, proteins, fats, vitamins, minerals, etc.

[0051] In step S20, determine the loss value between the predicted nutritional components and the actual nutritional components of the food.

[0052] Optionally, the difference between the predicted nutritional components and the actual nutritional components can be used as the loss value between the predicted nutritional components and the actual nutritional components of the food. The larger the loss value, the greater the gap between the predicted nutritional components predicted by the initial food model and the true actual nutritional components, and the lower the accuracy of the predicted nutritional components predicted by the initial food model. On the contrary, the higher the accuracy.

[0053] Among them, the predicted nutritional components and the actual nutritional components can be substituted into loss functions such as the cross-entropy loss function and the root mean square error to obtain the loss value, which will not be elaborated here.

[0054] Among them, the actual nutritional components are the accurate nutritional components of the food obtained by measuring the food multiple times.

[0055] In step S30, update the network parameters of the initial food model according to the loss value until the loss value meets the convergence condition.

[0056] The training of the initial food model includes a forward propagation stage and a backpropagation stage. In the forward propagation stage, the initial food model includes multiple network layers such as an input layer, a hidden layer, and an output layer. Image data and spectral data can be input into the input layer. After calculation by the hidden layer, the predicted nutritional components are output through the output layer, and then the loss value between the predicted nutritional components and the actual nutritional components is calculated. In the backpropagation stage, starting from the output layer, the gradient of the loss value with respect to the network parameters of each network layer is calculated, that is, the partial derivative of the loss value with respect to each network layer is calculated, and then the calculated gradient of the network parameters is superimposed on the basis of the network parameters to obtain the network parameters of the updated initial food model to reduce the loss value.

[0057] For example, the image data and spectral data of the previous round of training samples can be input into the initial food model to obtain the predicted nutritional components of the food, and then the loss value between the predicted nutritional components and the actual nutritional components is calculated. The loss value is used to inversely update the network parameters of each network layer in the initial food model to achieve the update of the initial food model. If the loss value of the previous round does not meet the convergence condition, the image data and spectral data of the next round of training samples are input into the updated initial food model, and then the predicted nutritional components of the food are obtained again, and the loss value between the predicted nutritional components and the actual nutritional components is calculated again. This is repeated iteratively until the finally obtained loss value meets the convergence condition.

[0058] Among them, the loss value meeting the convergence condition includes that the loss value is less than the preset value, the loss value no longer becomes smaller, or the training of the initial food model reaches the preset number of iterations.

[0059] In step S40, the initial food model when the loss value meets the convergence condition is used as the target food model.

[0060] When the loss value meets the convergence condition, it means that the difference between the predicted nutritional components output by the initial food model and the actual nutritional components is small, the predicted nutritional components output by the initial food model are close to the actual nutritional components, or equal to the actual nutritional components, and the accuracy of the predicted nutritional components output by the initial food model is relatively high. Therefore, the initial food model when the loss value meets the convergence condition can be used as the trained target food model.

[0061] In subsequent applications of the target food model, the collected image data related to the food and the spectral data of the food can be input into the target food model to obtain the accurate nutritional components of the food.

[0062] In related technologies, an infrared spectrometer is used to scan the nutritional components of food, but there are certain inaccuracies in the nutritional components of food scanned by the infrared spectrometer for the following reasons:

[0063] First, the nutritional components of food are complex, and the infrared absorption spectra of different components overlap with each other, making it difficult to analyze the spectral data and affecting the accuracy of the calculated nutritional components.

[0064] Secondly, there are differences between different samples of the same type of food. For example, differences in moisture content, fat content, and protein content will all affect the spectral signal and the accuracy of the calculated nutritional components.

[0065] Furthermore, there are also errors in the infrared spectrometer itself. For example, the calibration of the infrared spectrometer and environmental factors will affect the accuracy of the obtained spectral data.

[0066] Through the above technical solutions, image data and spectral data can be used as training samples to train the initial food model at the same time. Then, the trained target food model will learn the potential influence of image data and spectral data on the nutritional components of food. When the subsequent target food model obtains similar image data and spectral data, it will predict the accurate nutritional components of the food.

[0067] On the one hand, when the accuracy of the spectral data is low, the image data can also compensate for the spectral data, and the target food model can also predict the accurate nutritional components of the food based on the image data. For example, taking the food as pork, if the image data shows that the pork has more lean meat and less fat, then the corresponding nutritional components of the pork have more protein and less fat. When the spectral data cannot accurately identify the proportion of protein and fat, the image data can be used for compensation, and the target food model can accurately predict the proportion of the nutritional components of the food.

[0068] On the other hand, due to the introduction of the concept of the initial food model, although some training samples and training time will be consumed in training the initial food model in the early stage, after the initial food model is trained into the target food model, the target food model can quickly divide the nutritional components of the food based on the image data and spectral data.

[0069] Figure 2 This is an exemplary embodiment involved in the above step S10, which is used to interpret that the training samples for training the initial food model also include text data. Then, the above step S10 further includes the following steps:

[0070] In step S11, the image data, the text data, and the spectral data are input into the initial food model to obtain the predicted nutritional components of the food.

[0071] The text data is information that describes the food in the image data.

[0072] For example, if the image data includes pork, then the text data includes that the food shown in the image is pork, and this pork has more fat and less lean meat.

[0073] When training the initial food model, the image data, spectral data, and text data of the training samples in the previous round can be input into the initial food model to obtain the predicted nutritional components of the food. Then, the loss value between the predicted nutritional components and the actual nutritional components is calculated, and the loss value is used to inversely update the network parameters of each network layer in the initial food model to update the initial food model. If the loss value in the previous round does not meet the convergence condition, the image data, spectral data, and text data of the training samples in the next round are input into the updated initial food model, and then the predicted nutritional components of the food are obtained again, and the loss value between the predicted nutritional components and the actual nutritional components is calculated again. This process is repeated iteratively until the finally obtained loss value meets the convergence condition.

[0074] The image data is restricted by the shooting environment. Factors such as shooting light, angle, and background will affect the image quality, resulting in poor display quality of the food indicated by the image data. In this case, predicting the nutritional components of the food based on the image data with relatively poor image quality will also result in errors in the predicted nutritional components.

[0075] Through the above technical solution, in addition to using image data and spectral data as the training samples of the initial food model, text data describing the food can also be added. This text data describes the freshness of the food, the proportion of food distribution, etc. from the side, so as to assist the initial food model to understand the influence of factors such as the freshness of the food, the proportion of food distribution on the nutritional components of the food, that is, to learn the potential relationship between the freshness of the food, the proportion of food distribution and the nutritional components of the food, in order to obtain more accurate nutritional components.

[0076] For example, over time, the nutritional components in fresh food will gradually decrease. For example, after fresh fruits and vegetables are picked, the content of vitamins and other antioxidant substances in the fruits will decrease as the storage time prolongs. By adding text descriptions of the fruits, such as the freshness of the fruits is relatively low, then when training the initial food model, the image data, spectral data, and text data of the fruits will be used to train the initial food model, and the actual nutritional components of the fruits will be used as the training labels of the initial food model. After the initial food model is trained by these training samples, when facing similar fruits with relatively low freshness later, it will accurately predict the nutritional components of these fruits.

[0077] Figure 3An exemplary embodiment related to the above step S10 is shown, which is used to interpret the solution for predicting the nutritional components of food after fusing image data, spectral data, and text data, including the following steps:

[0078] In step S12, the image data, the spectral data, and the text data are fused to obtain a first feature.

[0079] Optionally, the image data, the spectral data, and the text data can be directly concatenated to obtain a first feature.

[0080] Optionally, weights can be set for the image data, the spectral data, and the text data, and the sum of the products of the image data and the first weight, the spectral data and the second weight, and the text data and the third weight can be calculated to obtain a first feature. In this process, the magnitudes of the first weight, the second weight, and the third weight can be determined according to actual requirements. For example, if the user believes that the spectral data is more accurate, the second weight can be increased to enhance the importance of the spectral data.

[0081] Optionally, the image data, the spectral data, and the text data can also be mapped from features of different modalities to the same feature space and then fused to obtain a first feature.

[0082] It can be seen that the first feature is the feature after the initial fusion of the image data, the spectral data, and the text data.

[0083] In step S13, the first feature is fused with the image data, the spectral data, and the text data respectively to obtain a first fusion feature, a second fusion feature, and a third fusion feature.

[0084] Optionally, the first feature can be fused with the image data to obtain a first fusion feature, such that the first fusion feature not only has the image data but also takes into account the influence of the first feature after the fusion of the image data, the spectral data, and the text data on the image data. Furthermore, the subsequent initial food model can better consider features of different dimensions to capture information of different dimensions and improve the robustness of the initial food model.

[0085] Optionally, the first feature can be fused with the spectral data to obtain a second fusion feature, such that the second fusion feature not only has the spectral data but also takes into account the influence of the first feature after the fusion of the image data, the spectral data, and the text data on the spectral data. Furthermore, the subsequent initial food model can better consider features of different dimensions to capture information of different dimensions and improve the robustness of the initial food model.

[0086] Optionally, the first feature can be fused with the text data to obtain a third fused feature, such that the third fused feature not only has the text data, but also takes into account the influence of the first feature after the fusion of the image data, spectral data and text data on the text data. Furthermore, the subsequent initial food model can better consider features in different dimensions to capture information in different dimensions and improve the robustness of the initial food model.

[0087] In step S14, the first fused feature, the second fused feature and the third fused feature are input into the initial food model to obtain the predicted nutritional components of the food.

[0088] After the first fused feature, the second fused feature and the third fused feature are input into the initial food model, when the initial food model judges the potential relationship between the first fused feature and the nutritional components of the food, it will not only judge the potential relationship between the image data in the first fused feature and the nutritional components of the food, but also judge the potential relationship between the information after the fusion of the image data, spectral data and text data and the nutritional components of the food, so that the nutritional components of the food obtained based on the first fused feature are more accurate.

[0089] When the initial food model judges the potential relationship between the second fused feature and the nutritional components of the food, it will not only judge the potential relationship between the spectral data in the second fused feature and the nutritional components of the food, but also judge the potential relationship between the information after the fusion of the image data, spectral data and text data and the nutritional components of the food, so that the nutritional components of the food obtained based on the second fused feature are more accurate.

[0090] When the initial food model judges the potential relationship between the third fused feature and the nutritional components of the food, it will not only judge the potential relationship between the text data in the third fused feature and the nutritional components of the food, but also judge the potential relationship between the information after the fusion of the image data, spectral data and text data and the nutritional components of the food, so that the nutritional components of the food obtained based on the third fused feature are more accurate.

[0091] Finally, the initial food model will obtain the nutritional components of the food in three cases based on the first fused feature, the second fused feature and the third fused feature. Finally, fusion operations such as weighted averaging are performed on the nutritional components in these three cases to obtain the final nutritional components of the food.

[0092] In the process of training the initial food model with image data, spectral data, and text data as training samples through the above technical solution, it does not simply use the image data, spectral data, and text data alone as training samples to train the initial food model. Instead, the first feature after fusing the image data, spectral data, and text data is fused with these three types of data again. Then, the obtained first fusion feature, second fusion feature, and third fusion feature are input into the initial food model to train the initial food model. It takes into account the mutual influence relationship among these three types of data, and also enables the initial food model to learn the influence relationship among the image data, spectral data, and text data, considering the influence of these three types of data on the nutritional components of food from multiple dimensions and perspectives to obtain more accurate nutritional components.

[0093] Figure 4 An exemplary embodiment related to the above step S10 is shown, which is used to illustrate that obtaining the predicted nutritional components of food will consider not only historical training samples but also real-time training samples, including the following steps:

[0094] In step S15, the real-time image data and spectral data, as well as the historical image data and spectral data, are input into the initial food model to obtain the predicted nutritional components of the food.

[0095] The historical training samples include historical image data, spectral data, and text data; the real-time training samples include real-time image data, spectral data, and text data.

[0096] After each use of the initial food model, the current real-time image data, spectral data, and text data will be added to the historical image data, spectral data, and text data to calibrate and update the initial food model, so that the initial food model can be updated in real time according to the image data, spectral data, and text data of each newly tested food.

[0097] In the process of testing the nutritional components of food, the types of food are different, and the nutritional components of the same type of food also vary. If fixed historical image data, spectral data, and text data are used to train the initial food model, the generalization of the trained initial food model will be poor, and the trained initial food model can only be applied to the prediction of the nutritional components of a few foods.

[0098] Through the above technical solution, after each new food to be tested is added, the image data, spectral data, and text data of the food are added to the historical image data, spectral data, and text data, that is, the real-time training samples are added to the historical training samples, and the network parameters of the initial food model are calibrated and updated in real time, so that the initial food model can be updated in real time after each new food to be tested is added, improving the generalization ability of the initial food model, and enabling the initial food model to accurately predict the nutritional components of the food in the case of different types of foods and different morphologies of the same type of food.

[0099] Figure 5 An exemplary embodiment related to the present invention is shown. This exemplary embodiment is used to interpret an exemplary scheme for collecting image data, and it includes the following steps:

[0100] In step S50, single-angle images of the food at various angles are collected.

[0101] The single-angle image can be an image of the food at a single angle.

[0102] Optionally, please refer to Figure 6 As shown, a tray and a plurality of image acquisition devices located outside the tray can be configured in the infrared spectrometer. For example, they are a plurality of cameras A to D. The plurality of image acquisition devices are distributed around the tray. After the food is placed on the tray, the plurality of image acquisition devices located around the tray can respectively collect the single-angle images of the food at various angles. The single-angle images at various angles are images of various perspectives of the food.

[0103] Optionally, a tray can also be configured in the infrared spectrometer. The tray is rotatable. Then, an image acquisition device is configured inside the infrared spectrometer. When the image acquisition device collects the image of the food, the rotation of the tray can be controlled to make the food on the tray rotate. In this way, the image acquisition device can collect the images of the food at various perspectives.

[0104] In step S60, according to the single-angle images of the food at various angles, a panoramic image of the food is obtained, and the panoramic image has image data.

[0105] The single-angle images of the food at various angles can be stitched together to obtain the panoramic image of the food, and the panoramic image has the above-mentioned image data.

[0106] Through the above technical solution, by collecting single-angle images of food from various angles and stitching the single-angle images from various angles into a panoramic image of the food, it is possible to make the panoramic image display the food distribution of the food from various perspectives. For example, the panoramic image can display the fat and lean distribution, fascia distribution, etc. of pork from various perspectives. After using the more comprehensive panoramic image as the training sample of the initial food model, the nutritional components output by the initial food model will be more comprehensive and accurate.

[0107] Figure 7 An exemplary embodiment related to the above step S10 is shown. This exemplary embodiment is used to interpret an exemplary solution for quickly obtaining predicted nutritional components, and it includes the following steps:

[0108] In step S16, target regions in the panoramic image with a change value greater than a preset value are extracted.

[0109] The change value includes change values in color and brightness. Target regions in the panoramic image with a change value greater than a preset value include that if the change value between two adjacent pixel points in the panoramic image is greater than the preset value, then these two adjacent pixel points are used as pixel points in the target region, and multiple groups of two adjacent pixel points form the target region.

[0110] Optionally, an edge detection algorithm can be used to calculate the gray-scale changes of pixel points in the panoramic image in the horizontal and vertical directions, so as to find the edge lines and dividing lines in the image and locate the target regions in the panoramic image with a change value greater than the preset value.

[0111] Optionally, a difference algorithm (such as forward difference, backward difference, and central difference, etc.) can be used to detect changes in the panoramic image by calculating the differences between adjacent pixel points, so as to locate the target regions in the panoramic image with a change value greater than the preset value.

[0112] Optionally, clustering operations can also be performed on the pixel points in the panoramic image to identify target regions with more obvious changes in the panoramic image.

[0113] In step S17, the preset image features in the target region are used as the image data in the panoramic image.

[0114] Among them, the preset image features include at least one of color features and texture features.

[0115] Exemplarily, if the food is pork and there is a lot of fat and lean distribution on the pork, then the target regions in the panoramic image with a change value greater than the preset value, such as the critical region between fat and lean, can be used as the target region. This target region contains both the color features and texture features of the lean part and the color features and texture features of the fat part.

[0116] Through the above solution, instead of using the entire panoramic image as the input parameter of the initial food model, the color features and texture features in the target area with obvious changes in the panoramic image are used as the input parameters of the initial food model. In this way, the initial food model can process the color features and texture features in the target area with obvious changes in the panoramic image instead of recognizing the entire panoramic image, so as to predict the nutritional components of the food, reducing the calculation amount and processing amount of the initial food model and enabling the initial food model to more quickly predict the nutritional components of the food.

[0117] Figure 8 FIG. 5 is a block diagram of a food model training device according to an exemplary embodiment. The food model training device 800 includes: a training module 810;

[0118] The training module 810 is configured to perform multiple rounds of iterative training on the initial food model using multiple groups of training samples to obtain a target food model; each group of training samples in the multiple groups of training samples includes image data and spectral data of the food, and each round of iterative training in the multiple rounds of iterative training includes the following steps:

[0119] Input the image data and the spectral data into the initial food model to obtain the predicted nutritional components of the food;

[0120] Determine the loss value between the predicted nutritional components and the actual nutritional components of the food;

[0121] Update the network parameters of the initial food model according to the loss value until the loss value meets the convergence condition;

[0122] Use the initial food model when the loss value meets the convergence condition as the target food model.

[0123] Optionally, the training module 810 is further configured to update the network parameters of the initial food model according to the loss value until the loss value meets the convergence condition when the loss value does not meet the convergence condition.

[0124] Optionally, the training sample further includes text data, and the text data is a descriptive text of the food; the training module 810 is further configured to input the image data, the text data and the spectral data into the initial food model to obtain the predicted nutritional components of the food.

[0125] Optionally, the training module 810 includes:

[0126] A first fusion sub-module configured to fuse the image data, the spectral data and the text data to obtain a first feature;

[0127] A second fusion sub-module, configured to fuse the first feature with the image data, the spectral data, and the text data respectively to obtain a first fusion feature, a second fusion feature, and a third fusion feature;

[0128] An input sub-module, configured to input the first fusion feature, the second fusion feature, and the third fusion feature into the initial food model to obtain the predicted nutritional components of the food.

[0129] Optionally, the training samples include historical training samples and real-time training samples. The historical training samples include historical image data and spectral data, and the real-time training samples include real-time image data and spectral data; the training module 810 is further configured to input the real-time image data and spectral data, and the historical image data and spectral data into the initial food model to obtain the predicted nutritional components of the food.

[0130] Optionally, the food model training device 800 includes:

[0131] An acquisition module, configured to acquire single-angle images of the food at various angles;

[0132] A stitching module, configured to obtain a panoramic image of the food according to the single-angle images of the food at various angles; the panoramic image has the image data.

[0133] Optionally, the training module 810 includes:

[0134] An extraction sub-module, configured to extract a target region in the panoramic image where the change value is greater than a preset value; the change value includes change values in color and brightness.

[0135] A determination sub-module, configured to use the preset image features in the target region as the image data in the panoramic image.

[0136] Optionally, the preset image features include at least one of color features and texture features in the target region.

[0137] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0138] Figure 9 is a block diagram of a first electronic device 900 shown according to an exemplary embodiment. As Figure 9As shown, the first electronic device 900 may include: a first processor 901 and a first memory 902. The first electronic device 900 may also include one or more of a multimedia component 903, a first input / output (I / O) interface 904, and a first communication component 905.

[0139] Among them, the first processor 901 is used to control the overall operation of the first electronic device 900 to complete all or part of the steps in the above food model training method. The first memory 902 is used to store various types of data to support the operation of the first electronic device 900. These data may include, for example, instructions for any application or method operating on the first electronic device 900, as well as application-related data, such as contact data, messages sent and received, pictures, audio, video, and so on. The first memory 902 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The multimedia component 903 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the first memory 902 or sent through the first communication component 905. The audio component also includes at least one speaker for outputting audio signals. The first input / output interface 904 provides an interface between the first processor 901 and other interface modules, and the above other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The first communication component 905 is used for wired or wireless communication between the first electronic device 900 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination of one or more of them. Accordingly, the first communication component 905 may include: a Wi-Fi module, a Bluetooth module, and an NFC module.

[0140] In an exemplary embodiment, the first electronic device 900 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, and is used to execute the above-mentioned food model training method.

[0141] In another exemplary embodiment, a computer-readable storage medium including program instructions is further provided. When the program instructions are executed by a processor, the steps of the above-mentioned food model training method are implemented. For example, the computer-readable storage medium may be the first memory 902 including the program instructions, and the above-mentioned program instructions may be executed by the first processor 901 of the first electronic device 900 to complete the above-mentioned food model training method.

[0142] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program that can be executed by a processor. When the computer program is executed by the processor, the steps of the above-mentioned food model training method are implemented.

[0143] Figure 10 is a block diagram of a second electronic device 1000 shown according to an exemplary embodiment. For example, the second electronic device 1000 may be provided as a server. Referring to Figure 10 , the second electronic device 1000 includes a second processor 1022, the number of which may be one or more, and a second memory 1032 for storing computer programs that can be executed by the second processor 1022. The computer programs stored in the second memory 1032 may include one or more modules each corresponding to a set of instructions. In addition, the second processor 1022 may be configured to execute the computer program to execute the above-mentioned food model training method.

[0144] In addition, the second electronic device 1000 may further include a power supply component 1026 and a second communication component 1050. The power supply component 1026 may be configured to perform power management of the second electronic device 1000, and the second communication component 1050 may be configured to enable communication of the second electronic device 1000, for example, wired or wireless communication. In addition, the second electronic device 1000 may further include a second input / output (I / O) interface 1058. The second electronic device 1000 may operate based on an operating system stored in the second memory 1032.

[0145] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided. When the program instructions are executed by a processor, the steps of the above food model training method are implemented. For example, the computer-readable storage medium may be the above-mentioned second memory 1032 including program instructions, and the above program instructions may be executed by the second processor 1022 of the second electronic device 1000 to complete the above food model training method.

[0146] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program capable of being executed by a processor. When the computer program is executed by the processor, the steps of the above food model training method are implemented.

[0147] The preferred embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solutions of the present invention, and these simple modifications all fall within the protection scope of the present invention.

[0148] In addition, it should be noted that, among the various specific technical features described in the above specific embodiments, they can be combined in any suitable manner without conflict. To avoid unnecessary repetition, the present invention will not separately describe various possible combination methods.

[0149] In addition, any combination can be made between various different embodiments of the present invention as long as it does not violate the idea of the present invention, and it should also be regarded as the content disclosed by the present invention.

Claims

1. A food model training method, characterized in that: The method comprises: using multiple groups of training samples to perform multiple rounds of iterative training on the initial food model to obtain the target food model; each group of training samples in the multiple groups of training samples comprises image data and spectral data of the food, and each round of iterative training in the multiple rounds of iterative training comprises the following steps: Inputting the image data and the spectral data into an initial food model to obtain predicted nutritional components of the food; Determining a loss value between the predicted nutritional content and the actual nutritional content of the food; updating the network parameters of the initial food model according to the loss value until the loss value satisfies a convergence condition; Taking the initial food model when the loss value satisfies the convergence condition as the target food model; The method further comprises: Collecting single-angle images of the food at various angles; Obtaining a panoramic image of the food according to the single-angle images of the food at various angles; the panoramic image contains the image data; The step of inputting the image data and the spectral data into an initial food model to obtain the predicted nutritional components of the food includes: Extracting a target area in the surround image whose change value is greater than a preset value; the change value includes a change value in color and brightness; The preset image features in the target area are used as image data in the surround view image.

2. The food model training method according to claim 1, characterized in that: The updating of the network parameters of the initial food model according to the loss value comprises: When the loss value does not satisfy the convergence condition, the network parameters of the initial food model are updated according to the loss value until the loss value satisfies the convergence condition.

3. The food model training method according to claim 1, characterized in that: The training sample also includes text data, and the text data is a description text of the food; The step of inputting the image data and the spectral data into an initial food model to obtain predicted nutritional components of the food comprises: The image data, the text data and the spectral data are input into the initial food model to obtain the predicted nutritional components of the food.

4. The food model training method according to claim 3, characterized in that: The step of inputting the image data, the text data and the spectral data into the initial food model to obtain the predicted nutritional components of the food includes: fusing the image data, the spectral data and the text data to obtain a first feature; The first feature is respectively fused with the image data, the spectrum data and the text data to obtain a first fused feature, a second fused feature and a third fused feature; The first fusion feature, the second fusion feature and the third fusion feature are input into the initial food model to obtain the predicted nutritional components of the food.

5. The food model training method according to claim 1, characterized in that: The training samples include historical training samples and real-time training samples, the historical training samples include historical image data and spectral data, and the real-time training samples include real-time image data and spectral data; the image data and the spectral data are input into the initial food model to obtain the predicted nutritional components of the food, including: The real-time image data and spectral data, as well as the historical image data and spectral data are input into the initial food model to obtain predicted nutritional components of the food.

6. The food model training method according to claim 1, characterized in that: The preset image feature includes at least one of a color feature and a texture feature in the target area.

7. A food model training device, characterized in that: The device comprises: The training module is configured to use multiple groups of training samples to perform multiple rounds of iterative training on the initial food model to obtain a target food model; each group of training samples in the multiple groups of training samples includes image data and spectral data of the food, and each round of iterative training in the multiple rounds of iterative training includes the following steps: Inputting the image data and the spectral data into an initial food model to obtain predicted nutritional components of the food; Determining a loss value between the predicted nutritional content and the actual nutritional content of the food; updating the network parameters of the initial food model according to the loss value until the loss value satisfies a convergence condition; Taking the initial food model when the loss value satisfies the convergence condition as the target food model; The device also includes: A collection module, configured to collect single-angle images of the food at various angles; a splicing module configured to obtain a panoramic image of the food according to the single-angle images of the food at various angles; the panoramic image contains the image data; Wherein, the training module includes: An extraction submodule is configured to extract a target area in the surround image whose change value is greater than a preset value; the change value includes a change value in color and brightness; The determination submodule is configured to use the preset image features in the target area as image data in the surround view image.

8. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the food model training method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Generating hyperspectral image database by machine learning and mapping of color images to hyperspectral domain

    US20190311230A1