Food nutrition evaluation method and system based on multi-modal fusion

Through the combination of a multi-scale dual-branch Transformer encoder and a depth perception enhanced attention fusion module, the accuracy and convenience of food nutrition assessment in the prior art are solved, and efficient and accurate prediction of food nutritional components is achieved.

CN120356204AActive Publication Date: 2025-07-22JIANGNAN UNIV

Patent Information

Application Number
CN202510356219.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-22
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The existing vision-based food nutrition assessment methods have challenges in accuracy and limitations in food occlusion applications. Monocular RGB images estimate food nutritional components are unfavorable, and deep image fusion methods lack information complementary effects on cross-modal features, resulting in performance bottlenecks.

Method used

The food RGB images and RGB-D depth image features were extracted in parallel by using a multi-scale dual-branch Transformer encoder, and cross-modal fusion was performed using the depth perception enhancement attention fusion module. Fusion features were generated through global and local attention mechanisms, and three-dimensional geometric parameters were mapped from the RGB-D features to construct nutrition prediction branches to output five macronutrient contents.

Benefits of technology

It improves the accuracy and convenience of food nutrition estimation, reduces errors, realizes effective processing of food shape, size and surface texture differences, enhances the perception of food volume information, and provides more efficient nutrition prediction capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356204A_ABST
    Figure CN120356204A_ABST
Patent Text Reader

Abstract

The invention discloses a food nutrition evaluation method and system based on multi-modal fusion. The method comprises the steps of 1, extracting multi-scale features of a food RGB image and an RGB-D depth image in parallel through a multi-scale double-branch Transform encoder; 2, performing cross-modal fusion on the multi-scale features by using a depth perception enhanced attention fusion module; step 3, constructing nutrition prediction branches based on the fused features, and outputting five macro nutrient contents; and step 4, training the training data set on the model in an end-to-end mode, and inputting the RGB image and the RGB-D image of the food to be detected into the trained model to predict the nutrient content. According to the method, perception of volume information of different food types is enhanced by using the RGB-D depth mode corresponding to the food RGB image, visual information in the food RGB image and spatial physical characteristics in the RGB-D image are fused through the depth perception attention enhancement module, so that fusion characteristics with identification information are generated, and the evaluation accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of food image recognition, and particularly to a food nutrition assessment method and system based on multimodal fusion. Background Art

[0002] Proper diet management requires making informed decisions about the choice and quantity of food to ensure that our bodies receive the necessary nutrients without overindulging. Research shows that an unbalanced diet or excessive food intake can lead to various adverse consequences, including obesity, diabetes, chronic diseases, and even mental health problems. A comprehensive diet structure rich in basic nutrients such as carbohydrates and proteins is the key to supporting the optimal function of the body and promoting health. Therefore, accurately understanding and assessing food nutrition is of great significance in guiding people to make scientific food choices.

[0003] In recent years, with the rapid development of artificial intelligence and computer vision technologies, automatic and reliable food nutrition assessment has become a reality. This vision-based nutrition assessment method allows users to capture food images through mobile devices to obtain nutrition information, greatly reducing the burden on users and improving the accuracy of estimation. However, due to the high complexity of diets, accurately recording the nutritional content of food remains a challenge. Specifically, the nutritional composition of food is a complex variable affected by various factors such as food type, ingredients, and food volume. Therefore, effectively extracting food features and volume information from food images is very important for improving the accuracy and reliability of nutritional prediction.

[0004] Vision-based nutritional assessment methods aim to analyze food images by using various machine learning (ML) models and associate their visual features with specific nutritional components. Compared with traditional methods, these vision-based solutions are generally more objective and user-friendly. However, they also face challenges such as poor assessment accuracy and application limitations of food occlusion. For example, Dehais et al. rely on multi-view images to reconstruct the three-dimensional structure and estimate its volume. Then, they combine the volume to calculate the nutritional information of the food. However, this multi-view method is cumbersome and inefficient because they require users to capture images from specific angles. In contrast, monocular image-based methods provide an end-to-end detection method and show good performance. Specifically, Shao et al. adopted the Swin-Transformer model to implement an end-to-end lossless detection scheme, improving the accuracy of nutritional estimation through multi-task learning. However, estimating the nutritional components of food from monocular RGB images is an ill-posed problem. This is because mapping food to a monocular image usually results in the loss of key 3D details, which are very important for food component estimation. Therefore, Wang et al. proposed a calorie estimation method based on energy density mapping, which maps RGB images to food density based on pixel energy. This energy density-based method has advantages in estimating calories, but it does not bring obvious advantages in predicting the content of other nutritional components. Depth images can provide three-dimensional shape and size information of objects, which is of great significance for nutritional estimation. Recently, some researchers have considered using depth images (RGB-D images) to supplement the lack of spatial information in monocular RGB images. For example, Thames et al. spliced the depth map onto the fourth channel of the RGB image for model input, improving the model's prediction ability of nutritional content, but the gain in model performance brought by this simple splicing method is limited. Min et al. proposed a strategy combining a multi-scale nutritional estimation network based on CBAM attention and multi-modal fusion, achieving remarkable results. Research shows that the designed RGB-D fusion model is superior to simple linear fusion models. However, this general attention fusion method does not fully consider the information complementary effect between cross-modal features, lacks the perception and enhancement of depth information, and brings performance bottlenecks to food nutritional estimation models. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention provides a food nutritional assessment method and system based on multi-modal fusion.

[0006] In the first aspect, the present invention provides a food nutritional assessment method based on multi-modal fusion, including the following steps:

[0007] Step 1: Extract multi-scale features of the food RGB image and the RGB-D depth image in parallel through a multi-scale dual-branch Transformer encoder;

[0008] Step 2: Use a depth-aware enhanced attention fusion module to perform cross-modal fusion on the multi-scale features, including:

[0009] Multi-modal feature enhancement: Generate fused features through global and local attention mechanisms;

[0010] Volume parameter perception: Map three-dimensional geometric parameters from RGB-D features to generate a volume parameter perception matrix;

[0011] Step 3: Construct a nutrition prediction branch based on the fused features and output the contents of five macronutrients;

[0012] Step 4: Train the training dataset on the model in an end-to-end manner, and input the food RGB image and RGB-D image to be detected into the trained model for nutrition content prediction.

[0013] In an embodiment of the present invention, the multi-scale dual-branch Transformer encoder in Step 1 is composed of two parallel PVT v2 models. Each PVT v2 model in each branch is divided into four stages, and each stage includes a Patch embedding layer and a linear Transformer encoder layer;

[0014] The process of each branch model extracting features of the food image includes the following steps:

[0015] In the first stage, an image of size 224×224×3 is evenly divided into 56 blocks of size 4×4×3. Each small block is linearly embedded using a zero-padding convolutional layer, and the obtained sequence of embedding vectors is fed into the Transformer encoder layer to obtain the output features of the first stage;

[0016] The feature is continuously input into the Patch embedding layer and the linear Transformer encoder layer in the second stage for feature extraction;

[0017] The subsequent stages complete the feature extraction process using the same steps;

[0018] Finally, the multi-scale features extracted from the RGB image and the RGB-D image by the multi-scale dual-branch Transformer encoder are {R1, R2, R3, R4} and {D1, D2, D3, D4} respectively.

[0019] In an embodiment of the present invention, the depth perception enhanced attention fusion module in step 2 includes multi-modal feature enhancement and volume parameter perception; the depth perception enhanced attention fusion module is used to fuse the multi-scale features {R1, R2, R3, R4} and {D1, D2, D3, D4} extracted by the multi-scale dual-branch Transformer encoder in each stage in step 1;

[0020] The multi-modal feature enhancement part consists of attention mechanisms at global and local scales; for each stage, first, the features R of the RGB image i and the features D of RGB-D i are concatenated on the channel to obtain the initial fusion feature U i ; then, in the global-scale attention mechanism part, the global spatial information of U i is compressed into the channel descriptor using the global average pooling operation to obtain a feature of size 1×1×2C; then, two convolutional layers with a kernel size of 1x1 are used as the global channel context aggregator to compress and select features in the channel dimension; among them, the first 1×1 convolutional layer is followed by a Batch Normalization layer and a ReLU activation function layer to stabilize the training process and increase the ability of non-linear expression; the second 1×1 convolutional layer is only followed by a Batch Normalization layer; after the global attention mechanism, a feature of size 1×1×2C is obtained The global attention mechanism is described as:

[0021]

[0022] In the formula, i represents the i-th stage of the multi-scale dual-branch Transformer encoder, δ represents the ReLU activation function, represents the Batch Normalization layer.

[0023] In an embodiment of the present invention, the local-scale attention mechanism uses two convolutional layers with a kernel size of 1×1 to screen the local channel information of U i ; the first 1×1 convolutional layer is followed by a BatchNormalizatio n layer and a ReLU activation function layer; the second 1×1 convolutional layer is only followed by a BatchNormalization layer; after the local attention mechanism, a feature of size H×W×C is obtained The local-scale attention mechanism is described as:

[0024]

[0025] Next, the global calibration feature and local calibration features are added together to obtain fused features Then, the Sigmoid function is used to compress the feature values between 0 and 1 to obtain an attention coefficient matrix; finally, this coefficient matrix is multiplied by the RGB feature R i and the RG-D feature D i to obtain enhanced features enhanced by the multimodal feature enhancement part and

[0026]

[0027] In an embodiment of the present invention, the volume parameter perception part performs feature mapping on the RGB-D feature D through three parallel convolutional layers with a convolutional kernel size of 1×1 i ; among them, after each 1×1 convolutional layer, there is a Batch Normalization layer and a ReLU layer for non-linear learning and preserving important information; they compress the number of channels of the input feature D i to 1 and obtain three parameter matrices X, Y, and Z; based on geometric principles, the numerical values at each position in these three parameter matrices and are approximate representations of the length, width, and depth of the object in a specific receptive field area of the original image respectively; then, coefficient matrices and are generated respectively through the Sigmoid function; then, the three obtained coefficient matrices are multiplied to form a volume parameter perception matrix V; each value in this matrix represents the weight coefficient of the volume; finally, the volume parameter perception matrix is multiplied by the feature to obtain the enhanced feature W i ; finally, the output feature F i of the depth perception enhanced attention fusion module is the concatenated feature of the feature W i and the feature in the channel dimension; the features obtained by enhancing and fusing the multi-scale features {R1, R2, R3, R4} and {D1, D2, D3, D4} extracted by the multi-scale double-branch Transformer encoder in each stage in step 1 using the depth perception enhanced attention fusion module are {F 1, F 2, F 3, F4}.

[0028] In an embodiment of the present invention, the nutritional prediction branch in step 3 is constructed for the fused features {F 1, F 2, F3, Estimate the contents of five macronutrients for F4; first, use convolutional layers with a convolutional kernel size of 1×1 to extract information from each feature F in turn, and then use global average pooling operations to compress their features respectively; then, concatenate the four compressed features in the channel dimension; the concatenated features are input into five nutritional prediction branches in turn, each branch consists of two fully connected layers, and finally output the contents of calories, mass, carbohydrates, fat, and protein. i In one embodiment of the present invention, step 4 includes:

[0029] In one embodiment of the present invention, step 4 includes:

[0030] Step 4.1: During the model training process, first adjust the RGB images and RGB-D images in the original food dataset to a resolution of 238×238, and then center-crop them to a resolution of 224×224 to enlarge the food part in the images. Then, perform dataset augmentation through random vertical and horizontal flipping and contrast adjustment operations; at the same time, use the pre-trained weights on the ImageNet1K dataset to initialize the PVT v2 model;

[0031] Step 4.2: The model training process quantifies the difference between the model prediction value and the true value through a loss function, and continuously minimizes the loss through the backpropagation process; the loss function is the sum of the calorie regression loss L cal , the mass regression loss L mass and the regression losses of three macronutrients L carb , L fat and L protein ; the specific formula of the loss function is:

[0032] L = L cal + L mass + L carb + L fat + L protein (3)

[0033] wherein, the loss function of each nutrient uses the MAPE loss, and the formula of the calorie regression loss L cal is:

[0034]

[0035] In the formula, and respectively represent the prediction result and the actual value of the food calorie content;

[0036] The formula of the mass regression loss L mass is:

[0037]

[0038] In the formula, and respectively represent the predicted result and the actual value of the food quality;

[0039] The regression loss L of carbohydrates carb has the formula:

[0040]

[0041] In the formula, and respectively represent the predicted result and the actual value of the carbohydrate content of the food;

[0042] The regression loss L of fat fat has the formula:

[0043]

[0044] In the formula, and respectively represent the predicted result and the actual value of the fat content of the food;

[0045] The regression loss L of protein fat has the formula:

[0046]

[0047] In the formula, and respectively represent the predicted result and the actual value of the protein content of the food.

[0048] In a second aspect, the present invention provides a food nutrition evaluation system based on multi-modal fusion, including:

[0049] A multi-scale feature extraction module for parallelly extracting multi-scale features of a food RGB image and an RGB-D depth image through a multi-scale dual-branch Transformer encoder;

[0050] A cross-modal fusion module for cross-modal fusion of the multi-scale features by using a depth perception enhanced attention fusion module, including: multi-modal feature enhancement: generating fusion features through global and local attention mechanisms; volume parameter perception: mapping three-dimensional geometric parameters from RGB-D features to generate a volume parameter perception matrix;

[0051] An output module for constructing a nutrition prediction branch based on the fused features and outputting the contents of five macronutrients;

[0052] The nutrient content prediction module is used to train the training data set on the model in an end-to-end manner, and input the RGB image and RGB-D image of the food to be detected into the trained model for nutrient content prediction.

[0053] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the food nutrition assessment method based on multi-modal fusion are implemented.

[0054] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the food nutrition assessment method based on multi-modal fusion are implemented.

[0055] The beneficial effects achieved by the present invention are as follows:

[0056] 1. Considering the huge differences in food shape, size, and surface texture, the present invention effectively extracts multi-scale information of food by constructing a multi-scale dual-branch Trensformer structure, improving the performance of food nutrition estimation.

[0057] 2. The present invention utilizes the RGB-D depth modality corresponding to the food RGB image to enhance the perception of the volume information of different food types. A new fusion module (Depth Perception Enhancement Attention, DPEA) is designed to fuse the visual information in the food RGB image and the spatial physical features in the RGB-D image to generate discriminative fusion features, further improving the accuracy of nutrition estimation.

[0058] 3. The present invention adopts an end-to-end training method, which is more convenient and has less error compared with the multi-stage training method. These contributions make the present invention have important scientific and application values in the field of food nutrition assessment, providing new methods and tools for related research.

[0059] In summary, the present invention can effectively combine the complementary enhancement information in the two modalities of food RGB and RGB-D, thereby more accurately predicting the content of the five macronutrients in food. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 It is a schematic diagram of the overall process of a food nutrition assessment method based on multi-modal fusion provided by the present invention. DETAILED DESCRIPTION

[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0062] As Figure 1 shown, the present invention provides a food nutrition assessment method based on multi-modal fusion, including the following steps:

[0063] Step 1: Parallelly extract multi-scale features of the food RGB image and the RGB-D depth image through a multi-scale double-branch Transformer encoder;

[0064] Step 2: Use a depth perception enhanced attention fusion module (DPEA) to perform cross-modal fusion on the multi-scale features, including:

[0065] Multi-modal feature enhancement: Generate fusion features through global and local attention mechanisms;

[0066] Volume parameter perception: Map three-dimensional geometric parameters from RGB-D features to generate a volume parameter perception matrix;

[0067] Step 3: Build a nutrition prediction branch based on the fused features and output the contents of five macronutrients;

[0068] Step 4: Train the training data set on the model in an end-to-end manner, and input the food RGB image and RGB-D image to be detected into the trained model for predicting the nutrient content.

[0069] For further illustration, the structure of the multi-scale double-branch Transformer encoder in Step 1 is described as follows:

[0070] The multi-scale dual-branch Transformer encoder consists of two parallel Pyramid Vision Transformer (PVTv2) models; specifically, each branch's PVT v2 model is divided into four stages, and each stage includes two parts: a Patch embedding layer and a linear Transformer encoder layer; the process of each branch's model extracting features from food images is as follows: First, in the first stage, an image of size 224×224×3 is evenly divided into 56 blocks of size 4×4×3, and each small block is linearly embedded using a zero-padding convolutional layer, and the obtained sequence of embedding vectors is fed into the Transformer encoder layer to obtain the output features of the first stage; then, this feature is continuously input into the Patch embedding layer and the linear Transformer encoder layer in the second stage for feature extraction; the same steps are also used in the subsequent stages to complete the feature extraction process; finally, the multi-scale features extracted by the multi-scale dual-branch Transformer encoder from RGB images and RGB-D images are {R1, R2, R3, R4} and {D1, D2, D3, D4}, respectively.

[0071] For further illustration, the structure of the depth perception enhanced attention fusion module in step 2 is described as follows:

[0072] The depth perception enhanced attention fusion module (DPEA) includes two parts: multi-modal feature enhancement and volume parameter perception; DPEA is used to fuse the multi-scale features {R1, R2, R3, R4} and {D1, D2, D3, D4} extracted by the multi-scale dual-branch Transformer encoder in each stage in step 1;

[0073] The multi-modal feature enhancement part consists of attention mechanisms at global and local scales; specifically, for each stage, first, the feature R i of the RGB image and the feature D i of the RGB-D are concatenated on the channel to obtain the initial fusion feature U i ; then, in the global-scale attention mechanism part, the global average pooling (GAP) operation is used to process U iThe global spatial information is compressed into the channel descriptor to obtain features of size 1×1×2C; then, two convolutional layers (PWConv) with a kernel size of 1x1 are used as the context aggregator for the global channels, so as to compress and select features in the channel dimension; among them, the first 1×1 convolutional layer is followed by a Batch Normalization layer and a ReLU activation function layer to stabilize the training process and increase the ability of non-linear expression; the second 1×1 convolutional layer is only followed by a Batch Normalization layer; features of size 1×1×2C are obtained through the global attention mechanism The global attention mechanism can be described as:

[0074]

[0075] In the formula, i represents the i-th stage of the multi-scale double-branch Transformer encoder, δ represents the ReLU activation function, represents the Batch Normalization layer;

[0076] The attention mechanism at the local scale directly uses two convolutional layers with a kernel size of 1×1 to filter the local channel information of U i ; the first 1×1 convolutional layer is followed by a Batch Normalization layer and a ReLU activation function layer; the second 1×1 convolutional layer is only followed by a Batch Normalization layer; features of size H×W×C are obtained through the local attention mechanism The attention mechanism at the local scale can be described as:

[0077]

[0078] Next, the globally calibrated features and the locally calibrated features are added together to obtain the fused features Then, the Sigmoid function is used to compress the feature values of between 0 and 1 to obtain the attention coefficient matrix; finally, the coefficient matrix is multiplied by the RGB feature R i and the RG-D feature D i respectively to obtain the enhanced features enhanced by the multi-modal feature enhanced part and

[0079]

[0080] The volume parameter perception part uses three parallel convolutional layers with a kernel size of 1×1 to process the RGB-D feature D iPerform feature mapping; among them, after each 1×1 convolutional layer, there is a Batch Normalization layer and a ReLU layer for non-linear learning and preserving important information; they compress the number of channels of the input feature D i to 1 and obtain three parameter matrices X, Y, and Z; based on geometric principles, the values at each position in these three parameter matrices and are approximate representations of the length, width, and depth (i.e., height) of the objects in a specific receptive field area of the original image respectively; then, coefficient matrices and are generated respectively through the Sigmoid function; Next, multiply the three obtained coefficient matrices to form a volume parameter perception matrix V; each value in this matrix represents the weight coefficient of the volume; finally, multiply the volume parameter perception matrix by the feature to obtain the enhanced feature W i ; ultimately, the output feature F i of the depth perception enhanced attention fusion module (DPEA) is the concatenated feature of the feature W i and the feature in the channel dimension; that is, using the DPEA module to enhance and fuse the multi-scale features {R1, R2, R3, R4} and {D1, D2, D3, D4} extracted by the multi-scale double-branch Transformer encoder in each stage in step 1, the resulting feature is {F 1, F 2, F 3, F4};

[0081] For further illustration, the structure of the nutrition prediction branch in step 3 is described as follows:

[0082] Construct a nutrition prediction branch to estimate the contents of five macronutrients for the fused feature {F 1, F 2, F 3, F4}; First, use convolutional layers with a kernel size of 1×1 to extract information from each feature F i in turn, and then use global average pooling (GAP) operations to compress their features respectively; Next, concatenate the four compressed features in the channel dimension; the concatenated feature is input into five nutrition prediction branches in turn, and each branch consists of two fully connected layers, and finally, the contents of calories, mass, carbohydrates, fat, and protein can be output.

[0083] For further illustration, the steps of the end-to-end training of the model in step 4 are described as follows:

[0084] Step 4.1: During the model training process, first, the RGB images and RGB-D images in the original food dataset need to be adjusted to a resolution of 238×238, and then center-cropped to a resolution of 224×224 to magnify the food part in the images. Next, dataset augmentation is performed through operations such as random vertical and horizontal flipping, contrast adjustment, etc.; at the same time, the PVT v2 model is initialized using the pre-trained weights on the ImageNet1K dataset; in addition, the adam optimizer is used during the training process, with a batch size of 8, a total of 200 Epochs are trained, and the learning rate is 1×e -4 , and the weight decay is 1×e -5 ;

[0085] Step 4.2: The model training process quantifies the difference between the model prediction value and the true value through the loss function and continuously minimizes the loss through the backpropagation process; the loss function is the calorie regression loss L cal , the quality regression loss L mass and the regression losses of three macronutrients L carb , L fat and L protein ; the specific formula of the loss function is:

[0086] L = L cal + L mass + L carb + L fat + L protein (3)

[0087] where the loss function for each nutrient uses the MAPE loss. For example, the formula for the calorie regression loss L cal is:

[0088]

[0089] In the formula, and represent the predicted result and the actual value of the food calorie content, respectively.

[0090] The formula for the quality regression loss L mass is:

[0091]

[0092] In the formula, and represent the predicted result and the actual value of the food quality, respectively.

[0093] The formula for the carbohydrate regression loss L carb is:

[0094]

[0095] In the formula, and respectively represent the predicted result and the actual value of the carbohydrate content of the food.

[0096] The formula for the fat regression loss L fat is as follows:

[0097]

[0098] In the formula, and respectively represent the predicted result and the actual value of the fat content of the food.

[0099] The formula for the protein regression loss L fat is as follows:

[0100]

[0101] In the formula, and respectively represent the predicted result and the actual value of the protein content of the food.

[0102] As shown in Table 1 below, Table 1 is a comparison table of the mean absolute percentage errors of the prediction results of the present invention and other nutritional estimation models. According to Table 1, in the comparison benchmark using only RGB images, the nutritional prediction by the method based on transfer learning verifies that the multi-scale double-branch Transformer encoder (i.e., the PVT v2 model) used in the present invention is significantly superior to other backbone models. Among them, in terms of the average prediction results of the five nutritional contents, the average error rate of PVT v2 is only 19.35%, which is reduced by up to 9.55% compared with other models; in the multi-modal prediction method using both RGB and RGB-D images, the present invention achieves the lowest error rates in calorie, mass, carbohydrate, and protein predictions, and the average error rate is reduced by up to 3.36% compared with other methods.

[0103] Table 1 Comparison table of the mean absolute percentage errors of the prediction results of the present invention and other nutritional estimation models

[0104]

[0105] In addition, the present invention also provides a food nutrition evaluation system based on multi-modal fusion, including:

[0106] A multi-scale feature extraction module for parallelly extracting multi-scale features of a food RGB image and an RGB-D depth image through a multi-scale double-branch Transformer encoder;

[0107] A cross-modal fusion module, which is used to perform cross-modal fusion on the multi-scale features by using a depth perception enhanced attention fusion module, includes: multi-modal feature enhancement: generating fused features through global and local attention mechanisms; volume parameter perception: mapping three-dimensional geometric parameters from RGB-D features to generate a volume parameter perception matrix;

[0108] An output module, which is used to construct a nutrition prediction branch based on the fused features and output the contents of five macronutrients;

[0109] A nutrition content prediction module, which is used to train the training data set on the model in an end-to-end manner, and input the RGB image and RGB-D image of the food to be detected into the trained model for nutrition content prediction.

[0110] A food nutrition assessment method and system based on multi-modal fusion provided by the present invention have unique advantages: First, the present invention takes into account the huge differences in food shape, size and surface texture, and effectively extracts the multi-scale information of food by constructing a multi-scale dual-branch Trensformer structure, improving the performance of food nutrition estimation. Second, the present invention uses the RGB-D depth modality corresponding to the food RGB image to enhance the perception of the volume information of different food types. A new fusion module (depth perception enhanced attention, DPEA) is designed to fuse the visual information in the food RGB image and the spatial physical features in the RGB-D image to generate discriminative fused features, further improving the accuracy of nutrition estimation. Finally, the present invention adopts an end-to-end training method, which is more convenient and has less error compared with the multi-stage training method. These contributions make the present invention have important scientific and application values in the field of food nutrition assessment, providing new methods and tools for related research.

[0111] In addition, the present invention also provides a computer device, which may include a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the processor executes the steps of the food nutrition assessment method based on multi-modal fusion as described in any of the above embodiments.

[0112] For the working process, working details and technical effects of the computer device provided in this embodiment, reference may be made to the embodiments of the food nutrition assessment method based on multi-modal fusion in the above text, which will not be elaborated here.

[0113] In addition, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the food nutrition assessment method based on multimodal fusion in any of the above embodiments are implemented. Wherein, the computer-readable storage medium refers to a carrier for storing data, and may include, but is not limited to, floppy disks, optical discs, hard disks, flash memories, USB flash drives, and / or memory sticks, etc. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0114] For the working process, working details, and technical effects of the computer-readable storage medium provided in this embodiment, reference may be made to the embodiments of the food nutrition assessment method based on multimodal fusion in the foregoing text, and details are not repeated herein.

[0115] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Wherein, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM).

[0116] Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing embodiments, or equivalently replace some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A food nutrition assessment method based on multimodal fusion, characterized in that, It includes the following steps: Step 1: Parallelly extract multi-scale features of the food RGB image and the RGB-D depth image through a multi-scale double-branch Transformer encoder; Step 2: Use a depth-aware enhanced attention fusion module to perform cross-modal fusion on the multi-scale features, including: Multi-modal feature enhancement: Generate fused features through global and local attention mechanisms; Volume parameter perception: Map three-dimensional geometric parameters from RGB-D features to generate a volume parameter perception matrix; Step 3: Construct a nutrition prediction branch based on the fused features and output the contents of five macronutrients; Step 4: Train the training dataset on the model in an end-to-end manner, and input the food RGB image and RGB-D image to be detected into the trained model for nutrition content prediction.

2. The method for evaluating food nutrition based on multi-modal fusion according to claim 1, wherein The multi-scale double-branch Transformer encoder in Step 1 consists of two parallel PVT v2 models. Each PVT v2 model in each branch is divided into four stages, and each stage includes a Patch embedding layer and a linear Transformer encoder layer; The process of each branch model extracting features of the food image includes the following steps: In the first stage, an image of size 224×224×3 is evenly divided into 56 blocks of size 4×4×3. A zero-padding convolutional layer is used to linearly embed each small block, and the obtained sequence of embedding vectors is fed into the Transformer encoder layer to obtain the output features of the first stage; The feature is continuously input into the Patch embedding layer and the linear Transformer encoder layer in the second stage for feature extraction; The subsequent stages complete the feature extraction process using the same steps; Finally, the multi-scale features extracted from the RGB image and the RGB-D image by the multi-scale double-branch Transformer encoder are {R1, R2, R3, R4} and {D1, D2, D3, D4} respectively.

3. The method for food nutrition assessment based on multimodal fusion according to claim 2, wherein The depth-aware enhanced attention fusion module in Step 2 includes multi-modal feature enhancement and volume parameter perception; the depth-aware enhanced attention fusion module is used to fuse the multi-scale features {R1, R2, R3, R4} and {D1, D2, D3, D4} extracted by the multi-scale double-branch Transformer encoder in each stage in Step 1; The multi-modal feature enhancement part consists of attention mechanisms at the global and local scales; for each stage, first, the features R of the RGB image i and the features D of the RGB-D i are concatenated on the channel dimension to obtain the initial fusion feature U i ; then, in the global-scale attention mechanism part, the global average pooling operation is used to compress the global spatial information of U i into the channel descriptor to obtain features of size 1×1×2C; then, two convolutional layers with a kernel size of 1x1 are used as the global-channel context aggregator to compress and select features in the channel dimension; among them, the first 1×1 convolutional layer is followed by a Batch Normalization layer and a ReLU activation function layer to stabilize the training process and increase the ability of non-linear expression; the second 1×1 convolutional layer is only followed by a Batch Normalization layer; features of size 1×1×2C are obtained after the global attention mechanism The global attention mechanism is described as: where \(i\) represents the \(i\)-th stage of the multi-scale double-branch Transformer encoder, \(\delta\) represents the ReLU activation function, represents the Batch Normalization layer.

4. A food nutrition assessment method based on multimodal fusion according to claim 3, characterized in that, The local-scale attention mechanism uses two convolutional layers with a kernel size of 1×1 to filter the local channel information of U i For local channel information screening, behind the first 1×1 convolutional layer is a Batch Normalization layer and a ReLU activation function layer; behind the second 1×1 convolutional layer is only a Batch Normalization layer; The features of size H×W×C are obtained through the local attention mechanism The attention mechanism of the local scale is described as follows: Next, add the global calibration feature and the local calibration feature to obtain the fused feature Then, use the Sigmoid function to compress the feature values between 0 and 1 to obtain the attention coefficient matrix; finally, multiply this coefficient matrix by the RGB feature R i and the RG-D feature D i to obtain the features enhanced by the multi-modal feature enhancement part and 5. The food nutrition assessment method based on multimodal fusion according to claim 4, characterized in that, The volume parameter perception part performs feature mapping on the RGB-D feature D through three parallel convolutional layers with a convolutional kernel size of 1×1 i ; among them, after each 1×1 convolutional layer, there is a Batch Normalization layer and a ReLU layer for non-linear learning and preserving important information; they compress the number of channels of the input feature D i to 1 and obtain three parameter matrices X, Y, and Z; based on geometric principles, the values at each position in these three parameter matrices and are approximate representations of the length, width, and depth of the object in a specific receptive field area of the original image respectively; then, coefficient matrices and are generated through the Sigmoid function respectively. Then, the three obtained coefficient matrices are multiplied to form a volume parameter perception matrix V; each value in this matrix represents the weight coefficient of the volume; finally, the volume parameter perception matrix is multiplied by the feature to obtain the enhanced feature W i ; ultimately, the output feature F i of the depth perception enhanced attention fusion module is the concatenated feature of the feature W i and the feature in the channel dimension; the features {F1, F2, F3, F4} after enhancing and fusing the multi-scale features {R1, R2, R3, R4} and {D1, D2, D3, D4} extracted by the multi-scale double-branch Transformer encoder in each stage in step 1 using the depth perception enhanced attention fusion module are obtained.

6. The food nutrition assessment method based on multimodal fusion according to claim 5, characterized in that Construct the nutrition prediction branch in Step 3 to estimate the contents of five macronutrients for the fused features {F 1, F 2, F 3, F4}; First, use convolutional layers with a kernel size of 1×1 to extract information from each feature F i respectively, and then use global average pooling operations to compress their features; Next, concatenate the four compressed features in the channel dimension; The concatenated features are sequentially input into five nutrition prediction branches, each branch consists of two fully connected layers, and finally the contents of calories, mass, carbohydrates, fats and proteins are output.

7. A method for food nutrition assessment based on multi-modal fusion according to claim 6, characterized in that Step 4 includes: Step 4.1: During the model training process, first adjust the RGB image and the RGB-D image in the original food dataset to a resolution of 238×238, and then center-crop them to a resolution of 224×224 to enlarge the food part in the image. Then, perform dataset enhancement through random vertical and horizontal flipping and contrast adjustment operations; at the same time, use the pre-trained weights on the ImageNet1K dataset to initialize the PVT v2 model; Step 4.2: In the model training process, the loss function is used to quantify the difference between the model prediction value and the true value, and the backpropagation process is used to continuously minimize the loss; the loss function is the calorie regression loss L cal , the quality regression loss L mass and the regression losses of three macronutrients L carb , L fat and L protein ; the specific formula of the loss function is: L = L cal + L mass + L carb + L fat + L protein (3) Among them, the loss function of each nutrient uses the MAPE loss, and the formula for the calorie regression loss L cal is as follows: In the formula, and respectively represent the predicted result and the actual value of the food calorie content; Quality regression loss L mass The formula is as follows: In the formula, and respectively represent the predicted result and the actual value of the food quality; Carbohydrate regression loss L carb The formula is as follows: Wherein, and respectively represent the predicted result and the actual value of the food carbohydrate content; Fat regression loss L fat The formula is as follows: In the formula, and represent the predicted result and the actual value of the food fat content respectively; Protein regression loss L fat The formula is as follows: In the formula, and represent the predicted result and the actual value of the food protein content respectively.

8. A food nutrition assessment system based on multimodal fusion, characterized in that, It includes: A multi-scale feature extraction module for parallelly extracting multi-scale features of the food RGB image and the RGB-D depth image through a multi-scale double-branch Transformer encoder; A cross-modal fusion module for cross-modal fusion of the multi-scale features by using a depth perception enhanced attention fusion module, including: multi-modal feature enhancement: generating a fusion feature through a global and local attention mechanism; volume parameter perception: mapping three-dimensional geometric parameters from RGB-D features to generate a volume parameter perception matrix; An output module for constructing a nutrition prediction branch based on the fused features and outputting the contents of five macronutrients; A nutrition content prediction module for training a training data set on the model in an end-to-end manner and inputting a food RGB image and an RGB-D image to be detected into the trained model for nutrition content prediction.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the multi-modal fusion-based food nutrition assessment method according to any one of claims 1-7 are implemented.

10. A computer-readable storage medium, having stored thereon a computer program, wherein when the computer program is executed by a processor, the steps of the multi-modal fusion-based food nutrition assessment method according to any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Food nutritional ingredient content prediction method and system based on cross-modal attention mechanism

    CN114529790A

  • RGB-D saliency target detection method based on lightweight cross-modal fusion network

    CN116486112A

  • Food evaluation method and device, electronic equipment and storage medium

    CN116883991A

  • Food nutrition assessment method based on multi-scale information fusion

    CN118644850A

  • Efficient RGB-D saliency detection method using multiple information complementation

    CN118982655A

Cited By

  • Multi-dimensional intelligent food material analysis system based on artificial intelligence

    CN120976917A

  • Anaerobic fermentation gas production prediction method based on multi-modal image features

    CN121236515A

  • Cattail pollen content prediction method and system based on hyperspectral double-flow multi-scale CNN (Convolutional Neural Network)

    CN121527631A