A food quality risk early warning method and system
By combining visible light images and near-infrared imaging, and employing multi-scale feature fusion and Transformer networks, the problems of low accuracy and high cost in food quality detection have been solved, enabling efficient and stable detection of various foods and reducing production costs.
Patent Information
- Application Number
- CN202510366906.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-03-26
AI Technical Summary
Existing food quality testing methods suffer from problems such as long testing time, complex operation, high cost, low accuracy, and unstable results of single feature detection, making it impossible to achieve unified testing of multiple foods.
By combining visible light images and near-infrared imaging, and through an improved multi-scale feature fusion network and Transformer network, dual-modal dual-scale fusion features of food are extracted for food classification and quality risk early warning.
It improves the accuracy and stability of food quality testing, and can simultaneously test the appearance and internal quality of food, reducing the defect rate, lowering production costs, and is suitable for unified testing of various foods.
Smart Images

Figure CN120236277B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of food detection, and particularly relates to a food quality risk early warning method and system. BACKGROUND
[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute the prior art.
[0003] With the complication of the food industry chain, food quality problems are the focus of social attention. Food may be affected by various factors during production, processing, transportation and sales, leading to frequent quality problems. For example, microbial contamination, pesticide residues and abnormal states in food can seriously affect the health of consumers. Therefore, during food production, attention should be paid to food quality detection.
[0004] Traditional food quality detection methods mainly rely on chemical analysis, microbial culture and other technologies. Although these methods can provide relatively accurate results, they have problems such as long detection time, complex operation and high cost.
[0005] There are also quality detection methods such as image processing and biosensor technology in the prior art, but the detection effect of a single feature has a large error. For example, image recognition technology using a high-definition camera can only observe surface deterioration of food, and cannot accurately detect problems such as pesticide residues and excessive internal microorganisms. The detection method using spectroscopy technology has certain requirements for the transparency, color and moisture content of the sample. For samples with poor transparency, dark color or high moisture content, the detection effect may not be good. The non-uniformity of food may affect the propagation of light, leading to instability of the detection results. Therefore, quality detection of a single feature has the problems of low accuracy and instability. Secondly, a single feature not only loses part of the features, but also ignores the correlation between the appearance and the internal state of the food. Features such as surface scratches and microbial attachment are mutually influenced by internal deterioration and other abnormalities.
[0006] In addition, from the perspective of the detection object, most existing detection models are for the detection of a specific fixed food, and can only detect a single food product on a production line, and cannot realize unified detection of food samples from multiple production lines. SUMMARY
[0007] In order to overcome the above-mentioned deficiencies of the prior art, the present application provides a food quality risk early warning method and system, which fuses multiple images to cover the overall characteristics of the food, improves the detection accuracy and stability, proposes multiple fusion of multi-scale features, retains the shallow edge features while extracting the deep detail features, balances the global features and local features, and reduces the feature loss, in addition, proposes similarity analysis of multiple image features to fuse the features, increases the correlation analysis of food appearance abnormalities and internal abnormalities, improves the representation ability of the fused features, and further improves the accuracy of the quality risk.
[0008] In order to achieve the above-mentioned purpose, one or more embodiments of the present application provide the following technical solutions:
[0009] In the first aspect, a food quality risk early warning method is disclosed, comprising:
[0010] obtaining a visible light image and near-infrared imaging of the food;
[0011] based on the visible light image of the food, using an improved multi-scale feature fusion network to determine the category and segmentation mask of the food;
[0012] based on the segmentation mask, performing binary mask filtering on the visible light image and the near-infrared imaging to obtain image data, calling different food quality detection models according to the category of the food, and inputting the image data into the food quality detection model to determine the quality risk of the food;
[0013] wherein the food quality detection model is an improved Transformer network, the improved Transformer network comprises an encoder, a feature mapping module and a multilayer perception machine; the encoder extracts double-modal double-scale fusion features, second visible light features and second near-infrared features of the image data, the feature mapping module performs correlation analysis on the double-modal double-scale fusion features, the second visible light features and the second near-infrared features to obtain re-fusion features, and the multilayer perception machine outputs the food quality risk based on the re-fusion features.
[0014] In the second aspect, a food quality risk early warning system is disclosed, comprising:
[0015] a data acquisition module configured to obtain a visible light image and near-infrared imaging of the food;
[0016] an intelligent classification module configured to determine the category and segmentation mask of the food based on the visible light image of the food using an improved multi-scale feature fusion network;
[0017] The quality detection module is configured to: perform binary mask filtering on the visible light image and the near-infrared image based on the segmentation mask to obtain image data, call different food quality detection models according to the category of the food, and input the image data into the food quality detection model to determine the quality risk of the food.
[0018] The food quality detection model is an improved Transform network, and the improved Transform network comprises an encoder, a feature mapping module and a multilayer perception machine; the encoder extracts double-modal double-scale fusion features, second visible light features and second near-infrared features of the image data, the feature mapping module performs correlation analysis on the double-modal double-scale fusion features, the second visible light features and the second near-infrared features to obtain re-fusion features, and the multilayer perception machine outputs the quality risk of the food based on the re-fusion features.
[0019] In a third aspect, an electronic device is disclosed, comprising a memory and a processor, and computer instructions stored in the memory and running on the processor, when the computer instructions are run by the processor, the steps of the food quality risk early warning method are completed.
[0020] In a fourth aspect, a computer readable storage medium is disclosed, for storing computer instructions, when the computer instructions are executed by the processor, the steps of the food quality risk early warning method are completed.
[0021] Compared with the prior art, the beneficial effects of the present application are:
[0022] The present application provides a food quality risk early warning method, which fuses visible light images and near-infrared images, detects appearance and internal quality at the same time, improves detection accuracy, and enables enterprises to more efficiently control production processes and reduce the rate of defective products.
[0023] In the food classification model, multi-scale feature fusion is adopted when each image is subjected to feature extraction, so that global features and detail features are retained, loss of basic features and edge features is avoided, and therefore foods with large appearance differences can be more accurately and quickly identified, and the risk of falling into a small detail error is avoided.
[0024] The food classification model can output a segmentation mask while obtaining the food category, thereby providing a data processing method for the detection model, processing the input image of the detection model into image data containing only the food body, and reducing the influence of the background on quality detection.
[0025] The multi-modal image fusion and the fusion method considering the correlation features of two images are proposed in the quality detection model, the features for detection include single modal image features and fusion features of two modalities, the visible light image can effectively represent the appearance state of the food, the near-infrared imaging can effectively represent the internal state of the food, the fusion features adopt the correlation mapping method to extract the correlation matrix of two images, so as to extract the correlation features of the internal and appearance, the correlation feature fusion is used to detect the food quality, the complementary information of different modalities can be fully utilized, the classification performance is improved, the data efficiency and robustness are enhanced, and the correlation between the quality risks of multi-modal data is revealed, and the accuracy of quality detection is improved.
[0026] The method can be used for detecting quality risks at any link of food production, quality problems in production can be found in time through detection, unqualified products can be avoided to enter subsequent links, and raw material waste and production cost can be reduced.
[0027] The advantages of the additional aspects of the application will be partially given in the following description, partially become obvious from the following description, or be known by the practice of the application. BRIEF DESCRIPTION OF DRAWINGS
[0028] The drawings accompanying the specification of the application form a part of the specification and serve to further illustrate the illustrative embodiments of the application and to explain the principles of the application.
[0029] Figure 1 The flow chart of the food quality risk early warning method described in the embodiment one of the application.
[0030] Figure 2 The intelligent classification model structure diagram described in the embodiment one of the application.
[0031] Figure 3 The quality detection model structure diagram described in the embodiment one of the application. DETAILED DESCRIPTION
[0032] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the application. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the application belongs.
[0033] It should be noted that the terms used herein are only for the purpose of describing specific embodiments, and are not intended to limit the exemplary embodiments according to the application.
[0034] The embodiments in the application and the features in the embodiments can be combined with each other without conflict.
[0035] Embodiment one
[0036] In one or more embodiments, a food quality risk early warning method is disclosed, comprising the following steps:
[0037] Step S1: Obtain visible light images and near-infrared images of the food.
[0038] The visible light images can be obtained by a CMOS / CCD sensor, a digital camera, a surveillance camera, an industrial camera, etc. high-definition shooting equipment, which is not limited here;
[0039] The near-infrared images can be obtained by a short-wave infrared detector, a short-wave infrared camera, etc. equipment, which is not limited here;
[0040] The visible light image acquisition device and the near-infrared image acquisition device are installed in the same position and as close as possible to each other. Alternatively, a device capable of simultaneously acquiring visible light images and near-infrared images can be used. The specific device model is not limited here, as long as the visible light image and the near-infrared image have the same shooting distance and angle.
[0041] The visible light images can detect the appearance defects of the food, such as scratches, dents, color changes, and mold and foreign matter on the surface, while the near-infrared images can detect the internal composition, structure, and chemical composition and potential harmful substances in the food, such as pesticide residues and heavy metals. By combining the two image technologies, defects or pollutants that cannot be found by a single technology can be detected, for example, near-infrared images can detect internal moisture content, bruises, or impurities through packaging.
[0042] Step S2: Based on the visible light images of the food, an intelligent classification model is used to classify the food into specific categories. In this embodiment, the food is mainly classified into three categories, including poultry eggs, chickens and ducks, and aquatic products. When the type of food produced on the actual production line changes, other categories can be divided according to actual needs.
[0043] The intelligent classification model is specifically an improved multi-scale feature fusion network. The intelligent classification model is a network model trained by a training set constructed from historical data. Historical visible light image data of poultry eggs, chickens and ducks, and aquatic products are collected, and category labels are calibrated to construct a training data set and train the model.
[0044] The structure of the improved multi-scale feature fusion network includes a backbone network, a feature fusion network, and a detection network connected in sequence, as shown in Figure 2 .
[0045] The backbone network includes a plurality of single convolution modules and a plurality of multi-convolution modules, specifically a first single convolution module, a first multi-convolution module, a second single convolution module, a second multi-convolution module, and a third single convolution module connected in sequence.
[0046] The single convolution module comprises a convolution layer conv, a batch normalization layer BN and an activation function layer SiLU, and the expression is:
[0047]
[0048] wherein, is the output feature of the single convolution module, is an activation function, is batch normalization, is a convolution layer, is the input food feature map of the single convolution module.
[0049] The single convolution module is used for feature extraction of the food image. With the stacking of the single-layer convolution module, the network captures the local information of the food image, refines the food edge frame, and improves the expression of the detailed features.
[0050] The multi-convolution module comprises a first convolution layer, a second convolution layer, a third convolution layer, a first connection layer and a fourth convolution layer. The input feature map is input into the first convolution layer to obtain a first feature map. The first feature map is input into the second convolution layer to obtain a second feature map. The second feature map is input into the third convolution layer to obtain a third feature map. The first feature map, the second feature map and the third feature map are input into the first connection layer to be spliced to obtain a fourth feature map. The fourth feature map is input into the fourth convolution layer to obtain the output of the multi-convolution module.
[0051]
[0052] wherein, is the output feature of the multi-convolution module, is the input food feature map of the multi-convolution module.
[0053] The multi-convolution module extracts detailed features through three convolution layers in sequence, and then fuses three feature maps of different depths to enable the multi-convolution module to recognize feature information from each scale of the food image. This avoids the loss of shallow features and increases the mutual semantic relationship between the shallow features and the deep features, thereby improving the feature representation capability.
[0054] The feature fusion network is connected behind the backbone network and includes a down-sampling layer, a second connection layer, a third multi-convolution module, a third connection layer, and a fifth convolution layer connected in sequence. The input of the down-sampling layer is the output feature of the backbone network. The output feature of the second multi-convolution module is also input into the second connection layer. The second connection layer splices the output feature of the second multi-convolution module and the output feature of the down-sampling layer. The output feature of the first multi-convolution module is also input into the third connection layer. The third connection layer splices the output feature of the first multi-convolution module and the output feature of the third multi-convolution module. According to actual production needs, the target of the intelligent classification model is set to three types of poultry eggs, chickens and ducks, and aquatic products. Poultry eggs, chickens and ducks, and aquatic products are quite different in appearance. Therefore, the output feature maps of the first multi-convolution module and the second multi-convolution module are spliced with deep features through jump connection, so as to retain the most basic global features of each type of food and avoid the adverse effects of too local features on classification. Multi-level multi-scale feature fusion not only retains the shape difference of the three types of images globally, but also gradually deepens the convolution module to extract detailed information, provide local features for food bounding box detection, make the food bounding box more refined, improve the food positioning and boundary precision, and fuse the feature maps of different scales in the backbone network to improve the accuracy of food classification and the accuracy of the mask, so as to affect the subsequent quality detection model, eliminate irrelevant background for the quality detection model, and avoid the influence of the conveyor belt or the equipment for receiving food on the food quality detection.
[0055] The detection network adopts a coupled detection mode to simultaneously obtain a predicted category and a target bounding box. Specifically, the detection network includes a fourth single convolution module, a fifth single convolution module, a sixth convolution layer, a seventh convolution layer, a fourth connection layer, a sixth single convolution module, and an output layer. The output feature of the feature fusion network is simultaneously output to the fourth single convolution module and the fifth single convolution module. One branch is connected in sequence with the fourth single convolution module, the sixth convolution layer, and the seventh convolution layer. The other branch is the fifth single convolution module. The two branches are jointly input into the fourth connection layer, and then output through the sixth single convolution module. Finally, the food category and the segmentation mask are output by the output layer.
[0056] The output of the food category makes the model not only applicable to a single production line of various foods, without the need for manual distinction of food categories to set detection software, but also applicable to sampling detection of multiple foods. Different types of foods on different production lines are randomly sampled and mixed into one detection line to reduce the arrangement of detection lines. The model can automatically identify different types of foods on the detection line to achieve fully automated detection and reduce costs.
[0057] The segmentation mask output obtains the range of the food body image, shields the background noise, the auxiliary quality detection model extracts features for the food, reduces the calculation amount, and improves the efficiency of the overall food quality risk early warning.
[0058] S3: According to the category label obtained in step S2, the quality detection model of the corresponding category is called to perform image detection on the target food.
[0059] Based on the segmentation mask preprocessing of the visible light image and the near-infrared image, the image data is obtained, different food quality detection models are called according to the category of the food, and the quality risk of the food is determined based on the image data.
[0060] S3-1: First, the segmentation mask is interpolated to restore the mask to the original image size, and then the mask is superimposed on the visible light image and the near-infrared image. A binary mask filter is used to retain the image with a mask value of 1 and set the image with a mask value of 0 to 0 to obtain visible light image data and near-infrared image data containing only the food body. Visible light image data and near-infrared image data containing only the food body together constitute image data, which is input into the food quality detection model. Thus, the influence of environmental factors on the detection effect during food quality detection is eliminated, and the detection model focuses entirely on the food.
[0061] S3-2: The processed visible light image data and near-infrared image data are input into the quality detection model to obtain a quality score.
[0062] The quality detection model includes a poultry quality detection model pre-trained using historical visible light image data and historical near-infrared image data of poultry, a chicken and duck quality detection model pre-trained using historical visible light image data and historical near-infrared image data of chicken and duck, and an aquatic product quality detection model pre-trained using historical visible light image data and historical near-infrared image data of aquatic products. When the classification result in step S2 is poultry, the poultry quality detection model is called; when the classification result in step S2 is chicken and duck, the chicken and duck quality detection model is called; and when the classification result in step S2 is aquatic product, the aquatic product quality detection model is called. Thus, the appearance of the three different food types is greatly different, which cannot be simultaneously trained to extract features, thereby improving the accuracy of quality detection. The quality detection models of the three food types are of the same structure, and are only models trained with different parameters.
[0063] Specifically, as shown in Figure 3 The quality detection model uses an improved transformer network, including an encoder, a feature mapping module, and a multilayer perceptron. By designing a cross-modal fusion transformer network, one image can receive information from another image, thereby achieving more effective information fusion.
[0064] Firstly, the encoder includes two Restormer layers, two Transformer layers, a multi-layer perception (MLP) and a weighted connection layer connected in sequence. The encoder extracts shallow and deep features of a single food image on one hand, and fuses the visible light image and the near-infrared image to obtain correlation features on the other hand. The MLP model is used as a food feature extractor.
[0065] Firstly, the visible light image data I im and the near-infrared image data I in are input into the network to extract features.
[0066] The first Restormer layer outputs first visible light features and first near-infrared features , wherein is a shallow feature map of the food visible light image data, is a shallow feature map of the near-infrared imaging data, and C, H and W represent the channel number, height and width of the feature map. Since the shallow features and contain the basic details of the image, more edge details of the image are extracted.
[0067] Then, the first visible light features and the first near-infrared features are input into the two Transformer layers through the second Restormer layer to output intermediate features and , which capture long-distance feature information and are used to extract semantic information of deep features.
[0068] The features and are further obtained by the MLP to obtain deep features of each image, including second visible light features and second near-infrared features . Meanwhile, the deep features and are input into the weighted connection layer together with the shallow features and to obtain bimodal and double-scale fusion features, and the expression is as follows:
[0069]
[0070] wherein, is the bimodal and double-scale fusion feature; , , , are feature weight coefficients.
[0071] The encoder adopts a loss function of bimodal double-scale feature fusion:
[0072]
[0073] wherein, is a bimodal double-scale loss function; is a correlation degree of the first visible light feature and the first near-infrared feature; is a correlation degree of the second visible light feature and the second near-infrared feature; is a parameter greater than zero, that is, a positive number infinitely close to 0; is a structural similarity index SSIM.
[0074] Through the joint action of the above loss function and the weighted connection layer, the fusion coefficients of the shallow features and the deep features of the two images are trained to be optimal, which not only ensures that the images can mine semantic information and retain more detailed information, but also ensures that the basic edge information is not lost, balances the overall and local features, and makes the visible light image and the near-infrared image play their maximum roles.
[0075] The obtained bimodal double-scale fusion features , the second visible light feature , the second near-infrared feature , and the refusion features are obtained through correlation mapping:
[0076] The autocovariance matrix of each group of features is calculated , , ;
[0077] wherein, is the autocovariance matrix of the bimodal double-scale fusion features, is the autocovariance matrix of the second visible light feature, is the autocovariance matrix of the second near-infrared feature, is a transpose.
[0078] The cross-covariance matrix is calculated , , ;
[0079] wherein, is the cross-covariance matrix of the bimodal double-scale fusion features and the second visible light feature, is the cross-covariance matrix of the second visible light feature and the second near-infrared feature, is the cross-covariance matrix of the second near-infrared feature and the bimodal double-scale fusion features.
[0080] The mapping direction is solved, and the projection direction that maximizes the correlation of the three groups of features is found:
[0081]
[0082] wherein, , , is a projection direction of the feature correlation maximization; , , is a correlation coefficient.
[0083] The mapping feature is:
[0084]
[0085] wherein, is a visible light mapping feature, is a near-infrared mapping feature, is a bimodal and double-scale fusion mapping feature, , and is a projection matrix, and , , .
[0086] The mapping feature is spliced according to the third dimension to obtain a re-fusion feature :
[0087]
[0088] wherein, is a third dimension splicing function.
[0089] In this way, the features of the visible light image and the features of the near-infrared imaging are retained at the same time, the appearance abnormalities and internal abnormalities of the food are captured, and more importantly, the correlation between the visible light image and the near-infrared imaging is explored by the method of correlation mapping. The three features are fused. Usually, when the appearance of the food is knocked and other quality problems occur, the internal rotting and deterioration also occur at the same time. Therefore, although the visible light image and the near-infrared imaging can represent different angle features in food quality detection, the two actually exist a mutual influence relationship. For example, the mold and mucus on the surface of the food are caused by the reproduction of the mold and bacteria inside the food. Similarly, the mucus, high moisture and rich nutrients on the surface of aquatic food provide an ideal breeding environment for microorganisms. From the surface to the inside, the microorganisms gradually invade the muscle tissue after destroying the skin barrier, causing the internal meat to soften and rot. Therefore, considering the correlation features of the two images, the accuracy of the detection is higher than that of the fusion of only two features.
[0090] Finally, the re-fusion feature is input into the MLP to calculate the quality probability P, and the food quality risk level is obtained based on the quality probability P, and the expression is as follows:
[0091]
[0092] wherein, is the food quality risk level, is a floor operation, is the quality probability.
[0093] The MLP model includes one hidden layer and one softmax layer, and the expression is:
[0094]
[0095] wherein, is the input feature map of the MLP, is the hidden layer weight coefficient, is the softmax layer weight coefficient, is the hidden layer bias, is the softmax layer bias.
[0096] When the MLP obtains a quality probability of 1, the food quality is the highest, and the risk level is zero level, and the warning level is the lowest; when the MLP obtains a quality probability of 0, the food quality is the lowest, and the risk level is ten level, and the warning level is the highest.
[0097] Further, the embodiment can set a risk level threshold, and when the detected risk level is lower than the threshold, quality warning is performed.
[0098] The workers can take corresponding processing measures on the detected food according to the risk warning level, discover the quality problems in the production process in time through detection, reduce the rate of defective products, and improve the production efficiency.
[0099] Embodiment two
[0100] In one or more embodiments, a food quality risk warning system is disclosed, specifically comprising:
[0101] A data acquisition module configured to acquire a visible light image and near-infrared imaging of the food;
[0102] An intelligent classification module configured to determine the category and segmentation mask of the food based on the visible light image of the food using an improved multi-scale feature fusion network;
[0103] A quality detection module configured to perform binary mask filtering on the visible light image and the near-infrared imaging to obtain image data based on the segmentation mask, call different food quality detection models according to the category of the food, and input the image data into the food quality detection model to determine the quality risk of the food;
[0104] The food quality detection model is an improved Transform network, the improved Transform network comprises an encoder, a feature mapping module and a multilayer perception machine; the encoder extracts double-mode double-scale fusion features, second visible light features and second near-infrared features of image data, the feature mapping module performs correlation analysis on the double-mode double-scale fusion features, the second visible light features and the second near-infrared features to obtain re-fusion features, and the multilayer perception machine outputs food quality risks based on the re-fusion features.
[0105] Embodiment three
[0106] The embodiment provides an electronic device, comprising a memory and a processor, and computer instructions stored in the memory and running on the processor, when the computer instructions are run by the processor, the steps of the food quality risk early warning method are completed.
[0107] Embodiment four
[0108] The embodiment provides a computer readable storage medium for storing computer instructions, when the computer instructions are executed by the processor, the steps of the food quality risk early warning method are completed.
[0109] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system) and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device that realizes the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks Figure 1 The functions specified in one flow or multiple flows and / or blocks
[0110] These computer program instructions can also be stored in a computer readable memory capable of guiding the computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer readable memory produce a product comprising instruction devices, which realize the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks Figure 1 The functions specified in one flow or multiple flows and / or blocks
[0111] These computer program instructions can also be loaded into a computer or other programmable data processing device to perform a series of operation steps on the computer or other programmable device to produce a computer implemented process, so that the instructions executed on the computer or other programmable device provide a process for realizing the functions specified in the flowcharts and / or block diagrams.Figure 1 one or more processes and / or functions described in one or more blocks. Figure 1 one or more processes and / or functions described in one or more blocks.
[0112] The above description of the various embodiments can be focused on one or more aspects of the embodiments. Those aspects not described in detail for one embodiment can be found in the description of the other embodiments.
[0113] The above descriptions are only preferred embodiments of the application, not intended to limit the application. Those of ordinary skill in the art can make various modifications and changes to the application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the application shall be included in the protection scope of the application.
Claims
1. A method for early warning of food quality risks, characterized in that, include: Acquire visible light and near-infrared images of food; Based on visible light images of food, an improved multi-scale feature fusion network is used to determine the food category and segmentation mask; Image data is obtained by performing binary mask filtering on visible light images and near-infrared imaging based on the segmentation mask. Different food quality detection models are called according to the food category. The image data is input into the food quality detection model to determine the quality risk of the food. The food quality detection model is an improved Transformer network, which includes an encoder, a feature mapping module, and a multilayer perceptron. The encoder includes two Restormer layers, two Transformer layers, and a multilayer perceptron connected in sequence. The first Restormer layer outputs shallow features, and the shallow features are processed by a Restormer, two Transformer layers, and a multilayer perceptron to obtain deep features. The encoder further includes a weighted connection layer; the deep features and the shallow features are input together into the weighted connection layer to obtain dual-modal dual-scale fusion features; the encoder extracts dual-modal dual-scale fusion features, second visible light features, and second near-infrared features from the image data; the feature mapping module performs correlation analysis on the dual-modal dual-scale fusion features, second visible light features, and second near-infrared features to obtain re-fusion features; the multilayer perceptron outputs food quality risk based on the re-fusion features.
2. The food quality risk early warning method as described in claim 1, characterized in that, The improved multi-scale feature fusion network structure includes a backbone network, a feature fusion network, and a detection network; The backbone network includes a first single convolutional module, a first multi-convolutional module, a second single convolutional module, a second multi-convolutional module, and a third single convolutional module connected in sequence. The feature fusion network includes a downsampling layer, a second connection layer, a third multi-convolutional module, a third connection layer, and a fifth convolutional layer connected in sequence. The input of the downsampling layer is the output feature of the backbone network. The output feature of the second multi-convolutional module is also input to the second connection layer, which concatenates the output feature of the second multi-convolutional module and the output feature of the downsampling layer. The output feature of the first multi-convolutional module is also input to the third connection layer, which concatenates the output feature of the first multi-convolutional module and the output feature of the third multi-convolutional module. The detection network includes a fourth single convolutional module, a fifth single convolutional module, a sixth convolutional layer, a seventh convolutional layer, a fourth connecting layer, a sixth single convolutional module, and an output layer. The output features of the feature fusion network are simultaneously output from the fourth single convolutional module and the fifth single convolutional module, and are then fed into the fourth connecting layer along with the output of the fifth single convolutional module. The fourth connecting layer is then connected to the sixth single convolutional module and the output layer.
3. The food quality risk early warning method as described in claim 1, characterized in that, The deep features and the shallow features are input together into a weighted connect layer to obtain a dual-modal, dual-scale fused feature, expressed as: in, It is a dual-modal, dual-scale fusion feature; It is the first visible light characteristic; This is the first near-infrared feature; This is the second visible light characteristic; This is the second near-infrared feature; , , , These are the weight coefficients for each feature.
4. The food quality risk early warning method as described in claim 3, characterized in that, The loss function of the encoder is: in, It is a two-modal, two-scale loss function; The correlation between the first visible light feature and the first near-infrared feature; The correlation between the second visible light feature and the second near-infrared feature; Parameters that are greater than zero; This is a structural similarity index.
5. The food quality risk early warning method as described in claim 1, characterized in that, The obtained dual-modal dual-scale fused features, second visible light features, and second near-infrared features are correlated and mapped through a feature mapping module to obtain re-fused features. Calculate the autocovariance matrix of each set of features. , , ; in, The autocovariance matrix of the dual-modal, dual-scale fused features. Let be the autocovariance matrix of the second visible light feature. The autocovariance matrix of the second near-infrared feature. For transpose; Calculate the cross-covariance matrix , , ; in, This is the cross-covariance matrix of the dual-modal, dual-scale fused features and the second visible light feature. Let be the cross-covariance matrix of the second visible light feature and the second near-infrared feature. The cross-covariance matrix of the second near-infrared feature and the dual-modal dual-scale fused feature; Solve for the mapping direction to find the projection direction that maximizes the correlation between the three sets of features: in, , , The projection direction that maximizes feature correlation; , , The correlation coefficient; The mapping features are: in, Visible light mapping characteristics, It is a near-infrared mapping feature. It is a dual-modal, dual-scale fusion mapping feature. , and Let be the projection matrix, and , , ; The mapped features are fused to obtain the refused features: in, This is the third-dimensional splicing function.
6. The food quality risk early warning method as described in claim 5, characterized in that, The re-fused features are input into a multilayer perceptron to calculate the quality probability P. Based on the quality probability, the food quality risk level is obtained, as shown in the following expression: in, Food quality risk level, This is for floor function.
7. A food quality risk early warning system, characterized in that, include: The data acquisition module is configured to acquire visible light images and near-infrared images of food. The intelligent classification module is configured to: determine the food category and segmentation mask based on the visible light image of the food using an improved multi-scale feature fusion network; The quality inspection module is configured to: perform binary mask filtering on visible light images and near-infrared imaging based on the segmentation mask to obtain image data; call different food quality inspection models according to the food category; and input the image data into the food quality inspection model to determine the quality risk of the food. The food quality detection model is an improved Transformer network, which includes an encoder, a feature mapping module, and a multilayer perceptron. The encoder includes two Restormer layers, two Transformer layers, and a multilayer perceptron connected in sequence. The first Restormer layer outputs shallow features, and the shallow features are processed by a Restormer, two Transformer layers, and a multilayer perceptron to obtain deep features. The encoder further includes a weighted connection layer; the deep features and the shallow features are input together into the weighted connection layer to obtain dual-modal dual-scale fusion features; the encoder extracts dual-modal dual-scale fusion features, second visible light features, and second near-infrared features from the image data; the feature mapping module performs correlation analysis on the dual-modal dual-scale fusion features, second visible light features, and second near-infrared features to obtain re-fusion features; the multilayer perceptron outputs food quality risk based on the re-fusion features.
8. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the food quality risk early warning method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, complete the food quality risk early warning method according to any one of claims 1-6.
Citation Information
Patent Citations
Food material quality detection system based on nondestructive detection and detection method thereof
CN116012837A
Chicken quality detection method based on deep learning multi-source spectrum fusion and sorting equipment
CN116202978A