Food quality risk early warning method and system
By combining visible light images and near-infrared imaging, using multi-scale feature fusion network and Transformer network, the existing food quality detection problems are solved, and efficient and accurate early warning of food quality risk is achieved.
Patent Information
- Application Number
- CN202510366906.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-26
AI Technical Summary
The existing food quality testing methods have long inspection time, complex operation, high cost, low accuracy and unstable results. The single feature detection effect is poor, and it is impossible to achieve unified inspection of multiple foods, ignoring the correlation between appearance and internal state.
Using visible light images and near-infrared imaging, through an improved multi-scale feature fusion network and Transformer network, the categories and segmentation masks of food are extracted, and multiple image features are fused to provide food quality risk warnings.
It improves the accuracy and stability of food testing, can simultaneously detect the appearance and internal quality of food, reduce defective rates, and reduce production costs, and is suitable for unified testing of multiple foods.
Smart Images

Figure CN120236277A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of food detection, and particularly relates to a food quality risk warning method and system. Background Art
[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] With the complication of the food industry chain, food quality problems have become the focus of social attention. Food may be affected by various factors during production, processing, transportation and sales, resulting in frequent quality problems. For example, problems such as microbial contamination, pesticide residues, and abnormal states in food will seriously affect the health of consumers. Therefore, during the process of food production by enterprises, attention should be paid to the detection of food quality.
[0004] Traditional food quality detection methods mainly rely on technologies such as chemical analysis and microbial culture. Although these methods can provide relatively accurate results, they have problems such as long detection time, complex operation, and high cost.
[0005] There are also quality detection methods such as image processing and biosensor technology in the prior art, but the detection effect of a single feature has a large error. For example, the image recognition technology using a high-definition camera can only observe the spoilage problem on the surface of food, and cannot accurately detect problems such as its residual pesticides and excessive internal microorganisms. The detection method using spectroscopy technology has certain requirements for the transparency, color and moisture content of the sample. For samples with poor transparency, dark color or containing a large amount of moisture, the detection effect may not be good; the non-uniformity of food may affect the propagation of light, resulting in instability of the detection results. Therefore, the quality detection of a single feature has problems of low accuracy and instability. Secondly, a single feature not only loses some features, but also ignores the state correlation between the appearance and the inside of food. Features such as bumps and microbial attachment on the outer surface interact with abnormal internal spoilage and the like.
[0006] In addition, from the perspective of the detection object, most of the existing detection models are for the detection of a specific fixed food, and can only detect a single food product on a certain production line, and cannot achieve unified detection of foods sampled from multiple production lines. Summary of the Invention
[0007] To overcome the deficiencies of the above-mentioned existing technologies, the present invention proposes a food quality risk early warning method and system, which covers the comprehensive features of food by fusing multiple images to improve the detection accuracy and stability; proposes multiple fusions of multi-scale features, retains shallow edge features while extracting deep detail features, balances global features and local features, and reduces feature loss; in addition, proposes a similarity analysis fusion feature between multiple image features, increases the correlation analysis of food appearance anomalies and internal anomalies, improves the representation ability of the fusion feature, and further improves the accuracy of quality risk.
[0008] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions: In a first aspect, a food quality risk early warning method is disclosed, including: Obtaining a visible light image and a near-infrared image of the food; Based on the visible light image of the food, using an improved multi-scale feature fusion network to determine the category and segmentation mask of the food; Performing binary mask filtering on the visible light image and the near-infrared image based on the segmentation mask to obtain image data, calling different food quality detection models according to the category of the food, and inputting the image data into the food quality detection model to determine the food quality risk; Wherein, the food quality detection model is an improved Transformer network, and the improved Transformer network includes an encoder, a feature mapping module and a multi-layer perceptron; the encoder extracts the bimodal and two-scale fusion feature, the second visible light feature and the second near-infrared feature of the image data, the feature mapping module performs correlation analysis on the bimodal and two-scale fusion feature, the second visible light feature and the second near-infrared feature to obtain a re-fusion feature, and the multi-layer perceptron outputs the food quality risk based on the re-fusion feature.
[0009] In a second aspect, a food quality risk early warning system is disclosed, including: A data acquisition module configured to: obtain a visible light image and a near-infrared image of the food; An intelligent classification module configured to: based on the visible light image of the food, use an improved multi-scale feature fusion network to determine the category and segmentation mask of the food; A quality detection module configured to: perform binary mask filtering on the visible light image and the near-infrared image based on the segmentation mask to obtain image data, call different food quality detection models according to the category of the food, and input the image data into the food quality detection model to determine the food quality risk; Among them, the food quality detection model is an improved Transformer network, and the improved Transformer network includes an encoder, a feature mapping module, and a multi-layer perceptron; the encoder extracts the dual-modal and dual-scale fusion features, the second visible light feature, and the second near-infrared feature of the image data, the feature mapping module performs correlation analysis on the dual-modal and dual-scale fusion features, the second visible light feature, and the second near-infrared feature to obtain the re-fusion feature, and the multi-layer perceptron outputs the food quality risk based on the re-fusion feature.
[0010] In a third aspect, an electronic device is disclosed, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps of the above food quality risk warning method are completed.
[0011] In a fourth aspect, a computer-readable storage medium is disclosed, which is used to store computer instructions. When the computer instructions are executed by the processor, the steps of the above food quality risk warning method are completed.
[0012] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention provides a food quality risk warning method, which fuses visible light images and near-infrared imaging. By simultaneously detecting the appearance and internal quality, the detection accuracy is improved, and enterprises can more efficiently control the production process and reduce the defective rate.
[0013] In the food classification model, when extracting features for each type of image, multi-scale feature fusion is adopted to retain the global features and detailed features, avoiding the loss of basic features and edge features. Therefore, foods with significantly different appearances can be more accurately and quickly identified, avoiding getting stuck in the misunderstanding of minor details.
[0014] When the food classification model obtains the food category, it can also output a segmentation mask, thereby providing a data processing method for the detection model, processing the input image of the detection model into an image data containing only the food body, and reducing the influence of the background on the quality detection.
[0015] In the quality detection model, a multi-modal image fusion and a fusion method considering the correlation features of the two images are proposed. The features used for detection include the image features of a single modality and the fusion features of the two modalities. The visible light image can effectively represent the appearance state of the food, and the near-infrared imaging can effectively represent the internal state of the food. The fusion features use the method of correlation mapping to extract the matrix of the correlation between the two images, thereby extracting the relevant features of the internal and appearance. Using the correlation feature fusion to detect the food quality can make full use of the complementary information of different modalities, improve the classification performance, enhance the data efficiency and robustness, and help to reveal the relevant connections of the quality risks between multi-modal data, improving the accuracy of quality detection.
[0016] This method can be used to detect quality risks at any stage of food production. By timely detection of quality problems during production, it can prevent unqualified products from entering subsequent processes, reducing waste of raw materials and production costs.
[0017] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings, which form a part of this specification, are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not unduly limit the present invention.
[0019] Figure 1 It is a flowchart of the food quality risk warning method described in Embodiment 1 of the present invention.
[0020] Figure 2 It is a structural diagram of the intelligent classification model described in Embodiment 1 of the present invention.
[0021] Figure 3 It is a structural diagram of the quality detection model described in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0023] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.
[0024] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0025] Embodiment 1 In one or more embodiments, a food quality risk warning method is disclosed, including the following steps: Step S1: Obtain visible light images and near-infrared imaging of the food.
[0026] The visible light images can be obtained by high-definition shooting devices such as CMOS / CCD sensors, digital cameras, surveillance cameras, industrial cameras, etc., which are not limited herein; The near-infrared imaging can be obtained by devices such as short-wave infrared detectors and short-wave infrared cameras, which are not limited herein; The acquisition devices for visible light images and near-infrared imaging are installed at the same location and placed as close together as possible. It is also possible to use a device that can acquire visible light images and near-infrared imaging simultaneously. The specific device model is not limited here. It is only necessary that the shooting distances and shooting angles of the visible light images and near-infrared imaging are the same.
[0027] Visible light images can detect the appearance defects of food, such as scratches, dents, color changes, and surface molds, foreign objects, etc. Near-infrared images can detect internal components, structures, chemical components, and potential harmful substances in food, such as pesticide residues, heavy metals, etc. Combining the two image technologies can detect defects or contaminants that cannot be found by a single technology. For example, near-infrared images can detect the internal moisture content, bruises, or impurities through packaging.
[0028] Step S2: Based on the visible light image of the food, use an intelligent classification model to classify the food into specific categories. In this embodiment, it is mainly divided into three categories, including: poultry eggs, chickens and ducks, and aquatic products. When the food types on the actual food production line change, other categories can be divided according to actual needs.
[0029] The intelligent classification model is specifically an improved multi-scale feature fusion network. The intelligent classification model is a network model trained with a training set constructed from historical collected data. Collect historical visible light image data of poultry eggs, chickens and ducks, and aquatic products, and calibrate the category labels to construct a training data set and train the model.
[0030] The structure of the improved multi-scale feature fusion network includes a backbone network, a feature fusion network, and a detection network connected in sequence, as Figure 2 shown; The backbone network includes several single convolution modules and several multi-convolution modules, specifically the first single convolution module, the first multi-convolution module, the second single convolution module, the second multi-convolution module, and the third single convolution module connected in sequence.
[0031] The single convolution module includes a convolutional layer conv + a batch normalization layer BN + an activation function layer SiLU, and the expression is:
[0032] Among them, is the output feature of the single convolution module, is the activation function, is the batch normalization, is the convolutional layer, is the input food feature map of the single convolution module.
[0033] The single convolution module is used to extract features from food images. As the single-layer convolution modules are stacked, the network captures the local information of the food images, refines the food bounding boxes, and improves the expression of detailed features.
[0034] The multi-convolution module includes a first convolution layer, a second convolution layer, a third convolution layer, a first connection layer, and a fourth convolution layer. The input feature map is input into the first convolution layer to obtain a first feature map. The first feature map is input into the second convolution layer to obtain a second feature map. The second feature map is input into the third convolution layer to obtain a third feature map. The first feature map, the second feature map, and the third feature map are input into the first connection layer for splicing to obtain a fourth feature map. The fourth feature map is input into the fourth convolution layer to obtain the output of the multi-convolution module.
[0035]
[0036] Among them, is the output feature of the multi-convolution module, is the input food feature map of the multi-convolution module.
[0037] The multi-convolution module sequentially extracts detailed features through three convolution layers, and then fuses the feature maps of three different depths so that the multi-convolution module can recognize the feature information from various scales of the food image, avoiding the loss of shallow features, while increasing the mutual semantic relationship between shallow features and deep features, and improving the feature representation ability.
[0038] The backbone network is followed by a feature fusion network, which includes a downsampling layer, a second connection layer, a third multi-convolution module, a third connection layer, and a fifth convolution layer connected in sequence. The input of the downsampling layer is the output feature of the backbone network. The output feature of the second multi-convolution module is also input to the second connection layer, and the second connection layer concatenates the output feature of the second multi-convolution module and the output feature of the downsampling layer. The output feature of the first multi-convolution module is also input to the third connection layer, and the third connection layer concatenates the output feature of the first multi-convolution module and the output feature of the third multi-convolution module. In this embodiment, according to the actual production needs, the targets of the intelligent classification model are set as three types: poultry eggs, chickens and ducks, and aquatic products. There are significant differences in the shapes of poultry eggs, chickens and ducks, and aquatic products. Therefore, the output feature maps of the first multi-convolution module and the second multi-convolution module are respectively concatenated with the deep features through skip connections, so as to retain the most basic global features of each type of food, avoid the adverse effects of overly local features on classification, and the multi-level multi-scale feature fusion not only retains the shape differences of the three types of images globally, but also the gradually deepening convolution modules can extract detailed information, provide local features for food border detection, make the food border more refined, improve the food positioning and boundary accuracy. The feature fusion network fuses the feature maps of different scales in the backbone network, improves the accuracy of food classification and the accuracy of the mask at the same time, and thus acts on the subsequent quality detection model, eliminates the irrelevant background for the quality detection model, and avoids the influence of the conveyor belt or the equipment for receiving food on the food quality detection.
[0039] Detection network: Adopt a coupled detection method to obtain the predicted category and the target bounding box at the same time. Specifically, the detection network includes a fourth single-convolution module, a fifth single-convolution module, a sixth convolution layer, a seventh convolution layer, a fourth connection layer, a sixth single-convolution module, and an output layer. The output feature of the feature fusion network is output to the fourth single-convolution module and the fifth single-convolution module at the same time. One branch is the fourth single-convolution module, the sixth convolution layer, and the seventh convolution layer connected in sequence, and the other branch is the fifth single-convolution module. The two branches are jointly input to the fourth connection layer, and then output after passing through a sixth single-convolution module. Finally, the food category and the segmentation mask are output by the output layer.
[0040] The output of the food category enables this model to be used in a single production line for various foods, without the need for manual distinction of food categories and separate setting of detection software, and the same model can be uniformly arranged; it can also be applied to the sampling detection of various foods. Different types of foods from different production lines are randomly sampled and mixed into one detection line, reducing the layout of detection lines. This model can automatically identify different types of foods on the detection line, achieve fully automated detection, and reduce costs.
[0041] The segmentation mask output obtains the range of the food ontology image, blocks background noise, assists the quality detection model to extract features for the food, reduces the computational amount, and improves the efficiency of the overall food quality risk warning.
[0042] S3: According to the category label obtained in step S2, call the quality detection model corresponding to the category to perform image detection on the target food; Preprocess the visible light image and near-infrared imaging based on the segmentation mask to obtain image data, call different food quality detection models according to the category of the food, and determine the quality risk of the food based on the image data.
[0043] S3-1: First, perform interpolation processing on the segmentation mask to restore the mask to the original image size, then overlay the mask on the visible light image and near-infrared imaging, and use binary mask filtering to retain the image with the mask value of 1 and set the image with the mask value of 0 to 0, obtaining the visible light image data and near-infrared image data that only contain the food ontology. The visible light image data and near-infrared image data that only contain the food ontology together constitute the image data and are simultaneously input into the food quality detection model. Thus, when detecting the food quality, the influence of environmental factors on the detection effect is eliminated, and the attention of the detection model is completely focused on the food.
[0044] S3-2: Input the processed visible light image data and near-infrared image data into the quality detection model to obtain the quality score.
[0045] The quality detection model includes a quality detection model for poultry and eggs pre-trained with historical visible light image data and historical near-infrared image data of poultry and eggs, a quality detection model for chickens and ducks pre-trained with historical visible light image data and historical near-infrared image data of chickens and ducks, and a quality detection model for aquatic products pre-trained with historical visible light image data and historical near-infrared image data of aquatic products. When the classification result in step S2 is poultry and eggs, the quality detection model for poultry and eggs is called; when the classification result in step S2 is chickens and ducks, the quality detection model for chickens and ducks is called; when the classification result in step S2 is aquatic products, the quality detection model for aquatic products is called; thus, it is avoided that the appearance differences of the three different food types are relatively large and features cannot be extracted simultaneously by training, thereby improving the accuracy of quality detection. The quality detection models of the three food types have the same structure and are only models with different parameters trained with different training data.
[0046] Specifically, as Figure 3 shown, the quality detection model uses an improved transformer network, including an encoder, a feature mapping module, and a multi-layer perceptron. By designing a cross-modal fusion Transformer network, one image can receive information from another image, thereby achieving more effective information fusion.
[0047] First, the encoder includes two Restormer layers, two Transformer layers, a multi-layer perceptron, and a weighted connection layer connected in sequence. On the one hand, the encoder extracts shallow and deep features of a single food image, and on the other hand, it fuses visible light images and near-infrared imaging to obtain correlation features. The multi-layer perceptron MLP model serves as a food feature extractor.
[0048] First, the visible light image data I im and the near-infrared image data I in are respectively input into the network to extract features.
[0049] The features output by the first Restormer layer are the first visible light feature and the first near-infrared feature , where is the shallow feature map of the food visible light image data, is the shallow feature map of the near-infrared imaging data, and C, H, and W represent the number of channels, height, and width of the feature map. Since the shallow features and contain the basic details of the image, more edge details of the image are extracted.
[0050] Then, the first visible light feature and the first near-infrared feature are input into two Transformer layers, and the features output by the second Restormer layer are used to extract intermediate features and , capturing long-range feature information for extracting the semantic information of deep features.
[0051] The features and further obtain the deep features of each image through the MLP, including the second visible light feature and the second near-infrared feature . At the same time, the deep features and are jointly input into the weighted connection layer together with the shallow features and to obtain the bimodal and dual-scale fusion feature, and the expression is:
[0052] where is the bimodal and dual-scale fusion feature; , , , are the weight coefficients of each feature.
[0053] The encoder adopts a loss function for bimodal and dual-scale feature fusion:
[0054] Among them, is the bimodal and two-scale loss function; is the correlation between the first visible light feature and the first near-infrared feature; is the correlation between the second visible light feature and the second near-infrared feature; is a parameter greater than zero, that is, a positive number infinitely close to 0; is the structural similarity index SSIM.
[0055] Through the combined action of the above loss function and the weighted connection layer, the fusion coefficients of the shallow features and deep features of the two images are trained to be optimal, which not only ensures that the images can mine semantic information and retain more detailed information, but also ensures that the basic edge information is not lost, balances the global and local features, and enables the visible light images and near-infrared images to play their maximum roles.
[0056] The obtained bimodal and two-scale fusion features , the second visible light feature , the second near-infrared feature , are used to obtain the re-fusion features through correlation mapping: Calculate the autocovariance matrix of each group of features , , ; Among them, is the autocovariance matrix of the bimodal and two-scale fusion features, is the autocovariance matrix of the second visible light feature, is the autocovariance matrix of the second near-infrared feature, is the transpose.
[0057] Calculate the cross-covariance matrix , , ; Among them, is the cross-covariance matrix between the bimodal and two-scale fusion features and the second visible light feature, is the cross-covariance matrix between the second visible light feature and the second near-infrared feature, is the cross-covariance matrix between the second near-infrared feature and the bimodal and two-scale fusion features.
[0058] Solve the mapping direction to find the projection direction that maximizes the correlation of the three groups of features:
[0059] Among them, , , The projection direction for maximizing feature correlation; , , is the correlation coefficient.
[0060] Then the mapped features are:
[0061] where, is the visible light mapped feature, is the near-infrared mapped feature, is the dual-modal and dual-scale fusion mapped feature, , and are projection matrices, and , , .
[0062] The mapped features are concatenated according to the third dimension to obtain the re-fused feature :
[0063] where, is the third-dimension concatenation function.
[0064] In this way, the features of the visible light image and the near-infrared image are retained simultaneously, capturing the appearance abnormalities and internal abnormalities of the food. More importantly, by fusing the three features through the correlation mapping method, the correlation between the visible light image and the near-infrared image can be explored. Usually, when there are quality problems such as appearance bumps on the food, internal decay and deterioration also occur. Therefore, although the visible light image and the near-infrared image can respectively represent the features from different angles in food quality detection, there is actually a mutual influence relationship between the two. For example, the appearance of mildew spots, mucus, etc. on the food surface is due to the reproduction of internal molds or bacteria. Similarly, the mucus, high moisture, and rich nutrients on the surface of aquatic foods provide an ideal breeding environment for microorganisms. From the surface to the inside, after the microorganisms break through the epidermal barrier, they gradually invade the muscle tissue, resulting in softening and spoilage of the internal meat. Therefore, considering the correlation features of the two images is more accurate than only fusing the two features for detection.
[0065] Finally, the re-fused feature is input into the MLP to calculate the quality probability to obtain the quality probability P, and the food quality risk level is obtained based on the quality probability P. The expression is as follows:
[0066] where, is the food quality risk level, is the floor operation, is the quality probability.
[0067] The MLP model includes a hidden layer and a softmax layer, and its expression is:
[0068] Among them, the input feature map of the MLP, is the weight coefficient of the hidden layer, is the weight coefficient of the softmax layer, is the bias of the hidden layer, is the bias of the softmax layer.
[0069] When the MLP obtains a quality probability of 1, the food quality is the highest, and at this time the risk level is zero level and the warning level is the lowest; when the MLP obtains a quality probability of 0, the food quality is the lowest, and at this time the risk level is ten levels and the warning level is the highest.
[0070] Furthermore, in this embodiment, a risk level threshold can be set, and when the detected risk level is lower than this threshold, a quality alarm is issued.
[0071] The staff can take corresponding treatment measures for the detected food according to the level of risk warning, timely discover quality problems in the production process through detection, reduce the defective rate, and improve production efficiency.
[0072] Embodiment 2 In one or more embodiments, a food quality risk warning system is disclosed, which specifically includes: A data acquisition module, which is configured to: acquire visible light images and near-infrared imaging of food; An intelligent classification module, which is configured to: based on the visible light image of the food, determine the category and segmentation mask of the food by using an improved multi-scale feature fusion network; A quality detection module, which is configured to: perform binary mask filtering on the visible light image and near-infrared imaging based on the segmentation mask to obtain image data, call different food quality detection models according to the category of the food, and input the image data into the food quality detection model to determine the food quality risk; Among them, the food quality detection model is an improved Transformer network, and the improved Transformer network includes an encoder, a feature mapping module and a multi-layer perceptron; the encoder extracts the dual-modal and dual-scale fusion features, the second visible light feature, and the second near-infrared feature of the image data, the feature mapping module performs correlation analysis on the dual-modal and dual-scale fusion features, the second visible light feature, and the second near-infrared feature to obtain a re-fusion feature, and the multi-layer perceptron outputs the food quality risk based on the re-fusion feature.
[0073] Embodiment 3 This embodiment provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps of the above food quality risk warning method are completed.
[0074] Embodiment 4 This embodiment provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the above food quality risk warning method are completed.
[0075] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0076] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0077] These computer program instructions can also be loaded onto a computer or other programmable data processing device, and a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0078] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0079] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A food quality risk early warning method, characterized in that: include: Obtain visible light images and near-infrared imaging of food; Based on the visible light images of food, an improved multi-scale feature fusion network is used to determine the category and segmentation mask of the food; Based on the segmentation mask, the visible light image and the near infrared image are subjected to binary mask filtering to obtain image data, different food quality detection models are called according to the category of food, and the image data is input into the food quality detection model to determine the quality risk of the food; Among them, the food quality detection model is an improved Transformer network, which includes an encoder, a feature mapping module and a multi-layer perceptron; the encoder extracts the bimodal bi-scale fusion features, the second visible light features, and the second near-infrared features of the image data, the feature mapping module performs correlation analysis on the bimodal bi-scale fusion features, the second visible light features, and the second near-infrared features to obtain re-fusion features, and the multi-layer perceptron outputs food quality risks based on the re-fusion features.
2. A food quality risk early warning method according to claim 1, characterized in that: The structure of the improved multi-scale feature fusion network includes a backbone network, a feature fusion network and a detection network; The backbone network includes a first single convolution module, a first multi-convolution module, a second single convolution module, a second multi-convolution module and a third single convolution module which are connected in sequence; The feature fusion network includes a downsampling layer, a second connection layer, a third multi-convolution module, a third connection layer and a fifth convolution layer connected in sequence; the input of the downsampling layer is the output feature of the backbone network; the output feature of the second multi-convolution module is also input into the second connection layer, the second connection layer splices the output feature of the second multi-convolution module with the output feature of the downsampling layer, the output feature of the first multi-convolution module is also input into the third connection layer, the third connection layer splices the output feature of the first multi-convolution module with the output feature of the third multi-convolution module; The detection network includes a fourth single convolution module, a fifth single convolution module, a sixth convolution layer, a seventh convolution layer, a fourth connection layer, a sixth single convolution module and an output layer; the output features of the feature fusion network are simultaneously output to the fourth single convolution module and the fifth single convolution module, and are successively input to the fourth connection layer through the outputs of the fourth single convolution module, the sixth convolution layer, the seventh convolution layer and the output of the fifth single convolution module, and the fourth connection layer is connected to the sixth single convolution module and the output layer.
3. A food quality risk early warning method according to claim 1, characterized in that: The encoder includes two Restormer layers, two Transformer layers and a multi-layer perceptron connected in sequence; The first Restormer layer outputs shallow features, and the shallow features pass through a Restormer, two Transformer layers and a multi-layer perceptron to obtain deep features.
4. A food quality risk early warning method as claimed in claim 3, characterized in that: The encoder also includes a weighted connection layer; The deep layer features and the shallow layer features are input into the weighted connection layer to obtain the dual-modal dual-scale fusion features, which are expressed as: in, It is a dual-modal dual-scale fusion feature; is the first visible light feature; It is the first near-infrared feature; is the second visible light feature; It is the second near-infrared feature; , , , is the weight coefficient of each feature.
5. A food quality risk early warning method as claimed in claim 4, characterized in that: The loss function of the encoder is: in, is a bimodal and biscale loss function; is the correlation between the first visible light feature and the first near infrared feature; is the correlation between the second visible light feature and the second near-infrared feature; is a parameter greater than zero; is the structural similarity index.
6. A food quality risk early warning method according to claim 1, characterized in that: The obtained dual-modal dual-scale fusion features, the second visible light features, and the second near-infrared features are correlated and mapped through the feature mapping module to obtain the re-fusion features: Calculate the autocovariance matrix for each set of features , , ; in, is the autocovariance matrix of the bimodal and biscale fusion features, is the autocovariance matrix of the second visible light feature, is the autocovariance matrix of the second near-infrared feature, is transposed; Compute the cross-covariance matrix , , ; in, is the cross-covariance matrix of the dual-modal dual-scale fusion feature and the second visible light feature, is the cross-covariance matrix of the second visible light feature and the second near-infrared feature, is the cross-covariance matrix of the second near-infrared feature and the dual-modal dual-scale fusion feature; Solve the mapping direction and find the projection direction that maximizes the correlation between the three sets of features: in, , , The projection direction that maximizes feature correlation; , , is the correlation coefficient; The mapping characteristics are: in, is the visible light mapping feature, is the near infrared mapping feature, is the dual-modal dual-scale fusion mapping feature, , and is the projection matrix, and , , ; Fuse the mapped features to obtain the re-fused features: in, is the third-dimensional splicing function.
7. A food quality risk early warning method as claimed in claim 6, characterized in that: The re-integrated features are input into a multi-layer perceptron to calculate the quality probability to obtain the quality probability P, and the food quality risk level is obtained based on the quality probability. The expression is as follows: in, is the food quality risk level, This is a floor operation.
8. A food quality risk early warning system, characterized in that: include: A data acquisition module is configured to: acquire visible light images and near infrared images of food; An intelligent classification module is configured to: determine the category and segmentation mask of the food based on the visible light image of the food using an improved multi-scale feature fusion network; A quality detection module is configured to: perform binary mask filtering on the visible light image and the near infrared image based on the segmentation mask to obtain image data, call different food quality detection models according to the category of food, and input the image data into the food quality detection model to determine the quality risk of the food; Among them, the food quality detection model is an improved Transformer network, which includes an encoder, a feature mapping module and a multi-layer perceptron; the encoder extracts the bimodal bi-scale fusion features, the second visible light features, and the second near-infrared features of the image data, the feature mapping module performs correlation analysis on the bimodal bi-scale fusion features, the second visible light features, and the second near-infrared features to obtain re-fusion features, and the multi-layer perceptron outputs food quality risks based on the re-fusion features.
9. An electronic device, characterized in that: The method comprises a memory and a processor, and computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the food quality risk early warning method according to any one of claims 1 to 7 is completed.
10. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, complete the food quality risk early warning method described in any one of claims 1-7.
Citation Information
Patent Citations
Food material quality detection system based on nondestructive detection and detection method thereof
CN116012837A
Chicken quality detection method based on deep learning multi-source spectrum fusion and sorting equipment
CN116202978A
Infrared and visible light image fusion method combining Transform and CNN double encoders
CN117314808A
Beef carcass quality rating method based on image analysis in combination with machine learning
CN118537664A
Evaluation of food quality based on machine learning
CN118742928A
Cited By
Waterfowl phenotype recognition device
CN120496131A
Transform-based wet tissue surface defect real-time detection method
CN120976205A