A food detection method and system based on computer vision

By utilizing multimodal imaging technology and deep learning algorithms, comprehensive automation of food testing has been achieved, solving the problems of insufficient detection accuracy and adaptability in existing technologies and improving the accuracy and adaptability of food testing.

CN122368048APending Publication Date: 2026-07-10蒙阴县检验检测中心

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
蒙阴县检验检测中心
Filing Date
2026-06-02
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing food testing technologies suffer from problems such as high labor intensity, inconsistent testing standards, poor environmental adaptability, limited testing accuracy, insufficient model generalization ability, inability to perform full-surface testing, and insufficient foreign object identification capabilities, making it difficult to meet the diverse and rapidly iterating needs of the food industry.

Method used

A multimodal imaging device is used to simultaneously acquire visible light, near-infrared and hyperspectral images. Through frame synchronization and spatial registration processing, multi-scale and multimodal features are extracted. Combined with adaptive fusion and deep learning algorithms, integrated detection of food defects, foreign objects and freshness is achieved.

Benefits of technology

It enables comprehensive detection of food appearance defects, foreign objects, and internal quality, improving detection accuracy and automation, adapting to different food types, reducing false detection and false negative rates, and enhancing the model's generalization ability and update efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122368048A_ABST
    Figure CN122368048A_ABST
Patent Text Reader

Abstract

This invention discloses a computer vision-based food inspection method and system, belonging to the field of computer vision-based food inspection technology. The invention simultaneously acquires visible light, near-infrared, and hyperspectral images of the food to be inspected, completing frame synchronization and spatial registration. After preprocessing, the region of interest (ROI) of the food body is extracted and normalized. Multi-scale and multi-modal features are extracted and adaptively fused. The system sequentially completes food defect identification, foreign object detection, and quantitative assessment of freshness. Finally, the food is graded according to preset standards, and an inspection report is generated and stored. This invention achieves integrated detection of food appearance defects, foreign objects, and internal quality; adaptive fusion of multi-modal features improves detection accuracy; full-surface acquisition eliminates blind spots; and edge-cloud collaboration continuously optimizes the model, providing technical support for food production quality control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision-based food inspection technology, and more particularly to a computer vision-based food inspection method and system. Background Technology

[0002] Food quality and safety are core aspects of food production and distribution, directly impacting consumer health. With the large-scale development of the food industry, traditional manual visual inspection methods suffer from high labor intensity, inconsistent testing standards, and susceptibility to human condition, making them unsuitable for the testing needs of mass production lines. Early machine vision inspection methods based on single-modal visible light could only identify obvious surface defects, failing to detect internal compositional changes and quality deterioration. Their ability to identify latent defects such as early-stage mold and slight discoloration was limited, and they struggled to distinguish between natural surface textures and genuine defects, leading to frequent false positives and false negatives.

[0003] Existing multimodal food detection technologies mostly employ fixed-weight feature fusion methods, failing to dynamically adjust the contribution ratio of each modality based on the characteristic differences of different food types and the needs of different tasks such as defect detection and freshness assessment. This results in the ineffective utilization of feature information from some modalities, limiting overall detection accuracy. Freshness assessment often relies on single color or texture features, failing to fully incorporate hyperspectral information reflecting changes in the internal composition of food. The assessment results have low correlation with actual physicochemical indicators and cannot effectively predict remaining shelf life. Regarding foreign object detection, existing methods are insufficient in identifying low-contrast micro-foreign objects such as hair and small plastic fragments, making it difficult to meet the stringent foreign object control requirements of the food industry.

[0004] In real-world production environments, factors such as fluctuating lighting, food surface reflections, and complex background interference can significantly impact image acquisition quality. Existing detection systems exhibit poor environmental adaptability, requiring frequent parameter adjustments to maintain stable detection results. Most detection systems can only acquire images of a single side of the food, failing to achieve comprehensive surface inspection and easily missing defects on the sides and bottom. Furthermore, existing detection models lack generalization ability; adding new food types necessitates extensive retraining with numerous new samples, resulting in long model update cycles and hindering their adaptation to the diverse and rapidly evolving nature of the food industry. Summary of the Invention

[0005] This invention proposes a food detection method and system based on computer vision to solve the problems mentioned in the prior art.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: a food detection method based on computer vision, comprising the following steps: Visible light color images, near-infrared grayscale images, and hyperspectral cube images of the food to be tested are acquired simultaneously using a multimodal imaging device, and the acquired raw images are processed for frame synchronization and spatial registration. The registered multimodal images are preprocessed to remove image noise and illumination interference, extract the region of interest of the food subject and normalize its size; Color and shape features are extracted from visible light images, texture and edge features are extracted from near-infrared images, and spectral reflectance features are extracted from hyperspectral images to construct a multi-scale, multi-modal feature set. Adaptive fusion of multi-scale and multi-modal features is performed to generate a fused feature vector. The fused feature vector is then input into a pre-trained food defect classification model to identify the types and locations of damage, mold, and discoloration defects on the food surface. The fused feature vector is input into the anomaly detection model to detect foreign objects such as metal, plastic, and hair mixed in with food, and a foreign object detection mask is generated. A food freshness assessment model was established based on hyperspectral and color features to quantitatively calculate the freshness level of the food to be tested. Based on the results of defect detection, foreign object detection, and freshness grade, the food is graded according to preset standards, and a test report is generated and the test data is stored.

[0007] Furthermore, it also includes a multimodal feature adaptive fusion step, which determines the fusion weights by calculating the discriminative power and stability of different modal features. The calculation formula is as follows: ; in, For the first The fusion weights of each modal feature For the first Inter-class discriminative power of modal features For the first Intra-class stability of modal features For modal indexing, before fusion, sub-pixel-level spatial alignment of different modal features is completed based on feature points on the food surface. During the fusion process, the weight allocation strategy is dynamically adjusted according to the priority of the detection task. At the same time, a modal confidence verification mechanism is introduced. When the feature variance of a certain modality exceeds the preset threshold, its weight is automatically reduced and the complementary weights of the other modalities are increased.

[0008] Furthermore, it also includes a comprehensive quantitative assessment step for food freshness, which calculates a comprehensive freshness score by integrating changes in multi-dimensional characteristics. The calculation formula is as follows: ; in, The overall score for food freshness The weighting coefficient for color feature changes. The Euclidean distance is the color feature of the food to be tested compared with the color feature of a fresh reference sample. The weighting coefficients for texture feature changes. Let be the cosine distance between the texture features of the food to be detected and the texture features of the fresh reference sample. The weighting coefficients represent the changes in hyperspectral characteristics. This is the average of the absolute values ​​of the differences between the reflectance of the food sample in the characteristic band and the reflectance of the corresponding band of the fresh reference sample.

[0009] Furthermore, the multimodal image preprocessing steps include: using a bilateral filtering algorithm to remove Gaussian noise and salt-and-pepper noise from visible light and near-infrared images; using wavelet transform denoising to preserve spectral details for hyperspectral images; using an adaptive histogram equalization algorithm to compensate for image brightness differences; introducing a polarization correction step for food surfaces with severe reflectivity to eliminate highlight areas caused by specular reflection; using a background extraction method combined with semantic segmentation to segment the food subject and background, extracting the minimum bounding rectangle of the food subject as the region of interest; uniformly scaling all regions of interest to a fixed size; using morphological opening operations to remove fine noise and edge burrs in the image; and using a grayscale normalization algorithm to map image pixel values ​​to the 0 to 1 range.

[0010] Furthermore, the multi-scale multimodal feature extraction steps include: constructing a three-layer feature pyramid to extract features at low, medium, and high resolution scales; extracting the mean, variance, skewness, and kurtosis statistical features of each channel in the RGB, HSV, and Lab color spaces from the visible light image; extracting the perimeter, area, roundness, aspect ratio, and concavity / convexity shape features of the food outline; extracting local binary patterns, gray-level co-occurrence matrix texture features, and orientation gradient histogram edge features at different scales from the near-infrared image; using a genetic algorithm to select the 20 feature bands most correlated with changes in food composition from the hyperspectral image; extracting the average reflectance, reflectance gradient, and spectral absorption depth features of each feature band; and generating a multi-scale multimodal feature set with unified dimensions.

[0011] Furthermore, the food defect and foreign object detection steps include: constructing an improved convolutional neural network combining channel attention and spatial attention mechanisms as a defect classification model; introducing a multi-scale feature fusion module into the network to detect defects and foreign objects; training the model using a food image dataset labeled with defect type, location, and severity; inputting the fused feature vector into the trained defect classification model to output defect type, confidence level, and bounding box coordinates; constructing an anomaly detection model using an isolated forest algorithm combined with template matching; establishing a multimodal feature template library for normal food; comparing the fused feature vector of the food to be detected with the template library to identify foreign object regions that exceed the normal feature distribution and generating a detection mask.

[0012] Furthermore, the food freshness assessment steps include: establishing dedicated freshness assessment sub-models for different food types; collecting food samples from the same batch under different storage times and conditions; obtaining multimodal images and corresponding physicochemical test indicators; establishing a correlation model between multimodal features and physicochemical indicators using partial least squares regression; determining the weights and scoring thresholds of freshness features for different food types; inputting the fused feature vector of the food to be tested into the corresponding type of freshness assessment sub-model; calculating the comprehensive freshness score; and classifying the food into four grades—special grade, grade one, grade two, and unqualified—based on the scoring thresholds.

[0013] Furthermore, a computer vision-based food inspection system includes the following modules: The multimodal image acquisition module simultaneously acquires visible light, near-infrared, and hyperspectral images of the food to be tested, and controls the lighting conditions and shooting angle of the acquisition environment; The image preprocessing module performs denoising, illumination compensation, segmentation and normalization on the acquired multimodal images, and extracts the region of interest of the food subject. The multi-scale feature extraction module extracts color, shape, texture, edge, and spectral features from images of different modalities, generating a multi-scale, multi-modal feature set. The fusion detection and evaluation module adaptively fuses multimodal features, performs defect classification, foreign object detection, and freshness assessment, and generates detection results. The results output and storage module grades food based on test results, generates standardized test reports, and stores all test data and image information.

[0014] Furthermore, the multimodal image acquisition module includes a ring-shaped white light source, a near-infrared LED light source, a hyperspectral imager, a color industrial camera, a near-infrared industrial camera, an electric rotary stage, a polarizer assembly, and an environmental parameter acquisition unit. The ring-shaped white light source and the near-infrared LED light source adopt a zoned independent control design, automatically adjusting the brightness and illumination angle of each area according to the type of food to be detected. The polarizer assembly adjusts the polarization angle to eliminate surface reflection. The electric rotary stage and the three cameras are synchronously triggered, driving the food to be detected to rotate at a constant speed, realizing 360-degree image acquisition of the entire surface of the food without blind spots.

[0015] Furthermore, the integrated detection and evaluation module includes a GPU-accelerated computing unit, a model storage unit, a real-time inference unit, an edge-cloud collaboration unit, and an anomaly alarm unit. The GPU-accelerated computing unit uses a parallel computing architecture to process multimodal image data and feature calculations. The model storage unit uses a hierarchical storage architecture, storing detection models for commonly used food types locally and storing full-category models and training datasets in the cloud. The real-time inference unit performs real-time calculations for feature fusion, defect classification, foreign object detection, and freshness assessment, outputting detection results and generating visualized labeled images. The edge-cloud collaboration unit enables collaborative linkage between real-time detection at the edge and big data analysis in the cloud. The anomaly alarm unit automatically triggers audible and visual alarms when non-compliant food is detected, and sends sorting and control instructions to complete the sorting and removal of non-compliant food.

[0016] Compared with existing technologies, the beneficial effects of this invention are: This invention employs visible light, near-infrared, and hyperspectral multimodal imaging technologies to simultaneously acquire surface appearance and internal composition information of food, enabling integrated detection of appearance defects, foreign matter contamination, and internal quality, covering multiple dimensions of food quality inspection. Through an adaptive multimodal feature fusion method, the fusion weights of each modality feature can be dynamically adjusted according to the priority of the detection task and the food type, fully leveraging the advantages of different modal features in different detection tasks and improving the effectiveness and specificity of feature representation.

[0017] This invention assesses the freshness of food by region, corrects the assessment results by incorporating environmental temperature and humidity data, and establishes a correlation between freshness scores and remaining shelf life. This more accurately reflects the actual quality status of food and provides a reference for food storage and distribution. The introduction of multi-scale feature extraction and attention mechanisms effectively detects defects and foreign objects of different sizes. By combining the fusion of continuous multi-frame detection results and false defect discrimination rules, the probability of false detections and false negatives is reduced, improving the accuracy of defect and foreign object detection.

[0018] The detection system of this invention employs independently controlled light sources and polarization correction technology, enabling it to adapt to the imaging needs of different foods and eliminate the influence of surface reflections on the detection results. A motorized rotating stage achieves 360-degree image acquisition of the entire food surface, avoiding the omission of defects in various parts of the food. An edge-cloud collaborative architecture is adopted, with the cloud continuously optimizing model parameters based on accumulated detection data and distributing them to the edge, improving the model's generalization ability and update efficiency. The overall solution enhances the automation and intelligence level of food detection, providing reliable technical support for quality control in the food production process. Attached Figure Description

[0019] Figure 1 A complete logic diagram of the entire process of multimodal food detection; Figure 2Flowchart for image preprocessing and region of interest extraction; Figure 3 A diagram illustrating multi-scale, multi-modal feature fusion strategies; Figure 4 A graph illustrating the parallel analysis of defect identification, foreign object detection, and freshness assessment; Figure 5 This is a diagram illustrating the hardware architecture and edge-cloud collaboration logic of a food testing system. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0022] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified. Furthermore, the terms "installed," "connected," and "linked" should be interpreted broadly; for example, they may refer to a fixed connection, a detachable connection, or an integral connection; they may refer to a mechanical connection or an electrical connection; they may refer to a direct connection or an indirect connection through an intermediate medium; and they may refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances. The invention will now be described in further detail with reference to the accompanying drawings.

[0023] Reference Figures 1 to 5 A computer vision-based food detection method includes the following steps: Visible light color images, near-infrared grayscale images, and hyperspectral cube images of the food to be tested are acquired simultaneously using a multimodal imaging device, and the acquired raw images are processed for frame synchronization and spatial registration. The registered multimodal images are preprocessed to remove image noise and illumination interference, extract the region of interest of the food subject and normalize its size; Color and shape features are extracted from visible light images, texture and edge features are extracted from near-infrared images, and spectral reflectance features are extracted from hyperspectral images to construct a multi-scale, multi-modal feature set. Adaptive fusion of multi-scale and multi-modal features is performed to generate a fused feature vector. The fused feature vector is then input into a pre-trained food defect classification model to identify the types and locations of damage, mold, and discoloration defects on the food surface. The fused feature vector is input into the anomaly detection model to detect foreign objects such as metal, plastic, and hair mixed in with food, and a foreign object detection mask is generated. A food freshness assessment model was established based on hyperspectral and color features to quantitatively calculate the freshness level of the food to be tested. Based on the results of defect detection, foreign object detection, and freshness grade, the food is graded according to preset standards, and a test report is generated and the test data is stored.

[0024] This invention also includes a multimodal feature adaptive fusion step, which determines the fusion weights by calculating the discriminative power and stability of different modal features. The calculation formula is as follows: ; in, For the first The fusion weights of modal features are dimensionless. For the first Inter-class discriminant strength of modal features, dimensionless. For the first Intra-class stability of modal features, dimensionless. The modal index has values ​​of 1, 2, and 3, corresponding to visible light, near-infrared, and hyperspectral modes, respectively. Before fusion, sub-pixel-level spatial alignment of different modal features is performed based on feature points on the food surface to achieve a one-to-one correspondence between each modal feature in spatial position. During the fusion process, the weight allocation strategy is dynamically adjusted according to the priority of the detection task. The defect detection task increases the basic weight of texture and edge features, and the freshness assessment task increases the basic weight of color and spectral features. At the same time, a modal confidence verification mechanism is introduced. When the feature variance of a certain modality exceeds a preset threshold, its weight is automatically reduced and the complementary weights of other modalities are increased to adapt to the feature differences of different food types such as fruits and vegetables, nuts, and meats.

[0025] This invention also includes a comprehensive quantitative evaluation step for food freshness, which calculates a comprehensive freshness score by integrating changes in multi-dimensional features. The calculation formula is as follows: ; in, This is a comprehensive score for food freshness, dimensionless, with a value ranging from 0 to 100. The color feature variation weighting coefficient is dimensionless. The Euclidean distance between the color characteristics of the food to be tested and the color characteristics of a fresh reference sample is dimensionless. The weighting coefficients for texture feature variations are dimensionless. The cosine distance between the texture features of the food to be detected and the texture features of a fresh reference sample is dimensionless. The weighting coefficient for hyperspectral feature variations is dimensionless. The value is the average absolute value of the difference between the reflectance of the characteristic band of the food to be tested and the reflectance of the corresponding band of the fresh reference sample. It is dimensionless. During the evaluation process, the different functional areas of the food are extracted and scored independently. Then, the comprehensive score is obtained by weighting the influence of each area on freshness. At the same time, the freshness score is corrected by combining the ambient temperature and humidity data at the time of collection. A correlation mapping model between the freshness score and the remaining shelf life is established to output the expected storage time of the food to be tested. When the difference in scores of each area exceeds the preset threshold, it is marked as local deterioration and the warning level is raised.

[0026] In this invention, the multimodal image preprocessing steps specifically include: using a bilateral filtering algorithm to remove Gaussian noise and salt-and-pepper noise from visible light and near-infrared images; using a wavelet transform denoising method to preserve spectral details for hyperspectral images; using an adaptive histogram equalization algorithm based on the reflectivity of food surfaces to compensate for image brightness differences caused by uneven illumination; introducing a polarization correction step for food surfaces with severe reflectivity to eliminate highlight areas caused by specular reflection; using a background extraction method combined with semantic segmentation to accurately segment the food subject from complex backgrounds such as conveyor belts, packaging paper, and labels; extracting the minimum bounding rectangle of the food subject as the region of interest; uniformly scaling all regions of interest to a fixed size; using morphological opening operations to remove fine noise and edge burrs in the image; using a grayscale normalization algorithm to map image pixel values ​​to the 0-1 range; and achieving consistency of images acquired at different times through batch grayscale mean calibration.

[0027] In this invention, the multi-scale multimodal feature extraction steps specifically include: constructing a three-layer feature pyramid to extract features at low, medium, and high resolution scales; extracting the mean, variance, skewness, and kurtosis statistical features of each channel in the RGB, HSV, and Lab color spaces from the visible light image; extracting the perimeter, area, roundness, aspect ratio, and concavity / convexity shape features of the food outline; extracting local binary patterns, gray-level co-occurrence matrix texture features, and orientation gradient histogram edge features at different scales from the near-infrared image; using a genetic algorithm to select the 20 feature bands with the highest correlation to changes in food composition in the hyperspectral image; extracting the average reflectance, reflectance gradient, and spectral absorption depth features of each feature band; using the mutual information method to filter and remove redundant features; performing Z-score standardization on all retained features to eliminate differences in the dimensions of different features; and generating a multi-scale multimodal feature set with unified dimensions.

[0028] In this invention, the food defect and foreign object detection steps specifically include: constructing an improved convolutional neural network combining channel attention and spatial attention mechanisms as a defect classification model; introducing a multi-scale feature fusion module into the network to detect defects and foreign objects of different sizes; training the model using a food image dataset labeled with defect type, location, and severity; using a focus loss function to address the imbalance between positive and negative samples during training; inputting the fused feature vector into the trained defect classification model to output the defect type, confidence level, and bounding box coordinates; constructing an anomaly detection model using an isolated forest algorithm combined with template matching; establishing a multimodal feature template library for normal food; comparing the fused feature vector of the food to be detected with the template library to identify foreign object regions exceeding the normal feature distribution; generating a detection mask containing the location, size, and type of foreign objects; introducing a multi-frame detection result fusion mechanism; combining temporal information to eliminate false detection results from a single detection; and establishing a false defect discrimination rule to distinguish between natural textures and real defects on the food surface.

[0029] In this invention, the food freshness assessment steps specifically include: establishing dedicated freshness assessment sub-models for different food types such as leafy vegetables, melons and fruits, nuts, and meats; collecting food samples from the same batch under different storage times and conditions; acquiring their multimodal images and corresponding physicochemical test indicators; establishing a correlation model between multimodal features and physicochemical indicators using partial least squares regression; determining the freshness feature weights and scoring thresholds for different food types; inputting the fused feature vector of the food to be tested into the corresponding type of freshness assessment sub-model; calculating the comprehensive freshness score; classifying the food into four grades—special grade, grade one, grade two, and unqualified—based on the scoring thresholds; introducing an abnormal freshness detection mechanism to identify foods that have undergone chemical preservation treatment; determining whether the food has undergone artificial processing by comparing abnormal shifts in spectral features; and continuously updating the assessment model parameters using an incremental learning method. When a new food type is added, only the top-level parameters of the model need to be fine-tuned to complete the adaptation.

[0030] This invention includes the following modules: The multimodal image acquisition module is used to simultaneously acquire visible light, near-infrared and hyperspectral images of the food to be tested, and to control the lighting conditions and shooting angle of the acquisition environment; The image preprocessing module is used to perform noise reduction, illumination compensation, segmentation and normalization on the acquired multimodal images, and to extract the region of interest of the food subject. The multi-scale feature extraction module is used to extract color, shape, texture, edge and spectral features from images of different modalities, and generate a multi-scale multimodal feature set; The fusion detection and evaluation module is used to adaptively fuse multimodal features, perform defect classification, foreign object detection and freshness evaluation, and generate detection results; The results output and storage module is used to grade food based on test results, generate standardized test reports, and store all test data and image information.

[0031] In this invention, the multimodal image acquisition module includes a ring-shaped white light source, a near-infrared LED light source, a hyperspectral imager, a color industrial camera, a near-infrared industrial camera, an electric rotary stage, a polarizer assembly, and an environmental parameter acquisition unit. The ring-shaped white light source and the near-infrared LED light source adopt a zoned independent control design, which can automatically adjust the brightness and illumination angle of the light source in each area according to the type of food to be detected. The polarizer assembly is installed at the front of the camera lens and can be rotated to adjust the polarization angle to eliminate surface reflection. The hyperspectral imager has a spectral range covering 400nm to 1000nm, a spectral resolution of 5nm, and a spatial resolution of 1024×768 pixels. The color industrial camera and the near-infrared industrial camera both have a resolution of 5 million pixels and a frame rate of 30 frames / second. The electric rotary stage is synchronously triggered with the three cameras, driving the food to be detected to rotate at a constant speed, realizing 360-degree image acquisition of the entire surface of the food without blind spots. The environmental parameter acquisition unit collects temperature, humidity, and light intensity data of the detection environment in real time and transmits them to the fusion detection and evaluation module.

[0032] In this invention, the fusion detection and evaluation module includes a GPU-accelerated computing unit, a model storage unit, a real-time inference unit, an edge-cloud collaboration unit, and an anomaly alarm unit. The GPU-accelerated computing unit uses a parallel computing architecture to process multimodal image data and feature calculations, with feature extraction and fusion time for a single food image not exceeding 50 milliseconds. The model storage unit adopts a hierarchical storage architecture, storing detection models for commonly used food types locally and storing full-category models and training datasets in the cloud, supporting online updates and incremental training of models. The real-time inference unit performs real-time calculations for feature fusion, defect classification, foreign object detection, and freshness assessment, outputting detection results and generating visualized labeled images. The edge-cloud collaboration unit enables collaborative work between real-time detection at the edge and big data analysis in the cloud. The cloud periodically optimizes model parameters based on accumulated detection data and distributes them to the edge. The anomaly alarm unit automatically triggers an audible and visual alarm and sends a control signal to the sorting device when non-conforming food is detected, achieving automatic sorting and removal of non-conforming food.

[0033] This invention relates to a computer vision-based food inspection method and system. By fusing surface appearance and internal composition information of food using multimodal imaging technology, and combining adaptive feature fusion and deep learning algorithms, it achieves integrated detection of food defects, foreign objects, and freshness. The system is deployed on food production lines and can adapt to the inspection needs of different types of food. The inspection process is fully automated and requires no manual intervention. The invention will be further described in detail below with reference to two specific embodiments.

[0034] Example 1:

[0035] This embodiment is applied to a leafy vegetable processing line, with fresh spinach as the test subject. The line operates at a speed of one piece per second, and the tests include leaf damage, yellowing, mold, hair and plastic foreign matter contamination, as well as freshness grading. The system is installed at the front end of the sorting station on the production line and is linked with the automatic sorting device to achieve seamless integration of testing and sorting.

[0036] After system startup, the multimodal image acquisition module enters standby mode. When the spinach arrives at the inspection station via the conveyor belt, the photoelectric sensor triggers the color industrial camera, near-infrared industrial camera, and hyperspectral imager to simultaneously acquire images. The ring-shaped white light source and near-infrared LED light source automatically adjust the brightness and illumination angle of each zone of the light source according to the imaging characteristics of the spinach. The polarizer assembly rotates to a preset polarization angle to eliminate specular reflections caused by water droplets on the spinach surface. The motorized rotary table and the three cameras are triggered synchronously, causing the spinach to rotate one revolution at a constant speed, completing 360-degree image acquisition of the entire surface. The environmental parameter acquisition unit collects real-time data on temperature, humidity, and light intensity of the inspection environment and transmits it synchronously to the fusion detection and evaluation module.

[0037] The image preprocessing module receives the raw image data and first performs frame synchronization processing on the multimodal images. Based on the leaf vein feature points on the spinach surface, it completes sub-pixel-level spatial alignment of different modal images, ensuring a one-to-one correspondence between the features of each modality in spatial location. A bilateral filtering algorithm is used to remove Gaussian and salt-and-pepper noise from visible light and near-infrared images. For hyperspectral images, a wavelet transform denoising method is used to preserve spectral detail. An adaptive histogram equalization algorithm based on the reflectivity of food surfaces is used to compensate for image brightness differences caused by uneven illumination. A background extraction method combined with semantic segmentation is used to accurately segment the spinach subject and the conveyor belt background, extracting the minimum bounding rectangle of the spinach subject as the region of interest (ROI), and uniformly scaling all ROIs to a fixed size. Morphological opening operations are used to remove fine noise and edge burrs in the image. A gray-level normalization algorithm maps image pixel values ​​to the 0-1 range, and batch-wise gray-level mean calibration ensures consistency of images acquired at different times.

[0038] A multi-scale feature extraction module constructs a three-layer feature pyramid to extract features at low, medium, and high resolution scales. In visible light images, it extracts the mean, variance, skewness, and kurtosis statistical features of each channel in the RGB, HSV, and Lab color spaces, and extracts the perimeter, area, roundness, aspect ratio, and concavity / convexity shape features of the spinach outline. In near-infrared images, it extracts local binary patterns, gray-level co-occurrence matrix texture features, and histogram edge features of directional gradients at different scales. In hyperspectral images, a genetic algorithm is used to select the 20 feature bands most correlated with changes in spinach chlorophyll content, and the average reflectance, reflectance gradient, and spectral absorption depth features of each feature band are extracted. A mutual information method is used to filter and remove redundant features, and all retained features are Z-score standardized to eliminate dimensional differences between features, generating a multi-scale, multi-modal feature set with uniform dimensions.

[0039] The fusion detection and evaluation module first performs modal confidence verification, calculating the intra-class variance of each modality's features. When the variance of a certain modality exceeds a preset threshold, its fusion weight is automatically reduced while the complementary weights of other modalities are increased. The basic weight allocation is dynamically adjusted according to the priority of the detection task; for defect detection, the basic weights of texture and edge features are increased, while for freshness evaluation, the basic weights of color and spectral features are increased. The generated fusion feature vector is input into an improved convolutional neural network combining channel attention and spatial attention mechanisms. This network employs a multi-scale feature fusion module, capable of detecting defects of different sizes. The network outputs the defect types, confidence scores, and bounding box coordinates for leaf damage, yellowing, and mold. The fusion feature vector is then input into an isolated forest anomaly detection model and compared with a multi-modal feature template library of normal spinach to identify foreign object regions exceeding the normal feature distribution, generating a detection mask containing the location, size, and type of the foreign object. A fusion mechanism for continuous multi-frame detection results is introduced, combining temporal information to eliminate false detections from single detections, and establishing a false defect discrimination rule to distinguish between natural leaf veins and genuine damage in spinach.

[0040] During the freshness assessment, features were extracted and independently scored from the leaf and petiole areas of spinach. A comprehensive score was obtained by weighting the impact of the leaves and petioles on the freshness of the spinach. The freshness score was corrected by incorporating ambient temperature and humidity data collected during the assessment, and a correlation mapping model between the freshness score and remaining shelf life was established to output the estimated shelf life of the spinach. When the score difference between the leaf and petiole areas exceeded a preset threshold, it was marked as localized spoilage and the warning level was raised. Based on the freshness score threshold, the spinach was classified into four grades: premium, grade one, grade two, and unqualified.

[0041] The results output and storage module generates a standardized inspection report based on the inspection results, recording the inspection time, inspection items, grading results, and corresponding image information, and stores all data in a local database. When substandard spinach is detected, the anomaly alarm unit triggers an audible and visual alarm, sends a control signal to the automatic sorting device, and removes the substandard spinach to the corresponding collection box. The edge-cloud collaboration unit periodically uploads local inspection data to the cloud. The cloud uses the accumulated inspection data to incrementally train the model, and the optimized model parameters are distributed to the edge devices to achieve continuous updates to the inspection model.

[0042] Table 1: Comparison of Detection Results for Leafy Vegetables Table 1 compares the application effects of traditional single-modal visual inspection methods and the method of this invention in the inspection of leafy vegetables. The improved accuracy of surface defect detection stems from the multi-modal feature fusion and multi-scale detection module's ability to identify minute defects; the false defect discrimination rule effectively distinguishes between natural leaf veins and damaged areas. The improved accuracy of small foreign object detection benefits from the sensitivity of near-infrared and hyperspectral features to low-contrast foreign objects; the continuous multi-frame fusion mechanism reduces false detections caused by random factors. Freshness assessment incorporates hyperspectral features reflecting internal components; regional assessment and environmental parameter correction make the results more closely reflect the actual quality state of the food. The GPU parallel computing architecture ensures that the detection speed can match the production rhythm of the assembly line.

[0043] Example 2:

[0044] This embodiment is applied to the pre-packaging inspection of nuts, specifically walnuts in their shells. The production line operates at a speed of 2 pieces per second. Inspection includes checking for shell damage, cracks, internal mold, contamination with metal or plastic foreign objects, and freshness grading. The system is integrated into the front end of the automatic packaging machine; walnuts that pass inspection proceed directly to the packaging process, while unqualified walnuts are automatically rejected.

[0045] After system startup, the multimodal image acquisition module enters operational mode. When the walnuts arrive at the inspection station via conveyor belt, the position sensor triggers three cameras to simultaneously acquire images. The ring-shaped white light source and near-infrared LED light source are adjusted to suitable brightness and angle for imaging the walnut shell, and the polarizer assembly rotates to the corresponding angle to eliminate specular reflections on the walnut shell. The motorized rotary table rotates the walnuts one revolution, acquiring visible light, near-infrared, and hyperspectral images of the entire surface. The hyperspectral imager covers a spectral range of 400nm to 1000nm, capable of capturing spectral feature differences caused by changes in the internal composition of the walnut. The environmental parameter acquisition unit records the temperature and humidity data of the inspection environment in real time and transmits them to the fusion detection and evaluation module for freshness correction.

[0046] The image preprocessing module performs frame synchronization and spatial registration on the acquired raw images, and completes sub-pixel-level spatial alignment of images of different modalities based on the texture feature points of the walnut shell. Bilateral filtering is used to remove noise from visible light and near-infrared images, and wavelet transform is used to denoise the hyperspectral images while preserving spectral details. Adaptive histogram equalization compensates for brightness differences caused by uneven illumination, and a semantic segmentation network is used to segment the walnut body and the conveyor belt background. The minimum bounding rectangle of the walnut is extracted as the region of interest and uniformly scaled to a fixed size. Morphological opening operations are performed to remove small noise points in the images, pixel values ​​are normalized to the 0-1 range, and batch grayscale mean calibration ensures consistency between different batches of images.

[0047] A multi-scale feature extraction module constructs a three-layer feature pyramid to extract multimodal features at three resolution scales. Statistical features of the RGB, HSV, and Lab color spaces are extracted from visible light images, along with the perimeter, area, roundness, aspect ratio, and concavity / convexity shape features of the walnut shell outline. Local binary patterns, gray-level co-occurrence matrix texture features, and directional gradient histogram edge features at different scales are extracted from near-infrared images. In hyperspectral images, a genetic algorithm is used to select the 20 feature bands most correlated with walnut fat oxidation and mold growth, extracting the average reflectance, reflectance gradient, and spectral absorption depth features for each band. Redundant features are removed using the mutual information method, and the retained features are Z-score normalized to generate a multi-scale, multimodal feature set.

[0048] The fusion detection and evaluation module first performs modal confidence verification and dynamically adjusts the fusion weights based on the variance of each modal feature. The defect detection task increases the base weights of texture and edge features, while the freshness evaluation task increases the base weights of spectral and color features. The fused feature vector is input into an improved convolutional neural network, which combines channel attention and spatial attention mechanisms to enhance the feature response of defective regions. The network outputs the defect type, confidence score, and bounding box coordinates for shell breakage and cracking. The fused feature vector is then input into an isolated forest anomaly detection model, compared with a feature template library of normal walnuts, to detect foreign objects such as metal and plastic and generate detection masks. A multi-frame detection result fusion mechanism eliminates false detections from single detections, and a false defect discrimination rule distinguishes between the natural texture and cracked defects of the walnut shell.

[0049] During the freshness assessment process, features corresponding to the walnut shell and interior are extracted and scored independently. A comprehensive score is obtained by weighting the impact of shell and interior quality on the walnut's freshness. The freshness score is corrected by incorporating environmental temperature and humidity data, and a correlation mapping model between the freshness score and remaining shelf life is established to output the estimated shelf life of the walnuts. An abnormal freshness detection mechanism is introduced to identify walnuts that have undergone chemical preservation treatment by comparing abnormal shifts in spectral features. Based on the freshness score threshold, walnuts are classified into four grades: premium, first-grade, second-grade, and unqualified. For different food types such as leafy vegetables, fruits, nuts, and meats, the system establishes dedicated freshness assessment sub-models. When adding new food types, only minor adjustments to the top-level parameters of the model are needed for adaptation.

[0050] The results output and storage module generates a standardized inspection report and stores all inspection data and image information. When a substandard walnut is detected, the anomaly alarm unit triggers an audible and visual alarm, sends a control signal to the sorting device, and removes the substandard walnut. The edge-cloud collaboration unit uploads local inspection data to the cloud, where the cloud incrementally trains the model based on accumulated data, and the optimized model parameters are then distributed to the edge. The model storage unit adopts a hierarchical storage architecture, storing inspection models for commonly used food types locally, and storing all categories of models and training datasets in the cloud, supporting online updates and replacements of models.

[0051] Table 2: Comparison of Detection Results for Nuts Table 2 illustrates the advantages of the method of this invention in the detection of nut products. The improved accuracy of shell defect detection stems from multi-scale feature extraction and the attention mechanism's enhancement of edge details, enabling accurate identification of minute shell cracks. Internal mold detection is a unique capability of this invention; hyperspectral imaging can non-destructively capture the spectral features generated by changes in the internal composition of walnuts, solving the problem that traditional methods cannot detect internal quality. Foreign object detection shows good recognition results for foreign objects of different materials such as metal and plastic; the edge-cloud collaborative architecture allows for continuous model optimization, adapting to quality differences between different batches of nuts. The overall detection speed can meet the production needs of high-speed packaging lines.

[0052] Reference Figure 1 This figure illustrates the overall method flow of this invention. Starting with the initial synchronous acquisition of multimodal images, after image preprocessing and spatial registration, the system enters the feature extraction stage. Through an adaptive fusion mechanism, the system transforms multidimensional features into fused feature vectors, and then executes three core tasks in parallel: defect classification, foreign object detection, and freshness measurement. Finally, the system grades food based on the comprehensive detection results and generates standardized reports and stores the data, achieving fully automated food quality monitoring.

[0053] Reference Figure 2 This figure details the standardization process of multimodal images before they enter the algorithm. For images acquired from different light sources, the system employs bilateral filtering and wavelet transform for denoising. Adaptive histogram equalization and polarization correction are used to eliminate uneven illumination and reflection interference. Subsequently, semantic segmentation technology is used to remove background elements such as conveyor belts, accurately extracting the minimum bounding rectangle of the food subject, and normalizing its size and grayscale to ensure consistency in subsequent feature extraction.

[0054] Reference Figure 3 This figure illustrates how the system processes feature information from different dimensions. The system constructs a feature pyramid, extracting features such as color, shape, texture, edge, and spectral reflectance from three modalities. The core "adaptive fusion logic" dynamically adjusts weights based on the priority of the current detection task. Simultaneously, modality confidence verification is introduced; if data from a particular modality is abnormal, the system automatically reduces its weight and increases the contribution of complementary modalities, ensuring the accuracy and stability of the fused feature vector.

[0055] Reference Figure 4 This figure illustrates the flow of fused feature vectors across three independent evaluation models. The defect classification model uses a convolutional neural network to locate damage or mold; the anomaly detection model compares the sample with a normal template library and uses an isolated forest algorithm to identify foreign objects such as plastic and hair; the freshness assessment model combines a physicochemical index correlation model to score different areas and adjusts for environmental temperature and humidity to ultimately calculate the remaining shelf life. This parallel architecture ensures comprehensive and real-time detection.

[0056] Reference Figure 5 This diagram illustrates the hardware components and data processing architecture of the detection system. The hardware includes light source control, multi-camera synchronous acquisition, and an electric rotary table. Data processing employs an edge-cloud collaborative model: the edge relies on GPU-accelerated computing units to perform millisecond-level real-time inference and automatic sorting control; the cloud is responsible for storing the full-category large dataset, performing incremental model learning and parameter optimization. The optimized model is periodically deployed to the edge, enabling continuous evolution of the system's detection capabilities.

[0057] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A food detection method based on computer vision, characterized in that, Includes the following steps: Visible light color images, near-infrared grayscale images, and hyperspectral cube images of the food to be tested are acquired simultaneously using a multimodal imaging device, and the acquired raw images are processed for frame synchronization and spatial registration. The registered multimodal images are preprocessed to remove image noise and illumination interference, extract the region of interest of the food subject and normalize its size; Color and shape features are extracted from visible light images, texture and edge features are extracted from near-infrared images, and spectral reflectance features are extracted from hyperspectral images to construct a multi-scale, multi-modal feature set. Adaptive fusion of multi-scale and multi-modal features is performed to generate a fused feature vector. The fused feature vector is then input into a pre-trained food defect classification model to identify the types and locations of damage, mold, and discoloration defects on the food surface. The fused feature vector is input into the anomaly detection model to detect foreign objects such as metal, plastic, and hair mixed in with food, and a foreign object detection mask is generated. A food freshness assessment model was established based on hyperspectral and color features to quantitatively calculate the freshness level of the food to be tested. Based on the results of defect detection, foreign object detection, and freshness grade, the food is graded according to preset standards, and a test report is generated and the test data is stored.

2. The food detection method based on computer vision according to claim 1, characterized in that, It also includes a multimodal feature adaptive fusion step, which determines the fusion weights by calculating the discriminative power and stability of different modal features. The calculation formula is as follows: ; in, For the first The fusion weights of each modal feature For the first Inter-class discriminative power of modal features For the first Intra-class stability of modal features For modal indexing, before fusion, sub-pixel-level spatial alignment of different modal features is completed based on feature points on the food surface. During the fusion process, the weight allocation strategy is dynamically adjusted according to the priority of the detection task. At the same time, a modal confidence verification mechanism is introduced. When the feature variance of a certain modality exceeds the preset threshold, its weight is automatically reduced and the complementary weights of the other modalities are increased.

3. The food detection method based on computer vision according to claim 1, characterized in that, It also includes a comprehensive quantitative assessment step for food freshness, which calculates a comprehensive freshness score by integrating changes in multi-dimensional characteristics. The calculation formula is as follows: ; in, The overall score for food freshness The weighting coefficient for color feature changes. The Euclidean distance is the color feature of the food to be tested compared with the color feature of a fresh reference sample. The weighting coefficients for texture feature changes. Let be the cosine distance between the texture features of the food to be detected and the texture features of the fresh reference sample. The weighting coefficients represent the changes in hyperspectral characteristics. This is the average of the absolute values ​​of the differences between the reflectance of the characteristic band of the food to be tested and the reflectance of the corresponding band of the fresh reference sample.

4. The food detection method based on computer vision according to claim 1, characterized in that, The multimodal image preprocessing steps include: using a bilateral filtering algorithm to remove Gaussian noise and salt-and-pepper noise from visible light and near-infrared images; using wavelet transform denoising to preserve spectral details for hyperspectral images; using an adaptive histogram equalization algorithm to compensate for image brightness differences; introducing a polarization correction step for food surfaces with severe reflectivity to eliminate highlight areas caused by specular reflection; using a background extraction method combined with semantic segmentation to segment the food subject and background, extracting the minimum bounding rectangle of the food subject as the region of interest; uniformly scaling all regions of interest to a fixed size; using morphological opening operations to remove fine noise and edge burrs in the image; and using a grayscale normalization algorithm to map image pixel values ​​to the 0 to 1 range.

5. The food detection method based on computer vision according to claim 1, characterized in that, The multi-scale, multi-modal feature extraction steps include: constructing a three-layer feature pyramid to extract features at low, medium, and high resolution scales; extracting the mean, variance, skewness, and kurtosis statistical features of each channel in the RGB, HSV, and Lab color spaces from the visible light image; extracting the perimeter, area, roundness, aspect ratio, and concavity / convexity shape features of the food outline; extracting local binary patterns, gray-level co-occurrence matrix texture features, and orientation gradient histogram edge features at different scales from the near-infrared image; using a genetic algorithm to select the 20 feature bands most correlated with changes in food composition from the hyperspectral image; extracting the average reflectance, reflectance gradient, and spectral absorption depth features of each feature band; and generating a multi-scale, multi-modal feature set with unified dimensions.

6. The food detection method based on computer vision according to claim 1, characterized in that, The steps for detecting food defects and foreign objects include: constructing an improved convolutional neural network combining channel attention and spatial attention mechanisms as a defect classification model; introducing a multi-scale feature fusion module into the network to detect defects and foreign objects; training the model using a food image dataset labeled with defect type, location, and severity; inputting the fused feature vector into the trained defect classification model to output defect type, confidence level, and bounding box coordinates; constructing an anomaly detection model using an isolated forest algorithm combined with template matching; establishing a multimodal feature template library for normal food; comparing the fused feature vector of the food to be detected with the template library to identify foreign object regions that exceed the normal feature distribution and generating a detection mask.

7. The food detection method based on computer vision according to claim 1, characterized in that, The food freshness assessment steps include: establishing dedicated freshness assessment sub-models for different food types; collecting food samples from the same batch under different storage times and conditions; obtaining multimodal images and corresponding physicochemical test indicators; establishing a correlation model between multimodal features and physicochemical indicators using partial least squares regression; determining the weights and scoring thresholds of freshness features for different food types; inputting the fused feature vector of the food to be tested into the corresponding type of freshness assessment sub-model; calculating the comprehensive freshness score; and classifying the food into four grades—special grade, grade one, grade two, and unqualified—based on the scoring thresholds.

8. A computer vision-based food inspection system, applied to the computer vision-based food inspection method according to any one of claims 1-7, characterized in that, Includes the following modules: The multimodal image acquisition module simultaneously acquires visible light, near-infrared, and hyperspectral images of the food to be tested, and controls the lighting conditions and shooting angle of the acquisition environment; The image preprocessing module performs denoising, illumination compensation, segmentation and normalization on the acquired multimodal images, and extracts the region of interest of the food subject. The multi-scale feature extraction module extracts color, shape, texture, edge, and spectral features from images of different modalities, generating a multi-scale, multi-modal feature set. The fusion detection and evaluation module adaptively fuses multimodal features, performs defect classification, foreign object detection, and freshness assessment, and generates detection results. The results output and storage module grades food based on test results, generates standardized test reports, and stores all test data and image information.

9. A food inspection system based on computer vision according to claim 8, characterized in that, The multimodal image acquisition module includes a ring-shaped white light source, a near-infrared LED light source, a hyperspectral imager, a color industrial camera, a near-infrared industrial camera, an electric rotary stage, a polarizer assembly, and an environmental parameter acquisition unit. The ring-shaped white light source and the near-infrared LED light source adopt a zoned independent control design, automatically adjusting the brightness and illumination angle of each area according to the type of food to be detected. The polarizer assembly adjusts the polarization angle to eliminate surface reflection. The electric rotary stage and the three cameras are synchronously triggered, driving the food to be detected to rotate at a constant speed, realizing 360-degree image acquisition of the entire surface of the food without blind spots.

10. A food inspection system based on computer vision according to claim 8, characterized in that, The fusion detection and evaluation module includes a GPU-accelerated computing unit, a model storage unit, a real-time inference unit, an edge-cloud collaboration unit, and an anomaly alarm unit. The GPU-accelerated computing unit uses a parallel computing architecture to process multimodal image data and feature calculations. The model storage unit uses a hierarchical storage architecture, storing detection models for commonly used food types locally and storing models and training datasets for all categories in the cloud. The real-time inference unit performs real-time calculations for feature fusion, defect classification, foreign object detection, and freshness assessment, outputting detection results and generating visualized labeled images. The edge-cloud collaboration unit enables collaborative linkage between real-time detection at the edge and big data analysis in the cloud. The anomaly alarm unit automatically triggers audible and visual alarms when non-compliant food is detected, and sends sorting and control instructions to complete the sorting and removal of non-compliant food.