Real-time Food Defect Detection Method and System Based on Image Processing

By using multimodal data fusion and the YOLO-BioNet model, combined with Kalman filtering and deep reinforcement learning, the computational burden and accuracy issues of AI quality inspection systems in food defect detection have been resolved, achieving efficient and accurate defect detection and intelligent sorting.

CN120783334BActive Publication Date: 2025-11-14咸阳家友缘食品有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511248519.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2025-11-14
Estimated Expiration
2045-09-03

AI Technical Summary

Technical Problem

Existing AI quality inspection systems face problems such as heavy computational burden, decreased detection accuracy, high false alarm and false negative rates, and insufficient adaptive learning capabilities in food defect detection, making it difficult to meet the needs of food safety production and automated quality inspection.

Method used

Multimodal data fusion technology is adopted, which combines frequency domain filtering and CNN denoising. Defect detection is performed through the YOLO-BioNet model. Combined with Kalman filtering, LSTM time series modeling and deep reinforcement learning, abnormal defect patterns are dynamically identified. Defect probability is calculated through Bayesian network and judged by dynamic threshold, triggering sorting instructions and equipment linkage.

Benefits of technology

It significantly improves the detection accuracy and real-time performance under complex food surfaces, enhances adaptability, achieves efficient and accurate defect detection and intelligent sorting, and optimizes the intelligence level of the quality inspection system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120783334B_ABST
    Figure CN120783334B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of machine vision and food quality inspection, specifically a real-time food defect detection method and system based on image processing. Specifically, it integrates multi-modal data, adaptively adjusts weight coefficients to improve data quality, performs multi-level noise reduction, extracts key frames, and strengthens the key frame verification mechanism by controlling variables with 2FA. It constructs a YOLO-BioNet model, optimizes the detection efficiency and accuracy in complex scenarios through dynamic convolution kernel allocation and spiking neural networks. It integrates Kalman filtering, LSTM time series modeling, and deep reinforcement learning models to analyze the change trajectory of the food surface and predict the defect state. It fuses real-time data and historical information to calculate dynamic thresholds through Bayesian networks, conducts automatic sorting instructions and production line linkage responses, and uses distributed storage and graph neural networks to optimize the model and support defect traceability. The present invention has significant advantages in terms of real-time performance, detection accuracy, and adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of machine vision and food quality inspection technology, specifically to a method and system for real-time detection of food defects based on image processing. Background Technology

[0002] With the rapid development of artificial intelligence (AI) and computer vision technologies, intelligent defect detection systems are increasingly widely used in food processing, quality control, and packaging inspection. Traditional manual quality inspection methods mainly rely on visual inspection, which not only consumes a lot of human resources but is also prone to missed or incorrect inspections due to human negligence. In addition, traditional methods are difficult to accurately identify subtle defects when faced with high-speed production lines and complex food surfaces, and cannot meet the needs of real-time online sorting. In recent years, breakthroughs in artificial intelligence technologies such as deep learning, reinforcement learning, and adaptive learning have enabled AI-based image analysis to show great potential in defect detection, morphological recognition, and anomaly identification.

[0003] Chinese invention patent CN119006469B discloses an automatic detection method and system for substrate glass surface defects based on machine vision. The method includes acquiring and preprocessing multimodal data of the substrate glass surface; fusing extracted image features, acoustic features, and thermal imaging features across modalities; inputting the fused feature map into a pre-constructed graph neural network to obtain a topological feature map representing the defect structure; concatenating the topological feature map with the fused feature map and inputting it into a multi-task learning network to obtain segmented suspected defect regions and corresponding defect types; using a pre-trained target detection model to locate defects, determining the precise location of the defects through bounding box regression to obtain an image of the target defect region; uploading the target defect region image to the cloud, and using a defect detection model deployed in the cloud to accurately identify the defect type of the target defect region to obtain the final defect detection result.

[0004] Current AI-based quality inspection systems still face numerous challenges, such as heavy computational burden due to large image data volumes, decreased defect detection accuracy in environments with complex and variable food surface textures, limited intelligent judgment capabilities for subtle or internal defects, high false alarm and false negative rates, and a lack of adaptive learning capabilities. Therefore, there is an urgent need for a real-time food defect detection method based on image processing. This method should construct models using deep learning, reinforcement learning, edge computing, and big data analysis to achieve efficient and accurate defect detection, intelligent defect morphology analysis, real-time defect event judgment, intelligent sorting linkage, and long-term optimization learning. This would improve the intelligence level and sorting efficiency of the quality inspection system, meeting the needs of food safety production and automated quality inspection. Summary of the Invention

[0005] The purpose of this invention is to address the problems existing in the background technology by proposing a real-time food defect detection method and system based on image processing.

[0006] The technical solution of this invention: a real-time food defect detection method based on image processing, comprising the following specific implementation steps:

[0007] S1. Acquire multimodal data: RGB data, near-infrared image data, and X-ray data;

[0008] S2. Based on a multi-level denoising method, multi-modal data is dynamically fused through adaptive weight coefficients, combined with frequency domain filtering and CNN denoising, key frames are screened using surface change energy, and two-factor authentication control variables are generated based on hash functions and finite field matrix operations.

[0009] S3. By dynamically regulating the keyframe logic stability through two-factor authentication control variables, a YOLO-BioNet model is constructed to simulate the pulse coding of biological neurons and dynamic convolution kernel allocation. The biological dynamic processing module is chain-stacked to optimize multi-scale feature extraction, and adaptive Bi-FPN is used to enhance cross-layer semantic interaction. The classification and regression tasks are decoupled through dynamic sparse self-attention and deformable anchor box prediction, and key defect information {defect category, confidence, bounding box coordinates} is output.

[0010] S4. Smooth the trajectory of the defect area and predict the position change by Kalman filtering to generate defect features. Combine LSTM temporal modeling and deep reinforcement learning to dynamically identify abnormal defect patterns. Calculate the defect probability using a Bayesian network.

[0011] S5. Generate a comprehensive signal by integrating defect probability, real-time morphological features and historical defect data, and make defect judgment based on dynamic threshold.

[0012] S6 triggers sorting instructions, production line alarms, and equipment linkage, and optimizes sorting strategies by combining quality inspection feedback and reinforcement learning.

[0013] S7. Defect backtracking and model optimization are performed through distributed storage and graph neural networks.

[0014] Preferably, the structure of the YOLO-BioNet model is as follows:

[0015] The backbone network constructs a hybrid convolutional-spiking neural network module, which mimics the dynamic pulse coding segmentation of biological neurons to generate pulse activation maps and output feature maps for four time steps from the input keyframe image. The multi-scale feature extraction process is dynamically optimized by chaining biological dynamic processing modules.

[0016] The neck network uses bidirectional cross-scale attention to update multi-scale features and predicts the offset field through 3×3 convolution to achieve dynamic sampling and output of deformable feature maps.

[0017] The classification branch of the detection head uses dynamic sparse self-attention to calculate the non-zero activation ratio, the regression branch predicts the offset and dynamically adjusts the anchor box position size, and dynamically sets the positive sample allocation threshold based on the average IoU.

[0018] Output the key defect information detected {defect category, confidence level, bounding box coordinates}.

[0019] Preferably, the backbone network structure is as follows:

[0020] The convolutional-spiking neural network module includes:

[0021] Input: A 640×640 1×3 keyframe image based on the input;

[0022] Pulsed Convolutional Layer: The pulsed convolutional layer uses a 3×3 convolutional kernel to output 64 channels. Combined with LIF spiking neurons, the input is divided into 4 time steps. Pulses are fired through membrane potential accumulation and threshold triggering mechanism, and the frequency map of the output pulse activation intensity is generated.

[0023] Dynamic depthwise separable convolution: The number of convolution kernels is determined by the average activation δ of the pulse layer output: if δ>0.5, i.e., a high-complexity region, the number of convolution kernels KN=64; if δ≤0.5, i.e., a simple region, the number of convolution kernels KN=32; then, a 3×3 depthwise separable convolution is performed independently on each channel, followed by the SiLU activation function, outputting an 80×80×64 feature map;

[0024] The biological dynamic processing modules are stacked in a chain, including:

[0025] Input: 80×80×64 feature map output by dynamic depthwise separable convolution, with the scale progressively downsampled to 40×40 and 20×20;

[0026] Module structure: An attention-gated network is constructed to perform global average pooling and global max pooling on the input features, and a channel weight vector W is generated through a shared multilayer perceptron. c Furthermore, spatial attention is used to perform 3×3 convolution on the input features to generate a spatial weight matrix W. s And W c and W s Element-wise multiplication yields the dynamic weight W: ;

[0027] Where σ is the sigmoid activation function;

[0028] Conditional branch selection: If the mean of W exceeds the set threshold T w=0.7, activates the Transformer branch, the activation layer contains a branch of 2 Transformer encoder layers: based on multi-head self-attention and feedforward network, the output is a feature map with residual connection; conversely, activates dynamic routing, which adjusts the number of channels only through 1×1 convolution;

[0029] Output: Dynamically weighted fusion of original features and branch outputs, output feature map F out :

[0030] ;

[0031] Among them, F in For input features; F branch This indicates the branch output characteristics.

[0032] Preferably, the keyframe filtering process is as follows:

[0033] The surface change energy is used to calculate the change energy E between adjacent frames. m (t):

[0034] ;

[0035] In the formula, u t (x,y), v t (x,y) represent the horizontal and vertical components of the pixel (x,y) respectively;

[0036] Define the key threshold T m If E m (t)>T m If so, then the frame is a keyframe;

[0037] Output any keyframe IK i ;

[0038] Obtain the keyframe set SetK={IK1,IK2,…,IK i ,…,IK m}, where m is the total number of keyframes.

[0039] Preferably, the process for generating the two-factor authentication control variables is as follows:

[0040] The keyframes {IK1,IK2,…,IK i ,…,IK m The data is encoded by combining the following: data = H(IK1||IK2||…||IK) m )∈GF(p) n×n ;

[0041] Where || denotes a concatenation operation; H() is a predefined hash function that maps messages to an n×n matrix over GF(p); GF(p) n×n Let GF(p) be the set of all n×n matrices whose elements belong to GF(p); GF(p) is a predefined finite field where p is a predefined prime number and the field GF(p) contains p elements, i.e., integers 0, 1, 2, ..., p-1.

[0042] Calculate basic variables ;

[0043] Here, A and B are two predefined invertible n×n matrices whose elements belong to the finite field GF(p);

[0044] By concatenating the diagonal elements of data with the bit value Fv∈GF(p), the comprehensive variable SV=f is calculated. -1 (Fv) mod p;

[0045] Where f() represents a predefined quadratic polynomial, f(x) = ax 2 +bx+c∈GF(p)[x];GF(p)[x] denotes a polynomial ring with x as a variable and coefficients belonging to GF(p);

[0046] Generate the two-factor authentication control variable 2FV = (basic variable BV, comprehensive variable SV).

[0047] Preferably, the control process for dynamically adjusting the keyframe logic stability through two-factor authentication control variables is as follows:

[0048] Regenerate encoded data =H(IK1||IK2||…||IK m )∈GF(p) n×n ,Will Diagonal element concatenation bit value ∈GF(p);

[0049] Calculate basic constraints ;

[0050] Calculate the synthesis constraint SS = h(SV) mod p;

[0051] Where h() is a predefined analytic function, i.e., the composition polynomial h(x) = f(f(x)) mod p;

[0052] If BS = BV and SS = If the condition is met, it indicates that the logical stability of the keyframe has been constrained; otherwise, an immediate warning is issued.

[0053] The preferred method for calculating the anomaly probability is as follows:

[0054] Construct the location information of the defect area and perform temporal modeling of the morphological changes of the defect: extract the bounding box coordinates of the defect area, record the location changes of the defect area in chronological order, construct the trajectory of the defect area, define the position of the defect area at time t, use Kalman filtering to smooth and predict the target trajectory, and calculate the position of the target at the next time step.

[0055] Extracting morphological features from the trajectory of the defect area provides input for subsequent pattern learning. Specifically, this involves calculating the rate of change and acceleration of change in the region to generate morphological feature vectors.

[0056] The morphological feature vector of the defect is input into the Long Short-Term Memory (LSTM) network for temporal modeling, and the output of the LSTM model is fed into the Deep Reinforcement Learning (DRL) model. The long-term pattern of the target behavior is learned through the reward function to identify abnormal defect patterns. The policy update in the DRL is performed through the Q-learning method.

[0057] Based on the outputs of LSTM and DRL, the future form of defects is predicted, and defects are detected through a Bayesian network, outputting anomaly probability values.

[0058] Preferably, the anomaly detection process based on dynamic thresholds is as follows:

[0059] Obtain the defect probability at time t Extract morphological features, location, velocity, acceleration, and historical sequence data;

[0060] The weighted integration of defect probability, morphological feature position change rate V(t), change acceleration a(t), and historical data generates a comprehensive input signal I(t), with coefficients automatically adjusted through training;

[0061] Set the dynamic threshold function Td(t): ;

[0062] Among them, V max and A max These represent the maximum values ​​of the set rate of change and acceleration, respectively. , , Indicates the adjustment factor;

[0063] Defect events are determined based on the comprehensive input signal I(t) and the dynamic threshold Td(t);

[0064] If I(t) > Td(t), then a sorting alarm is triggered and a sorting instruction is sent.

[0065] If I(t)≤Td(t), then it is considered that there is no defect at present.

[0066] The preferred multi-stage noise reduction process is as follows:

[0067] The collected multimodal data is fused and modeled to obtain fused data I. F (x,y,t):

[0068] ;

[0069] ; ; ;

[0070] Where α, β, and γ represent weighting coefficients, based on the ambient brightness L. t Adaptive adjustment; L0 represents the preset threshold; λ and k represent adjustment parameters that determine the adaptability under different lighting conditions;

[0071] For image frame I F Fourier transform F of (x,y,t): ;

[0072] Filtering is performed using a Gaussian filter H(u,v): ;

[0073] ;

[0074] Where u and v represent frequency components; σ1 controls the filter strength;

[0075] Denoising is performed using a convolutional neural network (CNN), with the input noisy image I. F (x,y,t), the denoised output I is calculated using a CNN. D (x,y,t)=f θ (I F (x,y,t));

[0076] Among them, f θ () represents a CNN denoising network; θ represents the trained parameters;

[0077] The denoising network is trained based on the mean squared error loss function: ;

[0078] in, Represents the i-th noisy image; This represents a noise-free image; N represents the number of training samples.

[0079] The technical solution of the present invention: a real-time food defect detection system based on image processing, which is used to execute the above-mentioned real-time food defect detection method based on image processing, comprising:

[0080] The data acquisition module is used to collect real-time multimodal data;

[0081] The edge computing module is used to preprocess the acquired multimodal data;

[0082] The AI ​​analysis module analyzes the acquired monitoring data, including:

[0083] The defect detection unit is used to build the YOLO-BioNet model and detect critical defects.

[0084] The morphological analysis unit is used to construct defect morphological change models and analyze defect evolution patterns.

[0085] The defect determination unit is used to build a defect determination model and output defect detection results.

[0086] The sorting linkage module is used to automatically sort in conjunction with production line equipment.

[0087] The cloud computing and big data analytics module is used to manage historical defect data using distributed storage, supports fast indexing and backtracking, and combines time series analysis and graph neural networks for defect tracing and batch correlation analysis.

[0088] Compared with the prior art, the above-mentioned technical solution of the present invention has the following beneficial technical effects:

[0089] This invention designs a real-time food defect detection method and system based on image processing. It integrates RGB images, near-infrared spectroscopy, and X-ray transmission information through multimodal data fusion technology, and dynamically adjusts the fusion weight coefficients based on food characteristics (such as surface reflectivity and transmittance), significantly improving data quality and detection accuracy for complex food surfaces (such as varied textures, uneven lighting, and internal structures). Furthermore, it performs adaptive denoising, frequency domain joint denoising, and surface change energy-driven keyframe selection on the multimodal data. It combines 2FA control variables with finite-field matrix operations and hash verification on keyframes to constrain the logical stability of keyframes, optimizing real-time performance and anti-interference capabilities. An improved YOLO- The BioNet defect detection model enhances the efficiency of feature extraction from complex food surfaces through a biomimetic spiking neural network and a dynamic convolutional kernel allocation mechanism. It also improves the robustness of detecting subtle defects by utilizing self-supervised pre-training and a dynamic positive sample allocation strategy. By integrating Kalman filtering trajectory smoothing, LSTM temporal modeling, and deep reinforcement learning, it dynamically learns defect morphology evolution patterns and predicts defect states. Furthermore, it combines Bayesian networks and dynamic threshold functions to fuse real-time features with historical defect data. Through production line equipment, it achieves automatic responses to sorting instructions and alarms. Finally, by leveraging the distributed storage of a cloud computing platform and graph neural network optimization of model parameters, it supports defect batch backtracking and the recognition of novel defect patterns.

[0090] This invention achieves synergistic optimization in detection accuracy, real-time performance, adaptability, and security, possessing significant technical advantages and application value. Attached Figure Description

[0091] Figure 1 This is a system architecture diagram of a real-time food defect detection system based on image processing proposed in this invention;

[0092] Figure 2 This is a flowchart of a real-time food defect detection method based on image processing proposed in this invention. Detailed Implementation

[0093] Example 1, as Figure 1 As shown, the present invention proposes a real-time food defect detection system based on image processing, comprising: a data acquisition module, an edge computing module, an AI analysis module, a cloud computing and big data analysis module, and an alarm and sorting linkage module.

[0094] The data acquisition module acquires real-time image data of the food surface and interior using high-resolution industrial cameras (RGB, near-infrared, X-ray).

[0095] The edge computing module preprocesses the acquired image data;

[0096] The AI ​​analysis module analyzes the acquired image data, including: a defect detection unit, a morphological analysis unit, and a defect determination unit;

[0097] The defect detection unit uses a deep learning model to detect key defect areas, including but not limited to mold, insect infestation, damage, foreign objects, and internal cavities.

[0098] The morphological analysis unit constructs a defect morphological evolution model and analyzes defect change patterns, including but not limited to mold spread, damage expansion, and foreign object movement.

[0099] The defect determination unit constructs a defect determination model and outputs detection results, including but not limited to critical defects, minor defects, and suspected defects.

[0100] The sorting linkage module combines with production line equipment (including but not limited to robotic arms, sorting push rods, and alarms) to automatically sort and push defect information to operators through the MES system or management platform.

[0101] The cloud computing and big data analytics module uses distributed storage to manage historical defect data, supports fast indexing and backtracking, and combines time series analysis and graph neural networks (GNN) to achieve defect batch tracing and production line correlation analysis.

[0102] Example 2, as Figure 2As shown, the present invention proposes a real-time food defect detection method based on image processing, which is applied to a real-time food defect detection system based on image processing proposed in Example 1. The specific implementation steps are as follows:

[0103] S1. The data acquisition module continuously collects multimodal data using an industrial camera to obtain real-time information on food flowing through the detection area. This data is then combined with RGB images, near-infrared (NIR) spectra, and X-ray transmission to form multimodal data {RGB image frame data I...} RGB (x,y,t), near-infrared image data I NIR The multimodal data (x,y,t) and X-ray data X(x,y,t) are transmitted to the edge computing module.

[0104] Where x and y represent pixel coordinates; t represents the number of time frames.

[0105] S2, the edge computing module improves data acquisition quality through multimodal fusion, adaptive denoising, and edge computing optimization, providing high-quality input to the intelligent defect detection system and ensuring the accuracy of defect identification and morphological analysis. The specific implementation process is as follows:

[0106] S21. Perform fusion modeling on the collected multimodal data to obtain fused data I. F (x,y,t):

[0107] ;

[0108] ; ; ;

[0109] Where α, β, and γ represent weighting coefficients, based on the ambient brightness L. t Adaptive adjustment; L0 represents the preset threshold; λ and k represent adjustment parameters that determine the adaptability under different lighting conditions;

[0110] S22. Combining frequency domain denoising and CNN denoising, multi-level noise reduction is performed on edge devices to improve real-time performance while ensuring detail preservation. Specifically:

[0111] S2201, For image frame I F Fourier transform F of (x,y,t):

[0112] ;

[0113] Filtering is performed using a Gaussian filter H(u,v): ;

[0114] ;

[0115] Where u and v represent frequency components; σ1 controls the filtering strength, a larger value will lead to loss of detail, while a smaller value may not be able to remove noise;

[0116] S2202. For environments with strong noise, a convolutional neural network (CNN) is used for denoising. The input noisy image I... F (x,y,t), the denoised output I is calculated using a CNN. D (x,y,t)=f θ (I F (x,y,t));

[0117] Among them, f θ () represents a CNN denoising network; θ represents the trained parameters;

[0118] The denoising network is trained based on the mean squared error (MSE) loss function: ;

[0119] in, Represents the i-th noisy image; This represents a noise-free image; N represents the number of training samples.

[0120] Therefore, the parameter θ is optimized using the gradient descent method to achieve the best quality of the denoised video.

[0121] S23. Select keyframes through surface change energy calculation to optimize computational resources and improve the efficiency of subsequent AI analysis, specifically:

[0122] (1) Surface change energy calculation: The surface change energy is used to calculate the change energy E between adjacent frames. m (t):

[0123] ;

[0124] In the formula, u t (x,y), v t (x,y) represent the horizontal and vertical components of the pixel (x,y) respectively;

[0125] Define the key threshold T m If E m (t)>T m If so, then the frame is a keyframe;

[0126] (2) Output keyframe IK i ;

[0127] S24. Obtain the keyframe set SetK={IK1,IK2,…,IK…} i ,…,IK m}, where m is the total number of keyframes;

[0128] S25. Construct a 2FA (Two-Factor Authentication Control Variable) generation model to generate 2FA control variables for the keyframe set. The 2FA control variable generation process is as follows:

[0129] S2501. Combine the keyframes for encoding to obtain the encoded data: data=H(IK1||IK2||…||IK m )∈GF(p) n×n ;

[0130] Where || denotes a concatenation operation; H() is a predefined hash function that maps messages to an n×n matrix over GF(p); GF(p) n×n Let GF(p) be the set of all n×n matrices whose elements belong to GF(p); GF(p) is a predefined finite field where p is a predefined prime number and the field GF(p) contains p elements (i.e., integers 0, 1, 2, ..., p-1);

[0131] S2502, Calculate the basic variable BV:

[0132] ;

[0133] Here, A and B are two predefined invertible n×n matrices whose elements belong to the finite field GF(p);

[0134] S2503. Concatenate the diagonal elements of data to obtain the position value Fv∈GF(p), and calculate the comprehensive variable SV=f. -1 (Fv) mod p;

[0135] Where f() represents a predefined quadratic polynomial, f(x) = ax 2 +bx+c∈GF(p)[x];GF(p)[x] denotes a polynomial ring with x as a variable and coefficients belonging to GF(p);

[0136] S2504, Generate 2FA control variable 2FV = (basic variable BV, composite variable SV);

[0137] S26. Set the keyframe set SetK = {IK1,IK2,…,IK} to {2FV = (basic variable BV, composite variable SV). i ,…,IK m The data is then transmitted to the defect detection unit.

[0138] S3. The defect detection unit constructs the YOLO-BioNet defect detection model. By mimicking the dynamic adaptability of biological vision, introducing self-supervised prior knowledge, and employing a multi-scale interaction mechanism, it significantly improves detection performance in complex scenes while maintaining the real-time performance of the YOLOv8 model. The specific architecture of the model is as follows:

[0139] S31. Input Layer: Construct a 2FA (Two-Factor Authentication) control variable dynamic adjustment model, receiving {2FV = (basic variable BV, comprehensive variable SV)}, keyframe set SetK = {IK1, IK2, ..., IK}. i ,…,IK m Extract 2FV = (basic variable BV, comprehensive variable SV) and the keyframe set SetK = {IK1, IK2, ..., IK}. i ,…,IK m To constrain the logical stability of keyframes and ensure their functional thresholds, the 2FA control variable 2FV is dynamically adjusted using two-factor authentication. The adjustment process is as follows:

[0140] S3101, Regenerate encoded data =H(IK1||IK2||…||IK m )∈GF(p) n×n ,Will Diagonal element concatenation bit value ∈GF(p);

[0141] S3102. Calculate the basic constraint BS:

[0142] ;

[0143] S3103, Calculate the comprehensive constraint SS = h(SV) mod p;

[0144] Where h() is a predefined analytic function, i.e., the composition polynomial h(x) = f(f(x)) mod p;

[0145] S3104, If BS = BV and SS = This indicates that the logical stability of the keyframes has been constrained and the functional thresholds of the keyframes have been guaranteed. The keyframe set SetK={IK1,IK2,…,IK} is defined as follows. i ,…,IK m Input is fed into the YOLO-BioNet model; otherwise, an alert is issued immediately.

[0146] S32. Backbone Network: Extracts multi-scale features from the input keyframe images and dynamically adjusts the allocation of computing resources through a biomimetic mechanism to balance efficiency and accuracy. Specifically:

[0147] (1) Hybrid Conv-SNN Block (Hybrid Convolutional-Spiking Neural Network Module) mimics the dynamic pulse coding of biological neurons, dividing the input image into multiple time steps (T=4), and generating a pulse activation map through membrane potential accumulation and threshold triggering (LIF neurons), specifically:

[0148] (a) Input: A 640×640 1×3 keyframe image based on the input, including:

[0149] (b) Spiking Convolutional Layer:

[0150] Structure: Based on a 3×3 convolution kernel, with 64 output channels, followed by a Leaky Integrate-and-Fire (LIF) spiking neuron;

[0151] Pulse coding: The input image is divided into 4 time steps (T=4). The input at each time step is convolved and the membrane potential is accumulated. When the membrane potential exceeds a threshold (V... th When the value is 0.5, a pulse is emitted:

[0152] Output: Pulse firing frequency diagram (80×80×64):

[0153] ;

[0154] Among them, Spike(x) i ) represents the pulse output (0 or 1) at time step i; S(t) is the pulse firing frequency diagram, i.e., the pulse activation intensity at each position; x i This represents the input feature map at the i-th time step;

[0155] (c) Dynamic Depthwise Convolution:

[0156] Dynamic kernel number selection: The number of convolutional kernels is determined based on the average activation δ (range 0~1) of the pulse layer output.

[0157] If δ > 0.5 (high complexity region), the number of convolution kernels KN = 64;

[0158] If δ≤0.5 (simple region), the number of convolution kernels KN=32;

[0159] Operation: Perform a 3×3 depthwise separable convolution independently on each channel, followed by the SiLU activation function;

[0160] Output: 80×80×64 feature map;

[0161] (2) BDP (Bio-Dynamic Processing) module chain stacking: Through dynamic attention mechanism and conditional branch selection, the multi-scale feature extraction process is optimized to balance computational efficiency and detection accuracy. Its core function is to imitate the dynamic adaptability of the biological vision system and adjust the feature processing strategy according to different scene complexities.

[0162] (a) Input an 80×80×64 feature map output by a dynamically depth-separable convolution, and progressively downsample the scale to 40×40 and 20×20, including:

[0163] (b) BDP module structure:

[0164] Attention-gated network: Performs global average pooling (GAP) and global max pooling (GMP) on the input features, and generates channel weight vector W through a shared MLP (multilayer perceptron). c ;

[0165] Spatial attention: Perform a 3×3 convolution on the input features to generate a spatial weight matrix W. s ;

[0166] Fusion weights: W c and W s Element-wise multiplication yields the dynamic weight W:

[0167] ;

[0168] Where σ is the sigmoid activation function;

[0169] Conditional branch selection:

[0170] Transformer branch activation condition: If the mean of W exceeds a set threshold T w =0.7, the activation layer contains two branches of transformer encoder: based on multi-head self-attention (3 heads, hidden layer dimension 256) and feedforward network (FFN expansion ratio 4), the output is a feature map after residual connection;

[0171] Dynamic routing: When not active, skip the Transformer branch and adjust the number of channels only through 1×1 convolution;

[0172] Output: Dynamically weighted fusion of original features and branch outputs, output feature map F out : ;

[0173] Among them, F in For input features; Fbranch Indicates branch output characteristics;

[0174] S33. Neck Network: Adaptive Bi-FPN fuses multi-scale features to enhance cross-layer semantic interaction and spatial alignment capabilities, specifically:

[0175] (1) Bi-Cross Attention, the input is three scale feature maps from the backbone (P3: 80×80, P4: 40×40, P5: 20×20), specifically:

[0176] (a) High-level → Low-level semantic guidance (P5 → P4):

[0177] Query matrix generation: High-level feature P5 generates the Query matrix Q through 1×1 convolution. h ;

[0178] Key / Value Generation: Low-level feature P4 generates a key matrix K through 1×1 convolution. l and Value matrix V l ;

[0179] Attention calculation: ;

[0180] Where, d k The dimension representing the query and key (d in this example) k =64), used to scale the dot product result and prevent gradient explosion;

[0181] (b) Low-level → High-level detail compensation (P4 → P5): Repeat the above process repeatedly, generate a query from low-level features, generate a key / value from high-level features, and update the high-level features;

[0182] (2) Deformable feature fusion:

[0183] Input: Multi-scale features updated by Bi-Cross Attention;

[0184] Offset prediction: For each scale of feature map, the offset field Δp is predicted by 3×3 convolution;

[0185] Deformable convolution: Dynamically samples the feature map according to Δp. ;

[0186] Among them, w k represents the learnable weights; p represents the feature map location; K is the number of kernels in the deformable convolution, K=3;

[0187] Output: Aligned multi-scale feature maps (F3, F4, F5);

[0188] S34. Detection head: Decouples classification and regression tasks, dynamically optimizes positive sample allocation, specifically as follows:

[0189] (1) Classification branch: Dynamic Sparse Self-Attention (DS-SA):

[0190] Input: Multi-scale feature maps (F3, F4, F5) output by the neck network;

[0191] Feature sparsity calculation: the proportion of non-zero activations in the statistical feature map. ;

[0192] Where H represents the width of the feature, i.e. the number of pixels in the feature map in the vertical direction; W represents the width of the feature, i.e. the number of pixels in the feature map in the horizontal direction; C represents the number of channels, i.e. the number of different features contained in the feature map; ||F||0 represents the number of non-zero activations in the feature map, which measures sparsity; ρ represents the sparsity ratio, which is used to dynamically select the computation region.

[0193] (2) Regression branch: Deformable anchor frame prediction:

[0194] Input: Multi-scale feature maps (F3, F4, F5) output by the neck network;

[0195] Deformable convolution predicts offsets: Four offsets (Δx, Δy, Δw, Δh) are predicted using 3×3 convolution;

[0196] Dynamic anchor box adjustment: Prediction box = ;

[0197] Among them, (x,y,w) a ,h a This indicates the initial position and size of the preset anchor frame;

[0198] (3) Dynamic positive sample allocation:

[0199] Threshold adjustment: based on the average IoU of the predicted bounding boxes in the current batch (IoU). avg Dynamically adjust the positive sample threshold T s :T s =0.5 + 0.2 × IoU avg ;

[0200] Label assignment: Only when the IoU between the predicted bounding box and the ground truth bounding box exceeds T s When the sample is considered a positive sample;

[0201] S35, Training Strategy:

[0202] (1) Self-supervised pre-training

[0203] (a) Jigsaw puzzle reconstruction task:

[0204] Input: Unlabeled images are segmented into a 3×3 grid, and the order is randomly shuffled;

[0205] Objective: Predict the index of the original permutation (9! possibilities are simplified to 9 positions for classification);

[0206] Loss function: Cross-entropy loss;

[0207] (b) Dynamic occlusion simulation:

[0208] GAN-generated occlusion: Use a lightweight GAN to generate realistic occlusion blocks that are overlaid on the training images;

[0209] Enhance robustness: Force the model to learn locally visible features;

[0210] (2) Supervision and fine-tuning

[0211] Dynamic loss weights:

[0212] Small target enhancement: For targets smaller than 32×32, the regression loss weight is increased by 1.5 times;

[0213] Formula: Regression Loss Weight ;

[0214] Consistency regularization: For different occluded versions of the same image, force consistent classification output (KL divergence constraint);

[0215] Accordingly: the target detection model YOLO-BioNet outputs the detected key targets and transmits the key target information {defect category, confidence level, bounding box coordinates} to the defect determination unit.

[0216] S4. The morphological analysis unit combines temporal modeling, deep reinforcement learning, and adaptive defect pattern learning to analyze the morphological changes of defects. The specific implementation process is as follows:

[0217] S41. Construct the location information of the defect region and perform temporal modeling of the morphological changes of the defect, specifically:

[0218] S4101. Extract the bounding box coordinates of the defect region, record the position changes of the defect region in chronological order, construct the defect region trajectory, and define the position of the defect region at time t as P(t)={x(t),y(t),w(t),h(t)};

[0219] S4102. Kalman filtering is used to smooth and predict the trajectory of the defect area, and the position of the defect area at the next time step is calculated: ;

[0220] Where P(t) is the position of the defect region at time t; A is the state transition matrix; and u(t) is the control data vector, i.e., external influence. The predicted location of the defect area at time t+1;

[0221] Accordingly: Kalman filtering is used to smoothly predict the target trajectory, and the key defect areas detected are accurately modeled, which enhances the ability to capture changes in dynamic defect areas.

[0222] S42. Extract morphological features from the defect area trajectory to provide input for subsequent pattern learning. Specifically, calculate the area change rate V(t) and change acceleration A(t) to generate a morphological feature vector.

[0223] Accordingly: by calculating the rate of change and acceleration of change of the defect area and combining it with the relative positional changes of the environment, richer defect morphological features can be extracted, providing high-quality input for pattern learning;

[0224] S43. Using deep learning technology, the morphological patterns of defective regions are learned, and potential defects are identified, specifically:

[0225] S4301. Input the morphological feature vector of the defect into a Long Short-Term Memory (LSTM) network for temporal modeling to learn the behavior pattern of the target over a period of time.

[0226] S4302. The output of the LSTM model is fed into the Deep Reinforcement Learning (DRL) module. The long-term pattern of defect morphological changes is learned through the reward function to identify abnormal defect patterns. Policy updates in DRL can be performed using the Q-learning method.

[0227] ;

[0228] In the formula, Q(s) t ,a t ) represents state s t Take action a t The value function of r; t This represents the reward at the current moment; Indicates the learning rate; This represents the discount factor, which indicates the degree of influence of future rewards.

[0229] Accordingly: By combining LSTM with deep reinforcement learning (DRL), the system can learn from complex defect morphology change patterns and identify abnormal defect patterns by learning defect morphology change patterns and feeding back a reward mechanism.

[0230] S44. Predict the future form of a defect using the learned defect morphology changes, specifically:

[0231] S4401: Based on the outputs of LSTM and DRL, the future form of defects is predicted, and defect detection is performed using a Bayesian network.

[0232] ;

[0233] Where, P(anomaly|x1,x2,…,x n P(x1,x2,…,x) represents the probability of detecting a defect given the feature data. n P(x1,x2,…,x) represents the prior probability of the observed data. n |anomaly) represents the observed feature data x1,x2,…,x when a defect is detected. n The probability of a given defect occurring is the likelihood of the feature data appearing; P(anomaly) represents the probability of detecting a defect, independent of any specific observation data or feature, describing the probability of detecting a defect without any feature data; x1, x2, ..., x n The input target features include, but are not limited to, defect location and morphological features;

[0234] S4402, Output the defect probability value and transmit the defect probability value to the defect determination unit; Defect determination unit.

[0235] S5. The defect judgment unit integrates defect characteristics, defect probability, and historical defect data. Through defect event detection and intelligent judgment mechanisms, it identifies and sorts defects. The specific implementation process is as follows:

[0236] S51. Obtain the defect occurrence probability P anomaly (t) (i.e., the probability of detecting a defect at time t), extract the morphological features of the defect region {position P(t), rate of change V(t), and acceleration A(t)}, and extract the historical sequence data H of the defect. past (t)(the location, rate of change, and acceleration of change of defect features over a historical period);

[0237] S52, Integration Defect Probability P anomaly (t), the morphological characteristics of the defect region {location P(t), rate of change V(t), and acceleration of change A(t)}, and the defect data H past (t), generating a comprehensive input signal I(t), which is used for subsequent intelligent judgment:

[0238] ;

[0239] Where w1, w2, and w3 represent coefficients that are automatically adjusted during the training process;

[0240] S53. Based on the comprehensive input signal I(t), dynamically set the threshold and judge abnormal events, specifically as follows:

[0241] S5301, Set the dynamic threshold function Td(t): ;

[0242] Among them, V max and A max These represent the maximum values ​​of the set rate of change and acceleration, respectively. , , Indicates the adjustment factor;

[0243] S5302. Based on the comprehensive input signal I(t) and the dynamic threshold Td(t), perform defect event judgment;

[0244] If I(t) > Td(t), a sorting alarm is triggered, and a sorting instruction is sent to the sorting linkage module.

[0245] If I(t)≤Td(t), then it is considered that there is no defect at present.

[0246] S6. Once the sorting linkage module receives a sorting instruction, it immediately triggers the sorting or alarm mechanism, including but not limited to starting the robotic arm for sorting, triggering the alarm, marking defective batches, and linking with the production line control system to automatically remove defective products, suspend the production line, and record defect information. Quality inspectors provide feedback on the judgment results, and the system uses reinforcement learning to adjust the sorting judgment strategy and optimize the sorting accuracy.

[0247] S7, the cloud computing and big data analysis module enables defect batch backtracking and production line correlation analysis, improves the ability to trace the source of quality problems, and regularly analyzes historical defect data to optimize AI models, improve defect detection accuracy, reduce false alarm rate, and enhance the ability to identify new defect patterns.

[0248] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited thereto. Various changes can be made within the scope of knowledge possessed by those skilled in the art without departing from the spirit of the present invention.

Claims

1. A real-time food defect detection method based on image processing, characterized in that, The specific implementation steps include the following: S1. Acquire multimodal data: RGB data, near-infrared image data, and X-ray data; S2. Based on a multi-level denoising method, multi-modal data is dynamically fused through adaptive weight coefficients, combined with frequency domain filtering and CNN denoising, and key frames {IK1,IK2,…,IK} are selected using surface change energy. i ,…,IK m }, and generates two-factor authentication control variables based on hash functions and finite field matrix operations; The process of generating the two-factor authentication control variables is as follows: The keyframes {IK1,IK2,…,IK i ,…,IK m The data is encoded by combining the following: data = H(IK1||IK2||…||IK) m )∈GF(p) n×n ; Where || denotes a concatenation operation; H() is a predefined hash function that maps messages to an n×n matrix over GF(p); GF(p) n×n Let be the set of all n×n matrices whose elements belong to GF(p); GF(p) is a predefined finite field, where p is a predefined prime number, and the field GF(p) contains p elements, i.e., integers 0, 1, 2, ..., p-1; IK i For any keyframe; m is the total number of keyframes; Calculate basic variables ; Here, A and B are two predefined invertible n×n matrices whose elements belong to the finite field GF(p); By concatenating the diagonal elements of data with the bit value Fv∈GF(p), the comprehensive variable SV=f is calculated. -1 (Fv) mod p; Where f() represents a predefined quadratic polynomial, f(x) = ax 2 +bx+c∈GF(p)[x];GF(p)[x] denotes a polynomial ring with x as a variable and coefficients belonging to GF(p); Generate the two-factor authentication control variable 2FV = (basic variable BV, comprehensive variable SV); S3. By dynamically regulating the keyframe logic stability through two-factor authentication control variables, a YOLO-BioNet model is constructed to simulate the pulse coding of biological neurons and dynamic convolution kernel allocation. The biological dynamic processing module is chain-stacked to optimize multi-scale feature extraction, and adaptive Bi-FPN is used to enhance cross-layer semantic interaction. The classification and regression tasks are decoupled through dynamic sparse self-attention and deformable anchor box prediction, and key defect information {defect category, confidence, bounding box coordinates} is output. The process of dynamically adjusting the keyframe logic stability using two-factor authentication control variables is as follows: Regenerate encoded data =H(IK1||IK2||…||IK m )∈GF(p) n×n ,Will Diagonal element concatenation bit value ∈GF(p); Calculate basic constraints ; Calculate the synthesis constraint SS = h(SV) mod p; Where h() is a predefined analytic function, i.e., the composition polynomial h(x) = f(f(x)) mod p; If BS = BV and SS = If yes, it means that the logic stability of the keyframe has been constrained; otherwise, an immediate warning will be issued. The structure of the YOLO-BioNet model is as follows: The backbone network constructs a hybrid convolutional-spiking neural network module, which mimics the dynamic pulse coding segmentation of biological neurons to generate pulse activation maps and output feature maps for four time steps from the input keyframe image. The multi-scale feature extraction process is dynamically optimized by chaining biological dynamic processing modules. The neck network uses bidirectional cross-scale attention to update multi-scale features and predicts the offset field through 3×3 convolution to achieve dynamic sampling and output of deformable feature maps. The classification branch of the detection head uses dynamic sparse self-attention to calculate the non-zero activation ratio, the regression branch predicts the offset and dynamically adjusts the anchor box position size, and dynamically sets the positive sample allocation threshold based on the average IoU. Output the key defect information detected {defect category, confidence level, bounding box coordinates}; S4. Smooth the trajectory of the defect area and predict the position change by Kalman filtering to generate defect features. Combine LSTM temporal modeling and deep reinforcement learning to dynamically identify abnormal defect patterns. Calculate the defect probability using a Bayesian network. S5. Generate a comprehensive signal by integrating defect probability, real-time morphological features and historical defect data, and make defect judgment based on dynamic threshold. S6 triggers sorting instructions, production line alarms, and equipment linkage, and optimizes sorting strategies by combining quality inspection feedback and reinforcement learning. S7. Defect backtracking and model optimization are performed through distributed storage and graph neural networks.

2. The real-time food defect detection method based on image processing according to claim 1, characterized in that, The backbone network structure is as follows: The convolutional-spiking neural network module includes: Input: A 640×640 1×3 keyframe image based on the input; Pulsed Convolutional Layer: The pulsed convolutional layer uses a 3×3 convolutional kernel to output 64 channels. Combined with LIF spiking neurons, the input is divided into 4 time steps. Pulses are fired through membrane potential accumulation and threshold triggering mechanism, and the frequency map of the output pulse activation intensity is generated. Dynamic depthwise separable convolution: The number of convolution kernels is determined by the average activation δ of the pulse layer output: if δ>0.5, i.e., a high-complexity region, the number of convolution kernels KN=64; if δ≤0.5, i.e., a simple region, the number of convolution kernels KN=32; then, a 3×3 depthwise separable convolution is performed independently on each channel, followed by the SiLU activation function, outputting an 80×80×64 feature map; The biological dynamic processing modules are stacked in a chain, including: Input: 80×80×64 feature map output by dynamic depthwise separable convolution, with the scale progressively downsampled to 40×40 and 20×20; Module structure: An attention-gated network is constructed to perform global average pooling and global max pooling on the input features, and a channel weight vector W is generated through a shared multilayer perceptron. c Furthermore, spatial attention is used to perform 3×3 convolution on the input features to generate a spatial weight matrix W. s And W c and W s Element-wise multiplication yields the dynamic weight W: ; Where σ is the sigmoid activation function; Conditional branch selection: If the mean of W exceeds the set threshold T w =0.7, activates the Transformer branch, the activation layer contains a branch of 2 Transformer encoder layers: based on multi-head self-attention and feedforward network, the output is a feature map with residual connection; conversely, activates dynamic routing, which adjusts the number of channels only through 1×1 convolution; Output: Dynamically weighted fusion of original features and branch outputs, output feature map F out : ; Among them, F in For input features; F branch This indicates the branch output characteristics.

3. The real-time food defect detection method based on image processing according to claim 2, characterized in that, The process for selecting keyframes is as follows: The surface change energy is used to calculate the change energy E between adjacent frames. m (t): ; In the formula, u t (x,y), v t (x,y) represent the horizontal and vertical components of the pixel (x,y) respectively; Define the key threshold T m If E m (t)>T m If so, then the frame is a keyframe; Output any keyframe IK i ; Obtain the keyframe set SetK={IK1,IK2,…,IK i ,…,IK m }, where m is the total number of keyframes.

4. The real-time food defect detection method based on image processing according to claim 3, characterized in that, The method for calculating the defect probability is as follows: Construct the location information of the defect area and perform temporal modeling of the morphological changes of the defect: extract the bounding box coordinates of the defect area, record the location changes of the defect area in chronological order, construct the trajectory of the defect area, define the position of the defect area at time t, use Kalman filtering to smooth and predict the target trajectory, and calculate the position of the target at the next time step. Extracting morphological features from the trajectory of the defect area provides input for subsequent pattern learning. Specifically, this involves calculating the rate of change and acceleration of change in the region to generate morphological feature vectors. The morphological feature vector of the defect is input into the Long Short-Term Memory (LSTM) network for temporal modeling, and the output of the LSTM model is fed into the Deep Reinforcement Learning (DRL) model. The long-term pattern of the target behavior is learned through the reward function to identify abnormal defect patterns. The policy update in the DRL is performed through the Q-learning method. Based on the outputs of LSTM and DRL, the future form of defects is predicted, and defects are detected through a Bayesian network, outputting the defect probability value.

5. The real-time food defect detection method based on image processing according to claim 4, characterized in that, The anomaly detection process based on dynamic thresholds is as follows: Obtain the defect probability at time t, and extract morphological features, location, velocity, acceleration, and historical sequence data; Weighted integration defect probability The morphological feature position change rate V(t), change acceleration a(t), and historical data are used to generate a comprehensive input signal I(t), and the coefficients are automatically adjusted through training; Set the dynamic threshold function Td(t): ; Among them, V max and A max These represent the maximum values ​​of the set rate of change and acceleration, respectively. , , Indicates the adjustment factor; Defect events are determined based on the comprehensive input signal I(t) and the dynamic threshold Td(t); If I(t) > Td(t), then a sorting alarm is triggered and a sorting instruction is sent. If I(t)≤Td(t), then it is considered that there is no defect at present.

6. The real-time food defect detection method based on image processing according to claim 1, characterized in that, The multi-stage noise reduction process is as follows: The collected multimodal data is fused and modeled to obtain fused data I. F (x,y,t): ; ; ; ; Where α, β, and γ represent weighting coefficients, based on the ambient brightness L. t Adaptive adjustment; L0 represents the preset threshold; λ and k represent adjustment parameters that determine the adaptability under different lighting conditions; For image frame I F Fourier transform F of (x,y,t): ; Filtering is performed using a Gaussian filter H(u,v): ; ; Where u and v represent frequency components; σ1 controls the filter strength; Denoising is performed using a convolutional neural network (CNN), with the input noisy image I. F (x,y,t), the denoised output I is calculated using a CNN. D (x,y,t)=f θ (I F (x,y,t)); Among them, f θ () represents a CNN denoising network; θ represents the trained parameters; The denoising network is trained based on the mean squared error loss function: ; in, Represents the i-th noisy image; This represents a noise-free image; N represents the number of training samples.

7. A real-time food defect detection system based on image processing, used to execute the real-time food defect detection method based on image processing according to any one of claims 1 to 6, characterized in that, include: The data acquisition module is used to collect real-time multimodal data; The edge computing module is used to preprocess the acquired multimodal data; The AI ​​analysis module analyzes the acquired monitoring data, including: The defect detection unit is used to build the YOLO-BioNet model and detect critical defects. The morphological analysis unit is used to construct defect morphological change models and analyze defect evolution patterns. The defect determination unit is used to build a defect determination model and output defect detection results. The sorting linkage module is used to automatically sort in conjunction with production line equipment. The cloud computing and big data analytics module is used to manage historical defect data using distributed storage, supports fast indexing and backtracking, and combines time series analysis and graph neural networks for defect tracing and batch correlation analysis.

Citation Information

Patent Citations

  • Automatic Detection Method and System for Surface Defects of Substrate Glass Based on Machine Vision

    CN119006469B

  • Video key frame extraction method based on center offset

    CN111639600A

  • Pan-tilt image tracking interaction method and device, equipment and medium

    CN120469566A