A machine vision-based intelligent inspection method and system for conveyor lines

By integrating food image, speed, and type information into a multimodal detection model, and combining shape coding and attention enhancement techniques, the limitations of food vision inspection systems in recognizing various types of food and high-speed conveyor lines, as well as the high false positive rate, have been solved, enabling more refined quality inspection and evaluation.

CN120741483BActive Publication Date: 2025-10-31SHAANXI YIMING FOOD CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511212546.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-10-31
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Existing food visual inspection systems have limited recognition capabilities and high false positive rates when detecting defects in food, especially in multi-category food and high-speed conveyor lines, making them difficult to adapt to highly variable scenarios and speed changes.

Method used

A machine vision-based intelligent inspection method for conveyor lines is proposed. By fusing information on food images, conveyor line speed, and food type, a multimodal detection model is used, combined with shape encoding and attention enhancement techniques, to identify and score food defects. This method includes a multimodal input fusion module, a feature extraction backbone network, a speed perception and type adaptation embedding module, a spatial channel attention enhancement module, a shape encoding branch, a defect detection branch, and a surface coverage judgment branch.

Benefits of technology

It significantly improves the adaptability and accuracy of food conveyor inspection systems, enabling the identification of various food defects, especially problems such as shape distortion, surface damage, and uneven sauce coverage, providing a more comprehensive quality assessment capability, and is suitable for detecting minute defects in complex backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120741483B_ABST
    Figure CN120741483B_ABST
Patent Text Reader

Abstract

This invention discloses a machine vision-based intelligent inspection method and system for conveyor lines, relating to the field of product inspection. The method includes: acquiring the conveyor line's operating speed information and food type information; acquiring food image information of a target area on the conveyor line; and inputting the operating speed information, food type information, and food image information into a food defect detection model to obtain food defect information. The food defect detection model integrates the food image information, operating speed information, and food type information, combining shape encoding and attention enhancement to obtain multi-type food defect identification and scoring. This method, by integrating multimodal perception, structural modeling, and attention enhancement technologies, significantly improves the adaptability, accuracy, and practicality of the intelligent inspection system for food conveyor lines.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of product inspection, and more specifically, to a machine vision-based intelligent inspection method and system for conveyor lines. Background Technology

[0002] With the continuous improvement of automation in the food industry, visual inspection-based quality control methods have gradually replaced traditional manual inspection methods, becoming an important technological path to improve production efficiency and ensure food quality. Currently, most mainstream food visual inspection systems on the market focus on detecting the integrity of outer packaging, label printing quality, and barcode recognition accuracy, with applications primarily concentrated in the post-packaging outgoing process. These methods rely on traditional image processing techniques such as rule templates, edge localization, or image difference analysis. While they perform well in standardized packaging inspection, their ability to identify defects in the food itself (such as pies, smoked and braised products, etc.) is limited.

[0003] On the other hand, research on defect detection in food products is still in its early stages, especially in real-time detection of multiple types of food (such as products with different shapes, varieties, and speeds) on production lines, which still faces several technical challenges. For example, different foods have large differences in appearance and complex structures, and defects manifest in various forms (such as cracks, scorching, uneven coating, shape distortion, etc.), making it difficult for traditional models to adapt to highly variable scenarios. At the same time, changes in the speed of the conveyor line may lead to image blurring or misjudgment, and existing methods mostly do not consider the influence of speed and food type on defect patterns.

[0004] It is necessary to propose a machine vision-based intelligent inspection method and system for conveyor lines to at least solve some of the above-mentioned problems. Summary of the Invention

[0005] The summary section introduces a series of simplified concepts, which will be further explained in detail in the detailed description section. The summary section of this invention is not intended to limit the key features and essential technical features of the claimed technical solution, nor is it intended to determine the scope of protection of the claimed technical solution.

[0006] In a first aspect, the present invention proposes an intelligent inspection method for conveyor lines based on machine vision, the method comprising:

[0007] Obtain information on the conveyor line's operating speed and the type of food;

[0008] Acquire food image information of the target area of ​​the conveyor line;

[0009] The aforementioned operating speed information, food type information, and food image information are input into the food defect detection model to obtain food defect information. The food defect detection model integrates the aforementioned food image information, operating speed information, and food type information, and combines shape encoding and attention enhancement to obtain multi-category food defect identification and scoring.

[0010] In one feasible implementation, the above-mentioned food defect detection model includes:

[0011] The multimodal input fusion module is used to receive and fuse the aforementioned food image information, the aforementioned running speed information, and the aforementioned food type information to form a multimodal input vector, wherein the aforementioned multimodal input vector includes an image vector I and a speed vector I. and type vector ;

[0012] The feature extraction backbone network is used to extract the multi-scale image feature map group F from the image vector I above based on the built-in multi-scale convolutional neural network;

[0013] The velocity-aware and type-adaptive embedding module is used to embed the velocity vector based on the multi-scale image feature map group F and the velocity vector. and the above type vectors Generate a velocity type fusion feature map F';

[0014] The spatial channel attention enhancement module is used to generate enhanced feature maps from the aforementioned multi-scale image feature map group F and the aforementioned velocity type fusion feature map F'. ;

[0015] The shape encoding branch is used to encode the enhanced feature map described above. Food structural morphology is modeled using Fourier edge description and polar coordinate transformation, and deviations from standard shapes are evaluated to obtain a shape consistency score. ;

[0016] The defect detection branch is used to process the aforementioned enhanced feature maps. Defect identification is performed to obtain defect box set B, category set C, confidence score set P, and heatmap. ;

[0017] The surface cover determination branch is used to process the enhanced feature map mentioned above. The integrity of food surface coatings or sauces is assessed to obtain a coverage rating label. .

[0018] In one feasible implementation, the specific processing steps of the speed sensing and type adaptation embedding module mentioned above include:

[0019] Based on the first multilayer sensor, according to the aforementioned velocity vector Generate velocity embedding vector ;

[0020] Based on the second multilayer perceptron, according to the above type vector Generate type embedding vectors ;

[0021] Based on the above velocity embedding vector and the above embedding vector Generate control signals ;

[0022] The above control signals The above multi-scale image feature map group F is applied to each channel to form the above-mentioned velocity type fusion feature map F'.

[0023] In one feasible implementation, the specific processing steps of the aforementioned spatial channel attention enhancement module include:

[0024] The multi-scale image feature map group F and the velocity type fusion feature map F' are weighted and fused to generate a fused feature map. ;

[0025] For the above fusion feature map Perform channel attention modeling to obtain channel-enhanced feature maps. ;

[0026] Enhance the feature map of the above channels Spatial attention waveforms are constructed to generate the enhanced feature maps described above. .

[0027] In one feasible implementation, the specific processing steps of the above-mentioned shape encoding branch include:

[0028] For the above enhanced feature maps Line channel compression and edge detection operations are performed to generate an edge map. ;

[0029] Based on the center point coordinates of the target area mentioned above The above edge map Perform a polar coordinate transformation to obtain a polar coordinate edge map. And extract the radial contour function of the corresponding edge based on each angular direction. ;

[0030] For the above radial profile function Perform a Discrete Fourier Transform to generate the Fourier descriptor of the current product. ;

[0031] The above Fourier descriptor Compared with pre-stored standard template shape Fourier descriptors Perform frequency domain distance calculation to obtain the descriptor difference. ;

[0032] Based on the above descriptor differences The shape consistency score is calculated using a nonlinear mapping function. :

[0033]

[0034] in, The preset rating sensitivity adjustment coefficient, This indicates the degree of consistency between the current food product's structure and shape and the standard template.

[0035] In one feasible implementation, the specific processing steps of the above-mentioned defect detection branch include:

[0036] For the above enhanced feature maps Parallel feature regression and classification are performed using multiple convolutional prediction heads to generate a classification tensor. Regression Tensor and confidence tensor ;

[0037] Based on the above regression tensor Map each candidate point to a corresponding set of defect bounding boxes. , where each bounding box Including center point coordinates and width and height parameters;

[0038] For the above classification tensor With confidence tensor The overall confidence score for each candidate box is calculated jointly. The optimized defect category set is obtained by performing non-maximum suppression based on the score size. and confidence score set ;

[0039] Based on confidence score set The spatial distribution generates a heatmap of defect saliency. .

[0040] In one feasible implementation, the specific processing steps for the above-mentioned surface coverage determination branch include:

[0041] For the above enhanced feature maps Extracting color anomaly response heatmap ;

[0042] The above heat map The system is divided into multiple spatial sub-blocks. The mean and variance of the response are calculated for each sub-block, and the statistical values ​​of all sub-blocks are concatenated to form a spatial consistency descriptor vector. ;

[0043] The above description vector Input a multi-class classifier and output a label indicating the surface coverage integrity level of the food. .

[0044] In one feasible implementation, the above method further includes:

[0045] When the above food type information When indicating that the current target food product is a pie,

[0046] The aforementioned shape encoding branch enhances the feature map. Before performing structural modeling, the above methods include:

[0047] For the above enhanced feature maps Perform directional gradient mapping to obtain a multi-directional gradient map group;

[0048] Based on the gradient map of each orientation, the local extreme response region is calculated and the boundary is traced to extract the refined edge region;

[0049] The refined edge regions described above are input into the polar coordinate transformation and Fourier description process to improve the extraction accuracy and scoring stability of the shape boundaries of pie products.

[0050] In one feasible implementation, the specific training process of the above-mentioned food defect detection model includes:

[0051] A training dataset was collected, which included multiple sample images. The conveyor speed information and food type label corresponding to each training image constituted the triplet training input. ,in, For training images, For training speed values, To train food type labels;

[0052] For the above training images Image enhancement processing is performed, including at least one of random flipping, affine transformation, color perturbation, and blur noise, to expand the training dataset.

[0053] The target output for each training image is labeled, and the target output includes a set of defect bounding boxes. Defect category tag set Defect confidence score Shape consistency score and surface coverage rating label ;

[0054] The above training images The above training speed values and the above-mentioned training food type labels The inputs are combined and fed into the aforementioned food defect detection model;

[0055] Construct the total loss function, and then perform multi-task joint training on the above food defect detection model:

[0056] The food defect detection model is updated using a stochastic gradient descent optimizer to complete the model training.

[0057] Secondly, this invention proposes an intelligent inspection system for conveyor lines based on machine vision, comprising:

[0058] The first acquisition unit is used to acquire information on the operating speed of the conveyor line and information on the type of food.

[0059] The second acquisition unit is used to acquire food image information of the target area of ​​the conveyor line;

[0060] The third acquisition unit is used to input the above-mentioned running speed information, food type information and food image information into the food defect detection model to obtain food defect information. The food defect detection model integrates the above-mentioned food image information, running speed information and food type information, and combines shape encoding and attention enhancement to obtain multi-type food defect identification and scoring.

[0061] In summary, this invention breaks through the limitations of traditional vision systems that only focus on packaging defects, directly targeting the food itself for quality inspection. It can identify various common quality problems in actual production, such as shape distortion, surface damage, and uneven sauce coverage, significantly broadening the application scope of visual quality inspection and improving the granularity and coverage of product quality control. Secondly, this method effectively alleviates the problem of misjudgment that existing models easily encounter when dealing with multiple types and speeds of food by constructing a multimodal detection model that integrates food images, conveyor line speed, and food type information. The speed perception and type adaptation module can adaptively adjust the feature extraction process according to the actual production speed and food type, enhancing the model's robustness to dynamic operating conditions. This method innovatively introduces a shape encoding branch, accurately depicting the food structure boundaries and contours through Fourier description and polar coordinate modeling. It can accurately assess the shape consistency between the target product and the standard template, making it particularly suitable for products with strict structural requirements, such as pies and cookies, achieving more refined structural evaluation and scoring. Furthermore, the proposed spatial-channel attention enhancement module integrates image features and velocity / type control signals, further strengthening the response weights of defect areas through saliency modeling of channel and spatial dimensions, thus improving the model's ability to perceive subtle defects in complex backgrounds. Finally, the system supports multi-task branch collaborative detection, including defect detection, shape scoring, and surface coverage judgment. Compared to traditional single-dimensional detection schemes, it possesses more comprehensive quality assessment capabilities, providing food production enterprises with more reliable quality data support and decision-making basis. In summary, this method, by integrating multimodal perception, structural modeling, and attention enhancement technologies, significantly improves the adaptability, accuracy, and practicality of intelligent detection systems for food conveyor lines.

[0062] The intelligent inspection method for conveyor lines based on machine vision proposed in this invention will be partly apparent from the following description, and partly understood by those skilled in the art through study and practice of the invention. Attached Figure Description

[0063] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit this specification. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0064] Figure 1 A schematic flowchart of a machine vision-based intelligent detection method for conveyor lines is provided for an embodiment of the present invention.

[0065] Figure 2 This is a flowchart illustrating the specific processing steps of a speed sensing and type adaptation embedding module provided in an embodiment of the present invention.

[0066] Figure 3 This is a flowchart illustrating the specific processing steps of a spatial channel attention enhancement module provided in an embodiment of the present invention.

[0067] Figure 4 A flowchart illustrating the specific processing steps of a shape encoding branch provided in an embodiment of the present invention;

[0068] Figure 5 This is a flowchart illustrating the specific processing steps of a defect detection branch according to an embodiment of the present invention.

[0069] Figure 6 This is a flowchart illustrating the specific processing steps for determining a branch based on surface coverage, as provided in an embodiment of the present invention.

[0070] Figure 7 This is a flowchart illustrating the method for structural modeling of the enhanced feature map before the shape encoding branch is applied when the food type information belongs to the pie category, as provided in an embodiment of the present invention.

[0071] Figure 8 This is a schematic diagram illustrating the specific training process of a food defect detection model provided in an embodiment of the present invention.

[0072] Figure 9 This is a schematic diagram of a machine vision-based intelligent inspection system for a conveyor line, provided as an embodiment of the present invention. Detailed Implementation

[0073] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus. The technical solutions of the embodiments of this invention will now be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them.

[0074] Please see Figure 1 This is a flowchart illustrating a machine vision-based intelligent inspection method for conveyor lines, provided by an embodiment of the present invention. Specifically, it may include:

[0075] S110. Obtain information on the operating speed of the conveyor line and the type of food.

[0076] S120: Obtain food image information of the target area of ​​the conveyor line;

[0077] S130. Input the above-mentioned running speed information, food type information and food image information into the food defect detection model to obtain food defect information. The food defect detection model integrates the above-mentioned food image information, running speed information and food type information, and combines shape encoding and attention enhancement to obtain multi-type food defect recognition and scoring.

[0078] For example, this embodiment provides a machine vision-based intelligent inspection method for conveyor lines, which specifically includes the following steps:

[0079] In step S110, the system first acquires the operating speed information of the conveyor line and the type of food currently being conveyed. This information can be read by encoders, sensors, or the control system on the conveyor line, and the system uses a food identification database to label the current food category. For example, if the food being conveyed is a pie, the system labels its type information as "pie." If it is a smoked or braised product, it labels it as "smoked or braised."

[0080] Next, in step S120, food image information is acquired in real time using an industrial camera or image acquisition device installed in the target detection area of ​​the conveyor line. The acquired images cover key areas of the food on the conveyor line, ensuring that visual information such as its surface condition, contour shape, and appearance structure can be observed.

[0081] In step S130, the acquired running speed information, food type information, and image information are jointly input into the constructed food defect detection model. This model is a multimodal fusion architecture, capable of jointly processing multi-source information. Specifically, the model first transforms image, speed, and type information into a unified feature vector representation through a multimodal input fusion module; then, it extracts image features through a multi-scale convolutional network and combines the embedding vectors of speed and type for attention adjustment to obtain a more discriminative fusion feature map. Subsequently, the shape encoding module uses Fourier description and polar coordinate transformation to model the food structure and identify deviations from standard shapes; simultaneously, the defect detection module locates abnormal areas such as cracks, damage, and contamination on the product surface and generates defect categories and confidence scores; furthermore, the model can also evaluate the integrity of sauce coatings or spreads through a surface coverage judgment module. Finally, the system outputs comprehensive results including defect location, category, severity score, shape consistency score, and surface coverage level, providing support for food production quality monitoring and automatic rejection.

[0082] In summary, this invention breaks through the limitations of traditional vision systems that only focus on packaging defects, directly targeting the food itself for quality inspection. It can identify various common quality problems in actual production, such as shape distortion, surface damage, and uneven sauce coverage, significantly broadening the application scope of visual quality inspection and improving the granularity and coverage of product quality control. Secondly, this method effectively alleviates the problem of misjudgment that existing models easily encounter when dealing with multiple types and speeds of food by constructing a multimodal detection model that integrates food images, conveyor line speed, and food type information. The speed perception and type adaptation module can adaptively adjust the feature extraction process according to the actual production speed and food type, enhancing the model's robustness to dynamic operating conditions. This method innovatively introduces a shape encoding branch, accurately depicting the food structure boundaries and contours through Fourier description and polar coordinate modeling. It can accurately assess the shape consistency between the target product and the standard template, making it particularly suitable for products with strict structural requirements, such as pies and cookies, achieving more refined structural evaluation and scoring. Furthermore, the proposed spatial-channel attention enhancement module integrates image features and velocity / type control signals, further strengthening the response weights of defect areas through saliency modeling of channel and spatial dimensions, thus improving the model's ability to perceive subtle defects in complex backgrounds. Finally, the system supports multi-task branch collaborative detection, including defect detection, shape scoring, and surface coverage judgment. Compared to traditional single-dimensional detection schemes, it possesses more comprehensive quality assessment capabilities, providing food production enterprises with more reliable quality data support and decision-making basis. In summary, this method, by integrating multimodal perception, structural modeling, and attention enhancement technologies, significantly improves the adaptability, accuracy, and practicality of intelligent detection systems for food conveyor lines.

[0083] In one feasible implementation, the above-mentioned food defect detection model includes:

[0084] The multimodal input fusion module is used to receive and fuse the aforementioned food image information, the aforementioned running speed information, and the aforementioned food type information to form a multimodal input vector, wherein the aforementioned multimodal input vector includes an image vector I and a speed vector I. and type vector ;

[0085] The feature extraction backbone network is used to extract the multi-scale image feature map group F from the image vector I above based on the built-in multi-scale convolutional neural network;

[0086] The velocity-aware and type-adaptive embedding module is used to embed the velocity vector based on the multi-scale image feature map group F and the velocity vector. and the above type vectors Generate a velocity type fusion feature map F';

[0087] The spatial channel attention enhancement module is used to generate enhanced feature maps from the aforementioned multi-scale image feature map group F and the aforementioned velocity type fusion feature map F'. ;

[0088] The shape encoding branch is used to encode the enhanced feature map described above. Food structural morphology is modeled using Fourier edge description and polar coordinate transformation, and deviations from standard shapes are evaluated to obtain a shape consistency score. ;

[0089] The defect detection branch is used to process the aforementioned enhanced feature maps. Defect identification is performed to obtain defect box set B, category set C, confidence score set P, and heatmap. ;

[0090] The surface cover determination branch is used to process the enhanced feature map mentioned above. The integrity of food surface coatings or sauces is assessed to obtain a coverage rating label. .

[0091] For example, the food defect detection model provided by the present invention integrates image features, conveyor line operating speed information, and food type information. Through multi-module collaborative modeling, it improves the ability and adaptability to identify defects in different types of food during high-speed conveying.

[0092] Specifically, the multi-modal input fusion module first receives three types of input information: food image vectors. velocity vector and type vectors A multimodal input representation is constructed. In this way, the model, while perceiving image information, incorporates speed and category context related to production conditions, providing a foundation for subsequent feature tuning.

[0093] Next, the model internally uses a multi-scale convolutional neural network to process the image vectors. Feature extraction is performed to obtain multi-scale feature maps of the image. This is to enhance the model's ability to identify defects of different granularities.

[0094] Subsequently, the speed perception and type adaptation embedding module will embed the speed vector With type vector Embedding mapping is performed, and control signals are generated to adjust the feature map set. Generate an intermediate representation of fusion speed and type characteristics. This is to improve the model's adaptability to differences in operating conditions and product types.

[0095] To further enhance the feature response of key regions, the spatial channel attention enhancement module models attention weights based on the fusion of image features and velocity type features, thereby generating enhanced feature maps. This significantly improves the model's attention to the location of potential defects.

[0096] Enhanced feature maps The data is fed into three branch processing modules. The first is the shape encoding branch, which extracts edge information through median filtering, models the structural contour of the food using polar coordinate transformation and Fourier descriptor techniques, compares it with a standard shape template, and finally generates a shape consistency score. It is used to measure the structural integrity of food.

[0097] The second branch is the defect detection branch, which uses multiple convolutional prediction heads to enhance the feature map. Perform regression and classification operations to output a set of location boxes for food defects. Category set Confidence score set and spatial distribution heat map This enables precise location and qualitative analysis of food defects.

[0098] The third branch is the surface coverage assessment branch, which focuses on analyzing the surface distribution of sauces, coatings, etc. It combines color response heatmaps to statistically analyze the distribution characteristics of multiple spatial sub-blocks, and finally uses a classifier to determine the surface coverage integrity level label. It is suitable for determining whether there is uneven coverage, missed coating, or other issues.

[0099] In summary, this implementation method achieves deep integration of three types of features: speed-type adjustment, structural modeling, and defect thermal visualization. It possesses high robustness, high accuracy, and multi-task recognition capabilities, making it particularly suitable for industrial scenarios with strict requirements for food appearance quality and frequent changes in production lines.

[0100] In one feasible implementation, such as Figure 2 As shown, Figure 2 This is a flowchart illustrating the specific processing steps of a speed sensing and type adaptation embedding module according to an embodiment of the present invention. The specific processing steps of the speed sensing and type adaptation embedding module include:

[0101] S210, Based on the first multilayer sensor according to the above velocity vector Generate velocity embedding vector ;

[0102] S220, Based on the second multilayer perceptron according to the above type vector Generate type embedding vectors ;

[0103] S230, Based on the above velocity embedding vector and the above embedding vector Generate control signals ;

[0104] S240, The above control signal The above multi-scale image feature map group F is applied to each channel to form the above-mentioned velocity type fusion feature map F'.

[0105] For example, the speed perception and type adaptation embedding module proposed in this invention is used to convert the operating speed information of the conveyor line and the food type information into control signals for an adjustable image feature extraction process, thereby enabling the food defect detection model to have stronger adaptability and discrimination ability when dealing with different operating states and different food categories.

[0106] Specifically, this includes: the system first uses a first multilayer perceptron (MLP) ), the input velocity vector Mapped to a velocity embedding vector This process can be represented as:

[0107]

[0108] in: : Represents the original velocity feature vector; : The weight matrix of the first perceptron; : Bias term; : Nonlinear activation functions, such as ReLU or SiLU; : Velocity embedding representation, dimension is .

[0109] In parallel, the system uses a second multilayer perceptron ( ), food type vector Mapped to type embedding vector The specific formula is as follows:

[0110]

[0111] in: : Original food type vector; The weight matrix of the second perceptron; : Bias term; Type embedding vector, dimension is .

[0112] Subsequently, the two embedded vectors are fused to generate a control signal vector. This is used to adjust image channel features. Fusion methods can include weighted summation, linear transformation after stitching, Hadamard product, etc. The following is an example of implementing stitching + linear projection:

[0113]

[0114] in: : indicates the concatenation of two embedding vectors; Linear transformation weights of the control signal It is the number of channels in the image feature map; : Bias term; : Control signal vector, used for adjustment of each channel.

[0115] Finally, the control signal As a channel-level adjustment factor, it acts on multi-scale image feature maps. Each channel in the array, for example, through channel-by-channel multiplication:

[0116]

[0117] This yields an image representation of fusion speed and type features. .

[0118] This implementation method introduces a channel adjustment strategy based on running speed and food type, which realizes dynamic weighting at the image feature channel level. This allows the model to flexibly adjust the attention distribution and feature intensity when faced with different delivery speeds and food types (such as pies, smoked and braised foods), thereby improving the model's generalization recognition ability and adaptability to various defects.

[0119] In one feasible implementation, such as Figure 3 As shown, Figure 3 This is a flowchart illustrating the specific processing steps of a spatial channel attention enhancement module according to an embodiment of the present invention. The specific processing steps of the spatial channel attention enhancement module include:

[0120] S310. Perform weighted fusion of the above-mentioned multi-scale image feature map group F and the above-mentioned velocity type fusion feature map F' to generate a fused feature map. ;

[0121] S320, regarding the above-mentioned fusion feature map Perform channel attention modeling to obtain channel-enhanced feature maps. ;

[0122] S330, Enhance the feature map of the above channels Spatial attention waveforms are constructed to generate the enhanced feature maps described above. .

[0123] For example, the specific processing steps of the spatial channel attention enhancement module of the present invention are as follows. Its purpose is to combine multi-scale image features and velocity type fusion features, and through the attention modeling mechanism of channel dimension and spatial dimension, effectively improve the model's ability to perceive small defects and structural features on the surface of food.

[0124] First, multi-scale image feature maps Feature map fused with velocity type We perform weighted fusion along the channel dimension to obtain the fused feature map. The formula is as follows:

[0125]

[0126] in: To fuse the weight coefficients, they can be hyperparameters or obtained adaptively from the trained network; Multi-scale image feature maps extracted from images; : A fusion feature map generated by combining speed and food type; The fusion result retains image perception and speed adjustment information.

[0127] Using channel attention mechanism Modeling is performed. The commonly used Squeeze-and-Excitation (SE) structure is employed:

[0128] 1. Squeeze: for In spatial dimension Perform global average pooling to obtain the channel description vector:

[0129]

[0130] Let it be denoted as the channel description vector. .

[0131] 2. Excitation: Enhancement weights for each channel are generated using a two-layer fully connected network with an activation function.

[0132]

[0133] in: This is the weight matrix. This refers to the compression ratio; This represents the weighting coefficient for each channel; This is the Sigmoid function.

[0134] 3. Reweight: Reweights the channel weights. Applied to Each channel:

[0135]

[0136] Finally, the channel-enhanced feature map is obtained. .

[0137] Feature map after channel enhancement Subsequently, a spatial attention mechanism is introduced to enhance the response to key regions (such as defects or structural edges). A spatial attention structure similar to CBAM (Convolutional Block Attention Module) is employed:

[0138] 1. Channel aggregation: for Perform max pooling and average pooling, and then concatenate them along the channel dimension to obtain the spatial description graph:

[0139]

[0140] 2. Convolution modeling: ... Send in one Convolutional layers are followed by Sigmoid activation:

[0141]

[0142] 3. Spatial Weighting: The final spatial attention-enhanced feature map is calculated as follows:

[0143]

[0144] This embodiment employs a dual "channel-space" attention mechanism to achieve fine-tuning of the multi-scale image feature and speed / type embedding fusion results. Channel attention is used to select key semantic dimensions, while spatial attention enhances the response to local defects and shape edge regions, forming the final enhanced feature map. This provides more discriminative feature support for subsequent shape coding, defect identification, and surface judgment.

[0145] In one feasible implementation, such as Figure 4 As shown, Figure 4 This is a flowchart illustrating the specific processing steps of a shape encoding branch according to an embodiment of the present invention. The specific processing steps of the shape encoding branch include:

[0146] S410. The above-mentioned enhanced feature map Line channel compression and edge detection operations are performed to generate an edge map. ;

[0147] S420. Based on the center point coordinates of the target area mentioned above. The above edge map Perform a polar coordinate transformation to obtain a polar coordinate edge map. And extract the radial contour function of the corresponding edge based on each angular direction. ;

[0148] S430, Regarding the above radial contour function Perform a Discrete Fourier Transform to generate the Fourier descriptor of the current product. ;

[0149] S440, The above Fourier descriptor Compared with pre-stored standard template shape Fourier descriptors Perform frequency domain distance calculation to obtain the descriptor difference. ;

[0150] S450, Based on the above descriptor difference The shape consistency score is calculated using a nonlinear mapping function. :

[0151]

[0152] in, The preset rating sensitivity adjustment coefficient, This indicates the degree of consistency between the current food product's structure and shape and the standard template.

[0153] For example, the shape encoding branch described in this invention is used to extract the food structure edges from the enhanced feature map, and through Fourier description and polar coordinate transformation, it matches and evaluates the actual shape of the food with a standard template to obtain a shape consistency score. This allows for the automated quantitative assessment of the structural integrity of food products. The specific processing flow is as follows:

[0154] First, for the enhanced feature map Channel compression and edge detection are performed. Channel compression typically uses weighted averaging or max pooling to convert multi-channel feature maps into two-dimensional images. Then, Canny or Sobel operators are used to extract edge contours and generate edge maps. .

[0155] Extracting edge maps Coordinates of the center point of the target area Convert the coordinates of all edge points from Cartesian coordinates to polar coordinates to obtain the polar coordinate edge map. For each angle Extract the radius of the edge point farthest from the center point in that direction. Construct the radial profile function:

[0156]

[0157] This function This means representing the morphological changes of the food edge in polar coordinates, capturing the overall outline features.

[0158] radial profile function Perform a Discrete Fourier Transform (DFT) to obtain the frequency domain shape descriptor. Typically, only low-frequency coefficients are retained as principal shape components.

[0159]

[0160] Fourier coefficients contain the periodic variation characteristics of the main outline in the food morphology and can effectively characterize the shape outline.

[0161] The extracted Fourier shape descriptor of the current food Fourier descriptors of pre-stored standard template shapes Perform Euclidean distance matching and calculate shape difference:

[0162]

[0163] in: : A standard description representing the ideal shape of food; : The shape difference value between the current food and the standard template.

[0164] Based on shape difference value Shape consistency score is calculated using an exponential nonlinear mapping function. :

[0165]

[0166] in: The coefficient is used to adjust the sensitivity of the scoring; the larger the value, the more sensitive the scoring is to differences. A value close to 1 indicates that the current product shape is consistent with the template height, while a value close to 0 indicates a large shape deviation.

[0167] This embodiment projects the edges of the food structure into polar coordinate space and uses Fourier spectral coding to achieve compressed shape modeling. Then, it measures the difference between the model and a standard template. This not only effectively reduces the shape's sensitivity to scale, rotation, and minute perturbations, but also outputs a quantitative scoring result. It provides a reliable basis for the quality control of food appearance and structure, and is particularly suitable for the batch automatic inspection of foods with regular shapes such as pies and shortbread.

[0168] In one feasible implementation, such as Figure 5 As shown, Figure 5 This is a flowchart illustrating the specific processing steps of a defect detection branch according to an embodiment of the present invention. The specific processing steps of the defect detection branch include:

[0169] S510. The above-mentioned enhanced feature map Parallel feature regression and classification are performed using multiple convolutional prediction heads to generate a classification tensor. Regression Tensor and confidence tensor ;

[0170] S520, Based on the above regression tensor Map each candidate point to a corresponding set of defect bounding boxes. , where each bounding box Including center point coordinates and width and height parameters;

[0171] S530, Regarding the above classification tensor With confidence tensor The overall confidence score for each candidate box is calculated jointly. The optimized defect category set is obtained by performing non-maximum suppression based on the score size. and confidence score set ;

[0172] S540, based on confidence score set The spatial distribution generates a heatmap of defect saliency. .

[0173] For example, the defect detection branch of the present invention is used to enhance the feature map. This branch performs target detection to identify various defects on food surfaces and outputs corresponding bounding boxes, category labels, confidence scores, and defect heatmaps. It employs a multi-branch convolutional prediction structure, combined with non-maximum suppression and spatial distribution modeling mechanisms, to achieve accurate identification and visualization of defect regions. The specific processing steps are as follows:

[0174] Enhanced feature maps Input multiple parallel convolutional prediction heads to perform classification prediction, bounding box regression, and confidence prediction, respectively. Output the following three tensors: classification tensor. : Represents the probability of the defect category corresponding to each location. Number of defect categories; regression tensor : Represents the predicted bounding box coordinates for each location, including the center point. With width and height ; - Confidence tensor : Indicates the confidence score of whether the current position is the center of the defect.

[0175] Based on the regression tensor The output at each position, combined with the corresponding center point and offset, generates a set of defect bounding boxes:

[0176]

[0177] in, , representing the total number of all candidate boxes.

[0178] For each candidate bounding box By classifying tensors Classification probability at the corresponding position in the middle and confidence tensor Calculate the overall confidence level:

[0179]

[0180] All candidate boxes Based on overall confidence level The bounding boxes are sorted and non-maximum suppression is performed to remove redundant boxes with high overlap, retaining only the optimized detection results. Final output: a set of defect bounding boxes. Defect category set and defect confidence set .

[0181] Based on the spatial distribution of all preserved bounding boxes and their confidence scores, a heatmap of defect significance is generated. This allows for visualization of the spatial distribution of potential defective regions in current food images.

[0182] Heatmaps are typically calculated as follows:

[0183]

[0184] in: Center point for each reserved frame; The confidence score; Control the diffusion range of the heat map.

[0185] This embodiment uses a three-branch convolutional structure for end-to-end defect detection in food images and introduces classification confidence fusion and non-maximum suppression mechanisms, significantly improving the accuracy and robustness of defect detection; simultaneously, a defect heatmap... This provides an intuitive basis for manual review or subsequent control strategies. This branch is applicable to the automatic detection of minor surface defects in various types of food (such as broken crust, charred spots, peeling crust, etc.).

[0186] In one feasible implementation, such as Figure 6 As shown, Figure 6This is a flowchart illustrating the specific processing steps for a surface coverage determination branch according to an embodiment of the present invention. The specific processing steps for the surface coverage determination branch include:

[0187] S610. The above-mentioned enhanced feature map Extracting color anomaly response heatmap ;

[0188] S620, the above heat map The system is divided into multiple spatial sub-blocks. The mean and variance of the response are calculated for each sub-block, and the statistical values ​​of all sub-blocks are concatenated to form a spatial consistency descriptor vector. ;

[0189] S630, the above description vector Input a multi-class classifier and output a label indicating the surface coverage integrity level of the food. .

[0190] For example, in one feasible implementation, the surface coverage judgment branch of this invention aims to intelligently determine the integrity of the color or texture coverage of food surfaces, to assist in identifying problems such as whether the surface sauce in pie products is completely covered, and whether there are any omissions, gaps, or offsets. This branch uses enhanced feature maps... Using this as input, and through thermal mapping of abnormal areas, spatial consistency analysis, and statistical classification, the automatic assessment of food surface coverage level is achieved. The specific processing steps are as follows:

[0191] First, from the enhanced feature map Extract color channels, calculate color anomaly response map, and generate heatmap. This heatmap reflects the degree of color deviation of each pixel in the image, and is commonly estimated using the following methods:

[0192]

[0193] in: Point RGB or Lab color vectors; This is the standard color mean (such as the average color of the entire covered area).

[0194] Heatmap Divided into A spatial region (e.g.) (block), for each sub-block Calculate its average value with standard deviation Thus, a spatially consistent description vector is constructed:

[0195]

[0196] The difference between each block and the overall map mean can also be used as an indicator of local deviation.

[0197]

[0198] Finally, all statistics are concatenated into a feature vector. Or higher-dimensional fusion features.

[0199] The spatial consistency description vector constructed above The input is fed into a pre-trained multi-class classifier (such as random forest, multilayer perceptron, etc.), and the output is a level label for the integrity of the food surface coverage.

[0200]

[0201] This label can be used directly for displaying quality inspection results, downstream control logic, or alarm mechanisms.

[0202] In one feasible implementation, such as Figure 7 As shown, Figure 7 This invention provides a flowchart illustrating a method for structural modeling of the enhanced feature map using the shape encoding branch when the food type information belongs to the pie category. The method further includes: when the food type information... When the target food item is indicated to be a pie-type product, the enhanced feature map is added to the shape encoding branch mentioned above. Before performing structural modeling, the above methods include:

[0203] S710, regarding the above enhanced feature map Perform directional gradient mapping to obtain a multi-directional gradient map group;

[0204] S720. Based on the gradient map of each direction, calculate the local extreme response region and perform boundary tracking to extract the refined edge region;

[0205] S730. Input the above-mentioned refined edge region into the polar coordinate transformation and Fourier description process to improve the extraction accuracy and scoring stability of the shape boundary of pie products.

[0206] For example, when food type information If the current food item is determined to be a pie (such as red bean paste pastry, mooncake, or meat pie), the system will initiate a dedicated structural modeling process, performing shape contour enhancement based on orientation information to avoid edge extraction errors caused by directly relying on coarse contours. In the shape encoding branch, the enhanced feature map... Before structural modeling, multi-directional perception processing needs to be performed.

[0207] Enhanced feature maps Perform Sobel, Scharr, and other operator processing to extract gradient map sets in multiple directions:

[0208]

[0209] Indicates direction as gradient plot; Represents the directional differential operator; can form a group of directional gradient maps G. .

[0210] Gradient plot for each direction Perform nonmaximum suppression operation to extract local extrema regions and form a set of directional response boundaries:

[0211]

[0212] By merging and combining multi-directional boundaries, a refined edge region can be formed.

[0213]

[0214] This area offers more accurate positioning, making it particularly suitable for handling pie shapes with irregular edges and subtle color differences.

[0215] The above-mentioned refined edge region Projecting onto polar coordinate space, constructing polar coordinate boundary functions. Further Fourier transform is performed to obtain the frequency domain feature vector. This improves the stability of shape consistency assessment and reduces sensitivity to image rotation, scaling, or slight perturbations.

[0216] This embodiment introduces a food type trigger strategy, activating the enhanced modeling process only when the detected object is a pie-like food, balancing model efficiency and accuracy. Before encoding, directional gradient fusion, boundary refinement, and polar coordinate mapping are performed, effectively improving the recognition accuracy of circular or ring-like structural boundaries and enhancing edge stability extraction capabilities in low-contrast, wrinkle-interference scenarios. This structurally adaptive construction method makes this invention more suitable for food quality inspection tasks with complex physical forms, possessing significant industrial application value.

[0217] In one feasible implementation, as shown in the figure... Figure 8 This is a schematic diagram illustrating the specific training process of a food defect detection model provided in an embodiment of the present invention. The specific training process of the food defect detection model includes:

[0218] S810. Collect the training dataset. The training dataset includes multiple sample images, and the conveyor speed information and food type label corresponding to each training image constitute the triplet training input. ,in, For training images, For training speed values, To train food type labels;

[0219] S820, Regarding the above training images Image enhancement processing is performed, including at least one of random flipping, affine transformation, color perturbation, and blur noise, to expand the training dataset.

[0220] S830. Label the target output corresponding to each training image. The target output includes a set of defect bounding boxes. Defect category tag set Defect confidence score Shape consistency score and surface coverage rating label ;

[0221] S840, Transfer the above training images The above training speed values and the above-mentioned training food type labels The inputs are combined and fed into the aforementioned food defect detection model;

[0222] S850. Construct the total loss function, and perform multi-task joint training on the above food defect detection model:

[0223] S860. Update the above food defect detection model using a stochastic gradient descent optimizer to complete model training.

[0224] For example, to train the aforementioned food defect detection model, a multi-modal fusion training process combining image information, transport speed information, and food type information is proposed. This process fully utilizes the impact of speed and product type differences on the detection task, enhancing the model's robustness to food defect identification under different working conditions.

[0225] Collect training sample images and their associated information to construct triplet training data:

[0226]

[0227] in: For color training images; R is the speed of the conveyor line during image acquisition; Z represents the food type label (such as "pie", "sauce", etc.); ultimately forming the training set. .

[0228] For training images Random augmentations can be performed to improve the model's generalization ability. Augmentation operations include: random rotation; flipping; color jitter; blurring; or simulated noise injection.

[0229] Let the enhancement operation be A, then the enhanced image is represented as:

[0230]

[0231] For each training image, the detected targets are manually or semi-automatically labeled to generate label information, including: a set of defect bounding boxes. Each box Defect type set C Each Predefined defect categories (e.g., cracks, bubbles, burnt); defect confidence set Indicates model confidence; shape-consistency score Surface Coverage Rating Labels High, Medium, Low .

[0232] Image ,speed Food types The input model is synchronously processed, encoded separately to form feature embeddings, and then fused together for joint learning of visual and contextual factors.

[0233]

[0234] Design the joint loss function It combines multiple sub-tasks, such as:

[0235]

[0236] in: Bounding boxes and category classification loss; Shape consistency prediction loss (such as SmoothL1); : Cross-entropy loss for surface cover level prediction; This is the weighting factor.

[0237] Use a stochastic optimizer (such as Adam or SGD) to iteratively optimize the model parameters using a learning rate adjustment strategy:

[0238]

[0239] in: Parameters of the food defect detection model; : No. Round learning rate; : The gradient of the total loss function.

[0240] This embodiment significantly improves the model's ability to identify food defects under different operating conditions by constructing a training dataset that integrates visual images, conveying speed, and food type. Simultaneously, a multi-task loss function is designed to achieve comprehensive supervision of defect box localization, type discrimination, shape judgment, and surface quality. This method effectively improves the training efficiency and accuracy of online defect detection models for industrial food products, exhibiting broad adaptability and high stability.

[0241] The second aspect, such as Figure 9 As shown, Figure 9 This is a schematic diagram of a machine vision-based intelligent inspection system for conveyor lines, provided by an embodiment of the present invention. The present invention proposes a machine vision-based intelligent inspection system for conveyor lines, comprising:

[0242] The first acquisition unit 21 is used to acquire information on the operating speed of the conveyor line and information on the type of food.

[0243] The second acquisition unit 22 is used to acquire food image information of the target area of ​​the conveyor line;

[0244] The third acquisition unit 23 is used to input the above-mentioned running speed information, food type information and food image information into the food defect detection model to obtain food defect information. The food defect detection model integrates the above-mentioned food image information, running speed information and food type information, and combines shape encoding and attention enhancement to obtain multi-type food defect identification and scoring.

[0245] Understandably, a machine vision-based intelligent inspection system for conveyor lines can also perform the steps of any of the methods described in the first aspect.

[0246] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A machine vision-based intelligent inspection method for conveyor lines, characterized in that, include: Obtain information on the conveyor line's operating speed and the type of food; Acquire food image information of the target area of ​​the conveyor line; The running speed information, the food type information, and the food image information are input into the food defect detection model to obtain food defect information. The food defect detection model integrates the food image information, the running speed information, and the food type information, and combines shape encoding and attention enhancement to obtain multi-type food defect identification and scoring. The food defect detection model includes: A multimodal input fusion module is used to receive and fuse the food image information, the running speed information, and the food type information to form a multimodal input vector, wherein the multimodal input vector includes an image vector I and a speed vector I. and type vector ; The feature extraction backbone network is used to extract the multi-scale image feature map group F from the image vector I based on the built-in multi-scale convolutional neural network; The velocity-aware and type-adaptive embedding module is used to embed the velocity vector based on the multi-scale image feature map group F. and the type vector Generate a velocity type fusion feature map F'; The spatial channel attention enhancement module is used to generate enhanced feature maps from the multi-scale image feature map group F and the velocity type fusion feature map F'. ; Shape encoding branch, used based on the enhanced feature map Food structural morphology is modeled using Fourier edge description and polar coordinate transformation, and deviations from standard shapes are evaluated to obtain a shape consistency score. ; The defect detection branch is used for the enhanced feature map. Defect identification is performed to obtain defect box set B, category set C, confidence score set P, and heatmap. ; Surface coverage determination branch, used for the enhanced feature map The integrity of food surface coatings or sauces is assessed to obtain a coverage rating label. ; The specific processing steps of the speed sensing and type adaptation embedding module include: Based on the first multilayer sensor, according to the velocity vector Generate velocity embedding vector ; Based on the second multilayer perceptron according to the type vector Generate type embedding vectors ; Based on the velocity embedding vector and the embedding vector Generate control signals ; The control signal The velocity type fusion feature map F' is formed by applying it to each channel of the multi-scale image feature map group F. The specific processing steps of the spatial channel attention enhancement module include: The multi-scale image feature map group F and the velocity type fusion feature map F' are weighted and fused to generate a fused feature map. ; For the fused feature map Perform channel attention modeling to obtain channel-enhanced feature maps. ; Enhance the channel feature map Spatial attention waveforms are constructed to generate the enhanced feature map. ; The specific processing steps for the shape encoding branch include: For the enhanced feature map Line channel compression and edge detection operations are performed to generate an edge map. ; Based on the center point coordinates of the target area For the edge map Perform a polar coordinate transformation to obtain a polar coordinate edge map. And extract the radial contour function of the corresponding edge based on each angular direction. ; For the radial profile function Perform a Discrete Fourier Transform to generate the Fourier descriptor of the current product. ; The Fourier descriptor Compared with pre-stored standard template shape Fourier descriptors Perform frequency domain distance calculation to obtain the descriptor difference. ; Based on the descriptor difference The shape consistency score is calculated using a nonlinear mapping function. : in, The preset rating sensitivity adjustment coefficient, This indicates the degree of consistency between the current food product's structure and shape and the standard template.

2. The intelligent inspection method for conveyor lines based on machine vision according to claim 1, characterized in that, The specific processing steps of the defect detection branch include: For the enhanced feature map Parallel feature regression and classification are performed using multiple convolutional prediction heads to generate a classification tensor. Regression Tensor and confidence tensor ; Based on the regression tensor Map each candidate point to a corresponding set of defect bounding boxes. , where each bounding box Including center point coordinates and width and height parameters; For the classification tensor With confidence tensor The overall confidence score for each candidate box is calculated jointly. The optimized defect category set is obtained by performing non-maximum suppression based on the score size. and confidence score set ; Based on confidence score set The spatial distribution generates a heatmap of defect saliency. .

3. The intelligent inspection method for conveyor lines based on machine vision according to claim 1, characterized in that, The specific processing steps for determining the surface coverage branch include: For the enhanced feature map Extracting color anomaly response heatmap ; The heat map The system is divided into multiple spatial sub-blocks. The mean and variance of the response are calculated for each sub-block, and the statistical values ​​of all sub-blocks are concatenated to form a spatial consistency descriptor vector. ; The description vector Input a multi-class classifier and output a label indicating the surface coverage integrity level of the food. .

4. The intelligent inspection method for conveyor lines based on machine vision according to claim 1, characterized in that, The method further includes: When the food type information When indicating that the current target food product is a pie, The shape encoding branch pairs the enhanced feature map Before performing structural modeling, the method includes: For the enhanced feature map Perform directional gradient mapping to obtain a multi-directional gradient map group; Based on the gradient map of each orientation, the local extreme response region is calculated and the boundary is traced to extract the refined edge region; The refined edge region is input into the polar coordinate transformation and Fourier description process to improve the extraction accuracy and scoring stability of the shape boundary of pie products.

5. The intelligent inspection method for conveyor lines based on machine vision according to claim 1, characterized in that, The specific training process of the food defect detection model includes: A training dataset is collected, comprising multiple sample images, and the conveyor speed information and food type label corresponding to each training image, forming a triplet training input. ,in, For training images, For training speed values, To train food type labels; For the training images Image enhancement processing is performed, including at least one of random flipping, affine transformation, color perturbation, and blur noise, to expand the training dataset; Label the target output corresponding to each training image, the target output including a set of defect bounding boxes. Defect category label set Defect confidence score Shape consistency score and surface coverage rating label ; The training image The training speed value and the training food type label The inputs are combined and fed into the food defect detection model. Construct the total loss function, and perform multi-task joint training on the food defect detection model: The food defect detection model is updated using a stochastic gradient descent optimizer to complete model training.

6. A machine vision-based intelligent inspection device for conveyor lines, characterized in that, include: The first acquisition unit is used to acquire information on the operating speed of the conveyor line and the type of food. The second acquisition unit is used to acquire food image information of the target area of ​​the conveyor line; The third acquisition unit is used to input the running speed information, the food type information and the food image information into the food defect detection model to obtain food defect information. The food defect detection model integrates the food image information, the running speed information and the food type information, and combines shape encoding and attention enhancement to obtain multi-type food defect identification and scoring. The food defect detection model includes: A multimodal input fusion module is used to receive and fuse the food image information, the running speed information, and the food type information to form a multimodal input vector, wherein the multimodal input vector includes an image vector I and a speed vector I. and type vector ; The feature extraction backbone network is used to extract the multi-scale image feature map group F from the image vector I based on the built-in multi-scale convolutional neural network; The velocity-aware and type-adaptive embedding module is used to embed the velocity vector based on the multi-scale image feature map group F. and the type vector Generate a velocity type fusion feature map F'; The spatial channel attention enhancement module is used to generate enhanced feature maps from the multi-scale image feature map group F and the velocity type fusion feature map F'. ; Shape encoding branch, used based on the enhanced feature map Food structural morphology is modeled using Fourier edge description and polar coordinate transformation, and deviations from standard shapes are evaluated to obtain a shape consistency score. ; The defect detection branch is used for the enhanced feature map. Defect identification is performed to obtain defect box set B, category set C, confidence score set P, and heatmap. ; Surface coverage determination branch, used for the enhanced feature map The integrity of food surface coatings or sauces is assessed to obtain a coverage rating label. ; The specific processing steps of the speed sensing and type adaptation embedding module include: Based on the first multilayer sensor, according to the velocity vector Generate velocity embedding vector ; Based on the second multilayer perceptron according to the type vector Generate type embedding vectors ; Based on the velocity embedding vector and the embedding vector Generate control signals ; The control signal The velocity type fusion feature map F' is formed by applying it to each channel of the multi-scale image feature map group F. The specific processing steps of the spatial channel attention enhancement module include: The multi-scale image feature map group F and the velocity type fusion feature map F' are weighted and fused to generate a fused feature map. ; For the fused feature map Perform channel attention modeling to obtain channel-enhanced feature maps. ; Enhance the channel feature map Spatial attention waveforms are constructed to generate the enhanced feature map. ; The specific processing steps for the shape encoding branch include: For the enhanced feature map Line channel compression and edge detection operations are performed to generate an edge map. ; Based on the center point coordinates of the target area For the edge map Perform a polar coordinate transformation to obtain a polar coordinate edge map. And extract the radial contour function of the corresponding edge based on each angular direction. ; For the radial profile function Perform a Discrete Fourier Transform to generate the Fourier descriptor of the current product. ; The Fourier descriptor Compared with pre-stored standard template shape Fourier descriptors Perform frequency domain distance calculation to obtain the descriptor difference. ; Based on the descriptor difference The shape consistency score is calculated using a nonlinear mapping function. : in, The preset rating sensitivity adjustment coefficient, This indicates the degree of consistency between the current food product's structure and shape and the standard template.

Citation Information

Patent Citations

  • Method and device for detecting defects of toughened glass insulator

    CN103149215A

  • Quick-frozen gristle crispy meat food prepared from poultry bones and meat and production method of food

    CN104000222A