Conveying line intelligent detection method and system based on machine vision
By constructing a multimodal detection model, combining food images, speed and type information, and using a multi-scale convolutional neural network and attention enhancement module, the problem of defect recognition in food vision inspection systems under multiple food types and multiple speed conditions is solved, and efficient and accurate detection of food entities is achieved, which is particularly suitable for products with strict structural forms such as pies.
Patent Information
- Application Number
- CN202511212546.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-28
AI Technical Summary
Existing food visual inspection systems have difficulty achieving efficient and accurate defect recognition when faced with multiple types of food and multiple speeds, especially their ability to detect defects in the food itself is limited, and traditional methods fail to effectively cope with complex appearances and dynamic working condition changes.
A machine vision-based intelligent conveyor line inspection method is constructed. By integrating a multimodal detection model with food images, conveyor line speed, and food type information, a multi-scale convolutional neural network, a speed perception and type adaptation embedding module, a shape encoding branch, and a spatial channel attention enhancement module are used to achieve multi-task detection of food defects.
The system significantly improves the adaptability and accuracy of food conveyor line inspection systems, enabling the identification of a variety of common quality issues encountered in actual production, such as shape distortion, surface damage, and uneven sauce coverage. It provides more comprehensive quality assessment capabilities and is suitable for high-speed inspection of various food types.
Smart Images

Figure CN120741483A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of product detection, and more specifically, the present invention relates to a method and system for intelligent detection of conveyor lines based on machine vision. Background Art
[0002] With the continuous improvement of automation in the food industry, quality control methods based on visual inspection have gradually replaced traditional manual quality control methods, becoming a key technical approach to improving production efficiency and ensuring food quality. Currently, mainstream food visual inspection systems on the market focus on inspecting aspects such as packaging integrity, label printing quality, and barcode recognition accuracy, with their primary application concentrated in the final packaging stage of shipment. These methods rely on traditional image processing techniques such as rule templates, edge location, or image differentiation. While effective for standardized packaging inspection, they have limited ability to identify defects in the food itself (such as pies and smoked sauce products).
[0003] On the other hand, research on food defect detection is still in its infancy. In particular, real-time inspection of diverse food types (e.g., products of varying shapes, types, and speeds) on production lines still faces numerous technical challenges. For example, different foods vary significantly in appearance and structure, and defects manifest in a variety of forms (such as cracks, burnt surfaces, uneven coatings, and distorted shapes). Traditional models struggle to adapt to high-variability scenarios. Furthermore, variations in conveyor speed can lead to blurred images or misjudgments, and existing methods often fail to consider the impact of speed and food type on defect patterns.
[0004] It is necessary to propose a conveyor line intelligent detection method and system based on machine vision to at least solve some of the above problems. Summary of the Invention
[0005] The Summary of the Invention introduces a series of simplified concepts that will be further described in the Detailed Description of the Invention. The Summary of the Invention is not intended to limit the key features and essential features of the claimed technical solution, nor is it intended to determine the scope of protection of the claimed technical solution.
[0006] In a first aspect, the present invention proposes a conveyor line intelligent detection method based on machine vision, the method comprising: Obtain the running speed information and food type information of the conveyor line; Obtain food image information in the target area of the conveyor line; The running speed information, the food type information and the food image information are input into a food defect detection model to obtain food defect information, wherein the food defect detection model integrates the food image information, the running speed information and the food type information, and combines shape coding and attention enhancement to obtain multi-category food defect recognition and scoring.
[0007] In one feasible implementation, the food defect detection model includes: The multimodal input fusion module is used to receive and fuse the food image information, the running speed information and the food type information to form a multimodal input vector, wherein the multimodal input vector includes an image vector I, a speed vector and type vector ; A feature extraction backbone network is used to extract a multi-scale image feature map group F from the above image vector I based on a built-in multi-scale convolutional neural network; Speed perception and type adaptation embedding module is used to embed the speed perception and type adaptation module based on the multi-scale image feature map group F and the speed vector and the above type vector Generate speed type fusion feature map F'; Spatial channel attention enhancement module, used for the above multi-scale image feature map group F and the above speed type fusion feature map F' to generate enhanced feature map ; Shape encoding branch, used to enhance the feature map according to the above Modeling food structural morphology through Fourier edge description and polar coordinate transformation and evaluating deviations from standard shapes to obtain shape consistency scores ; Defect detection branch, used to enhance the feature map Perform defect identification to obtain defect frame set B, category set C, confidence score set P and heat map ; Surface coverage judgment branch, used to enhance the feature map Determine the integrity of food surface coatings or sauces to obtain coverage rating labels .
[0008] In a feasible implementation manner, the specific processing steps of the speed perception and type adaptation embedded module include: Based on the first multilayer perceptron according to the above speed vector Generate velocity embedding vector ; Based on the second multilayer perceptron according to the above type vector Generate type embedding vector ; Based on the above velocity embedding vector and the above embedding vector Generate control signals ; The above control signal Applied to each channel of the above multi-scale image feature map group F to form the above speed type fusion feature map F'.
[0009] In a feasible implementation, the specific processing steps of the above-mentioned spatial channel attention enhancement module include: The multi-scale image feature map group F and the speed type fusion feature map F' are weighted fused to generate a fusion feature map ; For the above fusion feature map Perform channel attention modeling to obtain channel enhanced feature maps ; The above channel enhanced feature map Perform spatial attention wavelet building to generate the above enhanced feature map .
[0010] In a feasible implementation, the specific processing steps of the shape coding branch include: For the above enhanced feature map Row channel compression and edge detection operations to generate edge maps ; Based on the center point coordinates of the above target area , the above edge graph Perform polar coordinate transformation to obtain polar coordinate edge map , and extract the radial profile function of the corresponding edge based on each angular direction ; For the above radial profile function Perform discrete Fourier transform to generate the Fourier descriptor of the current product ; The above Fourier descriptor Fourier descriptors with pre-stored standard template shapes Perform frequency domain distance calculation to obtain the descriptor difference ; Based on the above descriptor difference , the above shape consistency score is calculated by nonlinear mapping function :
[0011] in, is the preset scoring sensitivity adjustment coefficient, Indicates the degree of consistency between the structural shape of the current food product and the standard template.
[0012] In a feasible implementation manner, the specific processing steps of the above-mentioned defect detection branch include: For the above enhanced feature map Perform parallel feature regression and classification through multiple convolutional prediction heads to generate classification tensors , regression tensor and the confidence tensor ; Based on the above regression tensor Map each candidate point to a corresponding set of defect bounding boxes , where each bounding box Including center point coordinates and width and height parameters; For the above classification tensor and the confidence tensor Jointly calculate the comprehensive confidence score of each candidate box , and perform non-maximum suppression processing according to the score size to obtain the optimized defect category set and confidence score set ; Based on confidence score set Distribution in spatial location to generate defect saliency heatmap .
[0013] In a feasible implementation manner, the specific processing steps of the surface coverage judgment branch include: For the above enhanced feature map Extracting color anomaly response heatmap ; The above heat map Divide the spatial region into multiple sub-blocks, calculate the response mean and variance of each sub-block respectively, and concatenate the statistical values of all sub-blocks into a spatial consistency description vector ; The above description vector Input a multi-class classifier and output a label of the coverage integrity level of the food surface .
[0014] In a feasible implementation manner, the above method further includes: When the above food type information When indicating that the current target food belongs to the pie category, In the above shape encoding branch, the feature map is enhanced Before structural modeling, the above method includes: For the above enhanced feature map Performing a directional gradient mapping operation to obtain a multi-directional gradient map group; Based on each directional gradient map, the local extreme response area is calculated and the boundary is traced to extract the refined edge area; The refined edge area is input into the polar coordinate transformation and Fourier description process to improve the extraction accuracy and scoring stability of the shape boundary of pie products.
[0015] In a feasible implementation, the specific training process of the above-mentioned food defect detection model includes: Collect training data sets, which include multiple sample images, conveyor line speed information and food type labels corresponding to each training image, forming a triplet training input ,in, For training images, is the training speed value, To train food type labels; For the above training images Performing image enhancement processing, wherein the image enhancement processing includes at least one of random flipping, affine transformation, color perturbation, and blur noise, to expand the training data set; Label the target output corresponding to each training image, which includes a set of defect bounding boxes , defect category label set , defect confidence score , shape consistency score and surface coverage grade labels ; The above training images , the above training speed value and the above training food type labels are jointly input into the above-mentioned food defect detection model; Constructing the total loss function, the above food defect detection model performs multi-task joint training: The above food defect detection model is updated through the stochastic gradient descent optimizer to complete the model training.
[0016] In a second aspect, the present invention proposes a conveyor line intelligent detection system based on machine vision, comprising: A first acquiring unit is used to acquire the running speed information of the conveyor line and the food type information; A second acquisition unit is used to acquire food image information of a target area of the conveyor line; The third acquisition unit is used to input the above-mentioned running speed information, the above-mentioned food type information and the above-mentioned food image information into the food defect detection model to obtain food defect information, wherein the above-mentioned food defect detection model integrates the above-mentioned food image information, the above-mentioned running speed information and the above-mentioned food type information, combines shape coding and attention enhancement, to obtain multi-category food defect recognition and scoring.
[0017] In summary, the present invention breaks through the limitation of traditional visual systems that only focus on packaging defects, and directly performs quality inspection on the food itself. It can identify a variety of common quality problems in actual production, such as shape distortion, surface damage, uneven sauce coverage, etc., significantly broadening the application scope of visual quality inspection and improving the granularity and coverage of product quality control. Secondly, this method effectively alleviates the problem that existing models are prone to misjudgment when facing multiple types and multiple speeds of food by constructing a multimodal detection model that integrates food images, conveyor line speed and food type information. Among them, the speed perception and type adaptation module can adaptively adjust the feature extraction process according to the actual production speed and food category, and enhance the robustness of the model to dynamic working condition changes. This method innovatively introduces a shape coding branch, and accurately depicts the boundaries and contours of food structure through Fourier description and polar coordinate modeling. It can accurately evaluate the shape consistency between the target product and the standard template. It is particularly suitable for products with strict structural morphology requirements such as pies and biscuits, and can achieve more refined structural evaluation and scoring. In addition, the proposed spatial-channel attention enhancement module fuses image features with speed / type control signals, and through saliency modeling in the channel and spatial dimensions, further strengthens the response weight of the defect area, improving the model's ability to perceive subtle defects in complex backgrounds. Finally, the system supports multi-task branch collaborative detection, including defect detection, shape scoring, and surface coverage judgment. Compared with traditional single-dimension detection solutions, it has more comprehensive quality assessment capabilities and can provide food production companies with more reliable quality data support and decision-making basis. In summary, this method significantly improves the adaptability, accuracy, and practicality of the intelligent detection system for food conveyor lines by integrating multimodal perception, structural modeling, and attention enhancement technologies.
[0018] The present invention proposes an intelligent detection method for conveyor lines based on machine vision. Other advantages, objectives and features of the present invention will be reflected in part through the following description, and in part will be understood by those skilled in the art through research and practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present description. The same reference symbols are used throughout the drawings to represent the same components. In the drawings: Figure 1 A schematic diagram of the process of a conveyor line intelligent detection method based on machine vision provided by an embodiment of the present invention; Figure 2 A schematic diagram illustrating the specific processing steps of a speed perception and type adaptation embedded module provided by an embodiment of the present invention; Figure 3 A schematic diagram illustrating the specific processing steps of a spatial channel attention enhancement module provided by an embodiment of the present invention; Figure 4 A schematic flow chart of specific processing steps of a shape coding branch provided by an embodiment of the present invention; Figure 5 A schematic diagram of the specific processing steps of a defect detection branch provided by an embodiment of the present invention; Figure 6 A schematic flow chart of specific processing steps for a surface coverage judgment branch provided by an embodiment of the present invention; Figure 7 A schematic diagram of a process for a shape coding branch to perform structural modeling on an enhanced feature map when food type information belongs to a pie product provided in an embodiment of the present invention; Figure 8 A schematic diagram illustrating a specific training process of a food defect detection model provided by an embodiment of the present invention; Figure 9 A schematic structural diagram of a conveyor line intelligent detection system based on machine vision provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0020] The terms "first", "second", "third", "fourth", etc. (if any) in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments.
[0021] See also Figure 1, which is a process diagram of a conveyor line intelligent detection method based on machine vision provided by an embodiment of the present invention, which may specifically include: S110, obtaining the running speed information of the conveyor line and the food type information; S120, obtaining food image information of a target area of the conveyor line; S130. Input the running speed information, the food type information, and the food image information into a food defect detection model to obtain food defect information. The food defect detection model integrates the food image information, the running speed information, and the food type information, and combines shape coding and attention enhancement to obtain multi-category food defect recognition and scoring.
[0022] For example, this embodiment provides a method for intelligent detection of conveyor lines based on machine vision, which specifically includes the following steps: In step S110, the system first obtains information about the conveyor line's operating speed and the type of food being conveyed. This information is read by encoders, sensors, or control systems on the conveyor line and, combined with the food identification database, labels the food's category. For example, if a pie is being conveyed, the system will label it as "pie." If a smoked sauce is being conveyed, it will be labeled "smoked sauce."
[0023] Next, in step S120, industrial cameras or image acquisition devices installed in the target detection area of the conveyor line acquire food image information in real time. The captured images cover key areas of the food on the conveyor line, ensuring that visual information such as its surface condition, contour shape, and appearance structure can be observed.
[0024] In step S130, the acquired speed information, food type information, and image information are input into a constructed food defect detection model. This model utilizes a multimodal fusion architecture, capable of jointly processing multiple sources of information. Specifically, the model first transforms the image, speed, and type information into a unified feature vector representation via a multimodal input fusion module. A multiscale convolutional network then extracts image features and performs attention adjustment on the embedded vectors of speed and type to produce a more discriminative fused feature map. Subsequently, the shape encoding module uses Fourier descriptors and polar coordinate transformations to model the food structure and identify deviations from the standard form. Simultaneously, the defect detection module locates possible surface anomalies such as cracks, damage, and contamination, generating a defect category and confidence score. Furthermore, the surface coverage assessment module assesses the integrity of sauce coatings or spreads. Ultimately, the system outputs comprehensive results, including defect location, category, severity score, shape consistency score, and surface coverage level, supporting food production quality monitoring and automated rejection.
[0025] In summary, the present invention breaks through the limitation of traditional visual systems that only focus on packaging defects, and directly performs quality inspection on the food itself. It can identify a variety of common quality problems in actual production, such as shape distortion, surface damage, uneven sauce coverage, etc., significantly broadening the application scope of visual quality inspection and improving the granularity and coverage of product quality control. Secondly, this method effectively alleviates the problem that existing models are prone to misjudgment when facing multiple types and multiple speeds of food by constructing a multimodal detection model that integrates food images, conveyor line speed and food type information. Among them, the speed perception and type adaptation module can adaptively adjust the feature extraction process according to the actual production speed and food category, and enhance the robustness of the model to dynamic working condition changes. This method innovatively introduces a shape coding branch, and accurately depicts the boundaries and contours of food structure through Fourier description and polar coordinate modeling. It can accurately evaluate the shape consistency between the target product and the standard template. It is particularly suitable for products with strict structural morphology requirements such as pies and biscuits, and can achieve more refined structural evaluation and scoring. In addition, the proposed spatial-channel attention enhancement module fuses image features with speed / type control signals, and through saliency modeling in the channel and spatial dimensions, further strengthens the response weight of the defect area, improving the model's ability to perceive subtle defects in complex backgrounds. Finally, the system supports multi-task branch collaborative detection, including defect detection, shape scoring, and surface coverage judgment. Compared with traditional single-dimension detection solutions, it has more comprehensive quality assessment capabilities and can provide food production companies with more reliable quality data support and decision-making basis. In summary, this method significantly improves the adaptability, accuracy, and practicality of the intelligent detection system for food conveyor lines by integrating multimodal perception, structural modeling, and attention enhancement technologies.
[0026] In one feasible implementation, the food defect detection model includes: The multimodal input fusion module is used to receive and fuse the food image information, the running speed information and the food type information to form a multimodal input vector, wherein the multimodal input vector includes an image vector I, a speed vector and type vector ; A feature extraction backbone network is used to extract a multi-scale image feature map group F from the above image vector I based on a built-in multi-scale convolutional neural network; Speed perception and type adaptation embedding module is used to embed the speed perception and type adaptation module based on the multi-scale image feature map group F and the speed vector and the above type vector Generate speed type fusion feature map F'; Spatial channel attention enhancement module, used for the above multi-scale image feature map group F and the above speed type fusion feature map F' to generate enhanced feature map ; Shape encoding branch, used to enhance the feature map according to the above Modeling food structural morphology through Fourier edge description and polar coordinate transformation and evaluating deviations from standard shapes to obtain shape consistency scores ; Defect detection branch, used to enhance the feature map Perform defect identification to obtain defect frame set B, category set C, confidence score set P and heat map ; Surface coverage judgment branch, used to enhance the feature map Determine the integrity of food surface coatings or sauces to obtain coverage rating labels .
[0027] Illustratively, the food defect detection model provided by the present invention integrates image features, conveyor line running speed information and food type information. Through multi-module collaborative modeling, it improves the defect recognition capability and adaptability of different types of food during high-speed transportation.
[0028] Specifically, the multi-modal input fusion module first receives three types of input information, namely, food image vectors , velocity vector and type vector , constructing a multimodal input representation. In this way, the model not only perceives image information but also introduces speed and category context related to production conditions, providing a basis for subsequent feature adjustment.
[0029] Next, the model uses a multi-scale convolutional neural network to transform the image vector Perform feature extraction to obtain a multi-scale feature map group of the image , to enhance the model's ability to identify defects of different granularity.
[0030] Then, the speed perception and type adaptation module embeds the speed vector With type vector Perform embedding mapping and generate control signals to adjust the feature map group , generating an intermediate representation that fuses speed and type features , in order to improve the model's adaptability to operating conditions and product type differences.
[0031] In order to further enhance the feature response of the key area, the spatial channel attention enhancement module models the attention weight based on the image feature and speed type fusion feature to generate an enhanced feature map , significantly improving the model's attention to potential defect locations.
[0032] Enhanced feature map The data is sent to three branch processing modules. The first one is the shape coding branch, which extracts edge information through median filtering, combines polar coordinate transformation with Fourier descriptor technology to model the structural outline of the food, and compares it with the standard shape template to finally generate a shape consistency score. , which is used to measure the structural integrity of food.
[0033] The second is the defect detection branch, which uses multiple convolutional prediction heads to enhance feature maps Perform regression and classification operations to output the location frame set of food defects , category set , confidence score set and spatial distribution heat map , to achieve accurate positioning and qualitative analysis of food defects.
[0034] The third branch is the surface coverage judgment branch, which focuses on analyzing the surface distribution status of sauces, coatings, etc., combining the color response heat map to count the distribution characteristics of multiple spatial sub-blocks, and finally determining the surface coverage integrity level label through the classifier. , suitable for judging whether there is uneven coverage, missing coating, etc.
[0035] In summary, this implementation method achieves a deep fusion of three types of features: speed-type adjustment, structural modeling, and defect thermal visualization. It has high robustness, high precision, and multi-task recognition capabilities, and is particularly suitable for industrial scenarios with strict requirements on food appearance quality and frequent production line changes.
[0036] In one possible implementation, Figure 2 As shown, Figure 2 A schematic flow chart of specific processing steps of a speed perception and type adaptation embedded module provided in an embodiment of the present invention includes the following specific processing steps: S210, based on the first multilayer perceptron according to the above speed vector Generate velocity embedding vector ; S220, based on the second multilayer perceptron according to the above type vector Generate type embedding vector ; S230, based on the above speed embedding vector and the above embedding vector Generate control signals ; S240, the above control signal Applied to each channel of the above multi-scale image feature map group F to form the above speed type fusion feature map F'.
[0037] Exemplarily, the speed perception and type adaptation embedded module proposed in the present invention is used to convert the operating speed information of the conveyor line and the food type information into a control signal that can adjust the image feature extraction process, thereby enabling the food defect detection model to have stronger adaptability and discrimination capabilities when processing different operating states and different food categories.
[0038] Specifically, the system first passes through a first multi-layer perceptron (MLP ), the input velocity vector Mapped into a velocity embedding vector The process can be expressed as:
[0039] in: : represents the original velocity feature vector; : the weight matrix of the first perceptron; : bias term; : non-linear activation function, such as ReLU or SiLU; : Velocity embedding representation, dimension is .
[0040] In parallel, the system uses a second multilayer perceptron ( ), the food type vector Mapping to type embedding vector , the specific formula is:
[0041] in: : original food type vector; : the weight matrix of the second perceptron; : bias term; : Type embedding vector, dimension is .
[0042] Then, the two embedding vectors are fused to generate the control signal vector , used to control image channel features. Fusion methods can use weighted summation, linear transformation after splicing, Hadamard product, etc. The following is an implementation example of splicing + linear projection:
[0043] in: : represents the concatenation of two embedding vectors; : Linear transformation weight of the control signal, is the number of channels of the image feature map; : bias term; : Control signal vector, used for adjustment of each channel.
[0044] Finally, the control signal As a regulating factor in the channel dimension, it acts on the multi-scale image feature map group Each channel in , for example by channel-by-channel multiplication:
[0045] Thus, an image representation of fusion speed and type features is obtained .
[0046] This implementation method achieves dynamic weighting of image features at the channel level by introducing a channel adjustment strategy based on running speed and food type. This enables the model to flexibly adjust attention distribution and feature strength when faced with different conveying speeds and food types (such as pies, smoked sauces, etc.), thereby improving the model's generalized recognition ability and adaptability to various defects.
[0047] In one possible implementation, Figure 3 As shown, Figure 3 This is a flow chart of specific processing steps of a spatial channel attention enhancement module provided by an embodiment of the present invention. The specific processing steps of the spatial channel attention enhancement module include: S310, performing weighted fusion on the multi-scale image feature map group F and the speed type fusion feature map F' to generate a fusion feature map ; S320, the above fusion feature map Perform channel attention modeling to obtain channel enhanced feature maps ; S330, the channel enhancement feature map Perform spatial attention wavelet building to generate the above enhanced feature map .
[0048] Exemplarily, the specific processing steps of the spatial channel attention enhancement module of the present invention are as follows. Its purpose is to combine multi-scale image features and speed type fusion features, and through the attention modeling mechanism of channel dimension and spatial dimension, effectively improve the model's perception of tiny defects and structural features on the food surface.
[0049] First, the multi-scale image feature map group Fusion feature map with speed type Perform weighted fusion on the channel dimension to obtain the fusion feature map , the formula is as follows:
[0050] in: is the fusion weight coefficient, which can be a hyperparameter or adaptively learned from the training network; : A set of multi-scale image feature maps extracted from an image; : Fusion feature map generated by combining speed and food type; : Fusion results, retaining image perception and speed regulation information.
[0051] Using channel attention mechanism Modeling is performed using the commonly used SE (Squeeze-and-Excitation) structure: 1. Squeeze: Yes In the spatial dimension Perform global average pooling on it to obtain the channel description vector:
[0052] Denoted as channel description vector .
[0053] 2. Excitation: Generate enhanced weights for each channel through a two-layer fully connected network with an activation function:
[0054] in: is the weight matrix, is the compression ratio; Represents the weight coefficient of each channel; is the Sigmoid function.
[0055] 3. Reweight: Reweight the channel Applicable to Each channel:
[0056] Finally, the channel enhanced feature map is obtained .
[0057] After obtaining the channel enhanced feature map After that, the spatial attention mechanism is further introduced to strengthen the response to key areas (such as defects or structural edges). A spatial attention structure similar to CBAM (Convolutional Block Attention Module) is adopted: 1. Channel aggregation: Perform maximum pooling and average pooling, and splice in the channel dimension to obtain a spatial description map:
[0058] 2. Convolutional modeling: Send in one After the convolution layer, Sigmoid activation is performed:
[0059] 3. Spatial Weighting: The final spatial attention enhanced feature map is calculated as follows:
[0060] This embodiment uses the "channel-space" dual attention mechanism to achieve fine-tuning of the fusion results of multi-scale image features and speed / type embedding. Among them, channel attention is used to select key semantic dimensions, and spatial attention strengthens the response to local defects and shape edge areas to form the final enhanced feature map. , providing more discriminative feature support for subsequent shape coding, defect recognition and surface judgment.
[0061] In one possible implementation, Figure 4 As shown, Figure 4 This is a flow chart of specific processing steps of a shape coding branch provided by an embodiment of the present invention. The specific processing steps of the shape coding branch include: S410, enhancing the feature map Row channel compression and edge detection operations to generate edge maps ; S420, based on the center point coordinates of the target area , the above edge graph Perform polar coordinate transformation to obtain polar coordinate edge map , and extract the radial profile function of the corresponding edge based on each angular direction ; S430, the radial profile function Perform discrete Fourier transform to generate the Fourier descriptor of the current product ; S440, the above Fourier descriptor Fourier descriptors with pre-stored standard template shapes Perform frequency domain distance calculation to obtain the descriptor difference ; S450, based on the above descriptor difference , the above shape consistency score is calculated by nonlinear mapping function :
[0062] in, is the preset scoring sensitivity adjustment coefficient, Indicates the degree of consistency between the structural shape of the current food product and the standard template.
[0063] Exemplarily, the shape coding branch of the present invention is used to extract the edges of the food structure from the enhanced feature map, and to match and evaluate the actual shape of the food with the standard template through Fourier description and polar coordinate transformation to obtain a shape consistency score. , in order to achieve automatic quantitative evaluation of the integrity of food appearance and structure. The specific processing flow is as follows: First, the enhanced feature map Perform channel compression and edge detection operations. Channel compression usually uses weighted average or maximum pooling to convert multi-channel feature maps into two-dimensional images, and then uses Canny or Sobel operators to extract edge contours to generate edge maps. .
[0064] Extract edge map Coordinates of the center point of the target area , convert all edge point coordinates from Cartesian coordinate system to polar coordinate form to obtain polar coordinate edge map For each angle , extract the radius of the edge point farthest from the center point in this direction , construct the radial profile function:
[0065] This function That is, it represents the morphological changes of the edge of the food in polar coordinates and captures the overall contour features.
[0066] The radial profile function Perform a discrete Fourier transform (DFT) to obtain a frequency domain shape descriptor , usually only low-frequency coefficients are retained as shape principal components:
[0067] The Fourier coefficients contain the periodic variation characteristics of the main contour of the food morphology and can effectively characterize the shape contour.
[0068] The extracted Fourier shape descriptor of the current food Fourier descriptor with pre-stored standard template shape Perform Euclidean distance matching and calculate shape differences:
[0069] in: : A standard leaflet describing the ideal food appearance; : The shape difference value between the current food and the standard template.
[0070] Based on shape difference value , using exponential nonlinear mapping function to calculate shape consistency score :
[0071] in: This is a coefficient that adjusts the sensitivity of the score; the larger the value, the more sensitive the score is to differences; A value close to 1 indicates that the current product shape is highly consistent with the template, while a value close to 0 indicates a large shape deviation.
[0072] This embodiment projects the edges of food structures into polar coordinate space, uses Fourier spectrum coding to achieve shape compression modeling, and then measures the difference with the standard template. This not only effectively reduces the sensitivity of the shape to scale, rotation and small disturbances, but also outputs quantitative scoring results. , providing a reliable basis for the quality control of food appearance and structure, and is particularly suitable for batch automatic inspection of regular-shaped foods such as pies and pastries.
[0073] In one possible implementation, Figure 5 As shown, Figure 5 A schematic flow diagram of specific processing steps of a defect detection branch provided in an embodiment of the present invention, wherein the specific processing steps of the defect detection branch include: S510, enhancing the feature map Perform parallel feature regression and classification through multiple convolutional prediction heads to generate classification tensors , regression tensor and the confidence tensor ; S520, based on the above regression tensor Map each candidate point to a corresponding set of defect bounding boxes , where each bounding box Including center point coordinates and width and height parameters; S530, the above classification tensor and the confidence tensor Jointly calculate the comprehensive confidence score of each candidate box , and perform non-maximum suppression processing according to the score size to obtain the optimized defect category set and confidence score set ; S540, based on confidence score set Distribution in spatial location to generate defect saliency heatmap .
[0074] Exemplarily, the defect detection branch of the present invention is used to enhance the feature map Performs object detection tasks to identify various defects on food surfaces and outputs corresponding bounding boxes, category labels, confidence scores, and defect heatmaps. This branch uses a multi-branch convolutional prediction structure, combined with non-maximum suppression and spatial distribution modeling mechanisms, to achieve accurate identification and visualization of defect areas. The specific processing steps are as follows: Enhanced feature map Input multiple parallel convolution prediction heads to perform classification prediction, bounding box regression and confidence prediction respectively. Output the following three tensors: classification tensor : represents the probability of defect category corresponding to each position, is the number of defect categories; regression tensor : Represents the bounding box coordinates of each location prediction, including the center point and width and height ; - Confidence tensor : Indicates the confidence score of whether the current position is the defect center.
[0075] According to the regression tensor The output of each position in is combined with the corresponding center point and offset to generate a set of defect bounding boxes:
[0076] in, , represents the total number of all candidate boxes.
[0077] For each candidate bounding box , by classifying the tensor The classification probability of the corresponding position in , and the confidence tensor , calculate the comprehensive confidence:
[0078] All candidate boxes Based on comprehensive confidence Sort and perform non-maximum suppression to remove redundant boxes with large overlap rates, and only retain the optimized detection results. Final output: defect bounding box set , defect category set and defect confidence set .
[0079] Generate a defect saliency heatmap based on the spatial distribution of all retained bounding boxes and their confidence scores , to visualize the spatial distribution of potential defect areas in the current food image.
[0080] The heat map calculation method is usually:
[0081] in: For each reserved box center point; is the confidence score; Control the diffusion range of the heat map.
[0082] This embodiment uses a three-branch convolutional structure to perform end-to-end detection of defects in food images, and introduces classification confidence fusion and non-maximum suppression mechanisms, which significantly improves the accuracy and robustness of defect detection; at the same time, the defect heat map This provides an intuitive basis for manual review or subsequent control strategies. This branch is suitable for the automatic detection of minor surface defects (such as broken puff pastry, burnt spots, and skin peeling) on various types of food.
[0083] In one possible implementation, Figure 6 As shown, Figure 6 A schematic flow chart of specific processing steps of a surface coverage determination branch provided in an embodiment of the present invention, wherein the specific processing steps of the surface coverage determination branch include: S610: Enhance the feature map Extracting color anomaly response heatmap ; S620, the above heat map Divide the spatial region into multiple sub-blocks, calculate the response mean and variance of each sub-block respectively, and concatenate the statistical values of all sub-blocks into a spatial consistency description vector ; S630, the above description vector Input a multi-class classifier and output a label of the coverage integrity level of the food surface .
[0084] For example, in a feasible embodiment, the surface coverage judgment branch of the present invention is intended to intelligently judge the color or texture coverage integrity of the food surface, so as to assist in identifying whether the surface sauce of a pie product is completely covered, whether there is any missing coating, empty coating or offset. As input, through abnormal area thermal mapping, spatial consistency analysis and statistical classification, the automatic assessment of food surface coverage level is achieved. The specific processing steps are as follows: First, from the enhanced feature map Extract color channels, calculate color anomaly response map, and generate heat map The heat map reflects the color deviation of each pixel in the image, and is usually estimated as follows:
[0085] in: Indicates a point RGB or Lab color vector; is the standard color mean (e.g., the average color of a fully covered area).
[0086] The heat map Divided into spatial regions (e.g. block), for each sub-block Calculate the average and standard deviation , thereby constructing a spatial consistency description vector:
[0087] You can also add the difference between each block and the whole image mean as a local deviation indicator:
[0088] Finally, all statistics are concatenated into feature vectors Or higher-dimensional fusion features.
[0089] The spatial consistency description vector constructed above Input into a trained multi-class classifier (such as random forest, multi-layer perceptron, etc.) and output the grade label of the food surface coverage integrity:
[0090] This tag can be directly used for quality inspection result display, downstream control logic or alarm mechanism.
[0091] In one possible implementation, Figure 7 As shown, Figure 7 The embodiment of the present invention provides a flow chart of a method for the shape coding branch to perform structural modeling on the enhanced feature map when the food type information belongs to a pie product. The method further includes: when the food type information When the current target food is a pie product, the shape encoding branch enhances the feature map Before structural modeling, the above method includes: S710, enhancing the feature map Performing a directional gradient mapping operation to obtain a multi-directional gradient map group; S720: Based on each directional gradient map, calculate the local extreme response area and perform boundary tracking to extract the refined edge area; S730: Input the refined edge region into the polar coordinate transformation and Fourier description process to improve the extraction accuracy and scoring stability of the shape boundary of the pie product.
[0092] For example, when food type information If the current food is determined to be a pie (such as red bean paste cake, moon cake, fresh meat pie, etc.), the system will start a dedicated structure modeling process to perform shape contour enhancement based on directional information to avoid edge extraction errors caused by direct reliance on rough contours. In the shape encoding branch, the enhanced feature map Before structural modeling, multi-directional perception processing needs to be performed.
[0093] Enhanced feature map Perform Sobel, Scharr and other operator processing to extract gradient map sets in multiple directions:
[0094] Indicates the direction Gradient map of Represents the directional differential operator; it can form the directional gradient graph group G .
[0095] Gradient map for each direction Perform non-maximum suppression operations to extract local extreme value regions and form a set of directional response boundaries:
[0096] Combine the multi-directional boundaries to form a refined edge region:
[0097] This area is more accurately positioned and is especially suitable for processing pie shapes with irregular edges and subtle color differences.
[0098] The above-mentioned refined edge area Project to polar coordinate space and construct polar coordinate edge function , further perform Fourier transform to obtain frequency domain feature vector , to improve the stability of shape consistency evaluation and reduce its sensitivity to image rotation, scaling or slight perturbations.
[0099] This embodiment introduces a food type trigger strategy, activating the enhanced modeling process only when the detection object is a pie-like food, taking into account both model efficiency and accuracy. Directional gradient fusion, boundary refinement, and polar coordinate mapping are performed before encoding, effectively improving the recognition accuracy of circular or ring-like structure boundaries and enhancing the ability to extract stable edges in low-contrast and wrinkle-interference scenarios. This structurally adaptive construction method makes the present invention more suitable for food quality inspection tasks with complex physical forms, and has significant industrial promotion value.
[0100] In a feasible embodiment, as shown in the figure, Figure 8 The following is a flow chart illustrating a specific training process of a food defect detection model provided in an embodiment of the present invention. The specific training process of the food defect detection model includes: S810, collect training data sets, the training data sets include multiple sample images, each training image corresponding to the conveyor line running speed information and food type label, constitute a triplet training input ,in, For training images, is the training speed value, To train food type labels; S820, the above training images Performing image enhancement processing, wherein the image enhancement processing includes at least one of random flipping, affine transformation, color perturbation, and blur noise, to expand the training data set; S830: Label the target output corresponding to each training image. The target output includes a set of defect bounding boxes. , defect category label set , defect confidence score , shape consistency score and surface coverage grade labels ; S840, the above training images , the above training speed value and the above training food type labels are jointly input into the above-mentioned food defect detection model; S850: Construct a total loss function and perform multi-task joint training on the above-mentioned food defect detection model: S860: Update the food defect detection model through a stochastic gradient descent optimizer to complete model training.
[0101] For example, to train the aforementioned food defect detection model, a multimodal fusion training process was proposed that combines image information, conveyor speed information, and food type information. This process leverages the impact of speed and product type differences on the detection task, enhancing the model's robustness in identifying food defects under diverse conditions.
[0102] Collect training sample images and their supporting information to construct triplet training data:
[0103] in: is a color training image; R is the conveyor line running speed during image acquisition; Z is the food type label (such as "pie", "sauce", etc.); the training set is finally formed .
[0104] For training images Perform random enhancement to improve the generalization ability of the model. Enhancement operations include: random rotation (Rotation); mirror flip (Flip); color perturbation (Color Jitter); blur (Blur) or simulated noise injection (Noise), etc.
[0105] Assume that the enhancement operation is A, then the enhanced image is expressed as:
[0106] Manually or semi-automatically mark the detection target for each training image and generate label information, including: defect bounding box set , each box ; Defect type set C , each Predefined defect categories (e.g., cracks, bubbles, burns); defect confidence level sets , indicating model confidence; shape-consistency score ; Surface coverage score label High, Medium, Low .
[0107] The image ,speed , food type Synchronize the input model, encode each feature into embeddings, and then fuse them together to jointly learn visual and contextual factors:
[0108] Designing a joint loss function , integrating multiple subtasks, such as:
[0109] in: : Bounding box and category classification loss; : shape consistency prediction loss (such as SmoothL1); : Surface coverage level prediction cross entropy loss; is the weight factor.
[0110] Use a stochastic optimizer (such as Adam or SGD) to iteratively optimize the model parameters using a learning rate adjustment strategy:
[0111] in: : Food defect detection model parameters; : No. Round learning rate; : Gradient of the total loss function.
[0112] This example significantly improves the model's ability to identify food defects under different working conditions by constructing a training dataset that integrates visual images, conveyor speed, and food type. A multi-task loss function is also designed to comprehensively monitor defect box location, type discrimination, shape judgment, and surface quality. This approach effectively improves the training efficiency and accuracy of industrial food online defect detection models, demonstrating broad adaptability and high stability.
[0113] Second, as Figure 9 As shown, Figure 9 The present invention provides a schematic diagram of the structure of a conveyor line intelligent detection system based on machine vision. The present invention provides a conveyor line intelligent detection system based on machine vision, including: A first acquiring unit 21 is used to acquire the running speed information of the conveyor line and the food type information; A second acquisition unit 22 is used to acquire food image information in a target area of the conveyor line; The third acquisition unit 23 is used to input the above-mentioned running speed information, the above-mentioned food type information and the above-mentioned food image information into the food defect detection model to obtain food defect information, wherein the above-mentioned food defect detection model integrates the above-mentioned food image information, the above-mentioned running speed information and the above-mentioned food type information, combines shape coding and attention enhancement, to obtain multi-category food defect recognition and scoring.
[0114] It is understandable that a conveyor line intelligent detection system based on machine vision can also perform the steps of any method described in the first aspect.
[0115] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A conveyor line intelligent detection method based on machine vision, characterized in that: include: Obtain the running speed information and food type information of the conveyor line; Obtain food image information in the target area of the conveyor line; The running speed information, the food type information and the food image information are input into a food defect detection model to obtain food defect information, wherein the food defect detection model integrates the food image information, the running speed information and the food type information, and combines shape encoding and attention enhancement to obtain multi-category food defect recognition and scoring.
2. The intelligent detection method for conveyor lines based on machine vision according to claim 1 is characterized in that: The food defect detection model includes: A multimodal input fusion module is used to receive and fuse the food image information, the running speed information and the food type information to form a multimodal input vector, wherein the multimodal input vector includes an image vector I, a speed vector and type vector ; A feature extraction backbone network, configured to extract a multi-scale image feature map group F from the image vector I based on a built-in multi-scale convolutional neural network; Speed perception and type adaptation embedding module, for based on the multi-scale image feature map group F, the speed vector and the type vector Generate speed type fusion feature map F'; Spatial channel attention enhancement module, used for generating enhanced feature maps from the multi-scale image feature map group F and the speed type fusion feature map F' ; Shape encoding branch, used to enhance the feature map Modeling food structural morphology through Fourier edge description and polar coordinate transformation and evaluating deviations from standard shapes to obtain shape consistency scores ; Defect detection branch, used to enhance the feature map Perform defect identification to obtain defect frame set B, category set C, confidence score set P and heat map ; Surface coverage judgment branch, used to enhance the feature map Determine the integrity of food surface coatings or sauces to obtain coverage rating labels .
3. The intelligent detection method for conveyor lines based on machine vision according to claim 2 is characterized in that: The specific processing steps of the speed perception and type adaptation embedded module include: Based on the first multilayer perceptron according to the velocity vector Generate velocity embedding vector ; Based on the second multilayer perceptron according to the type vector Generate type embedding vector ; Based on the velocity embedding vector and the embedding vector Generate control signals ; The control signal Applied to each channel of the multi-scale image feature map group F to form the speed type fusion feature map F'.
4. The intelligent detection method for conveyor lines based on machine vision according to claim 2 is characterized in that: The specific processing steps of the spatial channel attention enhancement module include: The multi-scale image feature map group F and the speed type fusion feature map F' are weightedly fused to generate a fusion feature map ; The fusion feature map Perform channel attention modeling to obtain channel enhanced feature maps ; The channel enhanced feature map Perform spatial attention wavelet building to generate the enhanced feature map .
5. The intelligent detection method for conveyor lines based on machine vision according to claim 2 is characterized in that: The specific processing steps of the shape coding branch include: The enhanced feature map Row channel compression and edge detection operations to generate edge maps ; Based on the center point coordinates of the target area , the edge graph Perform polar coordinate transformation to obtain polar coordinate edge map , and extract the radial profile function of the corresponding edge based on each angular direction ; For the radial profile function Perform discrete Fourier transform to generate the Fourier descriptor of the current product ; The Fourier descriptor Fourier descriptors with pre-stored standard template shapes Perform frequency domain distance calculation to obtain the descriptor difference ; Based on the descriptor difference , the shape consistency score is calculated by a nonlinear mapping function : in, is the preset scoring sensitivity adjustment coefficient, Indicates the degree of consistency between the structural shape of the current food product and the standard template.
6. The intelligent detection method for conveyor lines based on machine vision according to claim 2 is characterized in that: The specific processing steps of the defect detection branch include: The enhanced feature map Perform parallel feature regression and classification through multiple convolutional prediction heads to generate classification tensors , regression tensor and the confidence tensor ; Based on the regression tensor Map each candidate point to a corresponding set of defect bounding boxes , where each bounding box Including center point coordinates and width and height parameters; For the classification tensor and the confidence tensor Jointly calculate the comprehensive confidence score of each candidate box , and perform non-maximum suppression processing according to the score size to obtain the optimized defect category set and confidence score set ; Based on confidence score set Distribution in spatial location to generate defect saliency heatmap .
7. The intelligent detection method for conveyor lines based on machine vision according to claim 2 is characterized in that: The specific processing steps of the surface coverage judgment branch include: The enhanced feature map Extracting color anomaly response heatmap ; The heat map Divide the spatial region into multiple sub-blocks, calculate the response mean and variance of each sub-block respectively, and concatenate the statistical values of all sub-blocks into a spatial consistency description vector ; The description vector Input a multi-class classifier and output a label of the coverage integrity level of the food surface .
8. The intelligent detection method for conveyor lines based on machine vision according to claim 2 is characterized in that: The method further comprises: When the food type information When indicating that the current target food belongs to the pie category, In the shape encoding branch, the feature map is enhanced Before performing structural modeling, the method includes: The enhanced feature map Performing a directional gradient mapping operation to obtain a multi-directional gradient map group; Based on each directional gradient map, the local extreme response area is calculated and the boundary is traced to extract the refined edge area; The refined edge region is input into the polar coordinate transformation and Fourier description process to improve the extraction accuracy and scoring stability of the shape boundary of pie products.
9. The intelligent detection method for conveyor lines based on machine vision according to claim 1, characterized in that: The specific training process of the food defect detection model includes: Collect training data sets, which include multiple sample images, conveyor line speed information and food type labels corresponding to each training image, forming a triplet training input ,in, For training images, is the training speed value, To train food type labels; For the training image Performing image enhancement processing, wherein the image enhancement processing includes at least one of random flipping, affine transformation, color perturbation, and blur noise, to expand the training data set; Label the target output corresponding to each training image, which includes a set of defect bounding boxes , defect category label set , defect confidence score , shape consistency score and surface coverage grade labels ; The training image , the training speed value and the training food type label are jointly input into the food defect detection model; The total loss function is constructed, and the food defect detection model is trained on multiple tasks: The food defect detection model is updated through a stochastic gradient descent optimizer to complete model training.
10. A conveyor line intelligent detection device based on machine vision, characterized in that: include: A first acquiring unit is used to acquire the running speed information of the conveyor line and the food type information; A second acquisition unit is used to acquire food image information of a target area of the conveyor line; The third acquisition unit is used to input the running speed information, the food type information and the food image information into a food defect detection model to obtain food defect information, wherein the food defect detection model integrates the food image information, the running speed information and the food type information, and combines shape coding and attention enhancement to obtain multi-category food defect recognition and scoring.
Citation Information
Patent Citations
Method and device for detecting defects of toughened glass insulator
CN103149215A
Quick-frozen gristle crispy meat food prepared from poultry bones and meat and production method of food
CN104000222A
Plane bread online flaw detection method and device
CN115201208A
High-precision contact lens edge defect detection method and system based on machine vision
CN117269179A
Multi-defect-category insulator defect detection method based on multi-angle feature enhancement
CN118469946A