A real-time monitoring method and system for a snap-fastening installation assembly line

Through multi-angle image acquisition and multi-cascade depth separation inverse convolution processing, combined with spatial shuffling and rearrangement and bidirectional attention calculation, the complexity and real-time problems of snap installation pipeline monitoring in the prior art are solved, and accurate capture of snap features and real-time monitoring of high-speed pipelines are realized.

CN119941723BActive Publication Date: 2025-06-24XIAN QIAOLUMING TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510421759.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-06-24
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing automotive decorative panel clamp installation assembly line monitoring methods have problems such as high complexity, accuracy and real-time difficulty, especially when dealing with small target clamp features and adapting to light changes.

Method used

Multi-angle image acquisition and original size maintenance technology are adopted, combined with multi-cascade depth, reverse convolution processing, spatial shuffling and rearrangement, and bidirectional attention calculation, the geometric and appearance characteristics of the snaps are extracted, and the pipeline control instructions are output through a fixed time hierarchical decision-making mechanism analysis.

Benefits of technology

Effectively capture snap features of different scales, meet the real-time monitoring needs of high-speed assembly lines, and improve the accuracy of decision-making and the stability of the system under light changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941723B_ABST
    Figure CN119941723B_ABST
Patent Text Reader

Abstract

The present application relates to the field of image processing technology, and discloses a real-time monitoring method and system for a snap-fastener installation assembly line. The method includes: performing image acquisition and multi-stage cascaded depthwise separable transposed convolution processing on the snap-fastener installation area to obtain a multi-scale snap-fastener feature image; performing spatial shuffle rearrangement and bidirectional attention calculation on the multi-scale snap-fastener feature image to obtain a snap-fastener segmentation mask and a snap-fastener type identifier; extracting geometric and appearance features based on the snap-fastener segmentation mask and the snap-fastener type identifier, and calculating the deviation degree from the normal range to obtain a snap-fastener installation quality score and a defect heat map; performing fixed-time hierarchical decision-making mechanism analysis based on the snap-fastener installation quality score and the defect heat map, and outputting an assembly line control instruction. The present application effectively captures snap-fastener features at different scales, meets the real-time monitoring requirements of a high-speed assembly line, and at the same time maintains the accuracy of decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly to a real-time monitoring method and system for a snap-fastener installation assembly line. Background Art

[0002] The installation quality of snap-fasteners on automotive trim panels has an important impact on the overall vehicle assembly quality and driving safety. Traditional manual inspection methods have problems such as low efficiency and poor consistency, and it is difficult to meet the requirements of modern high-speed assembly line production. With the continuous improvement of the intelligent level of the automotive manufacturing industry, automated monitoring technology based on image processing has emerged. However, the current monitoring methods face multiple technical challenges: on the one hand, the supervised learning paradigm requires a large number of defect samples, and it is difficult to obtain various types of snap-fastener installation defect samples in actual production; on the other hand, although the anomaly detection paradigm does not require a large number of samples, it cannot accurately locate specific snap-fastener defect types. In addition, the snap-fasteners on the trim panel are small in size and have unclear features, which causes problems such as target semantic distortion and background false alarms in the application of existing image processing methods.

[0003] Existing monitoring methods for snap-fastener installation assembly lines on automotive trim panels generally have problems such as high complexity and difficulty in balancing accuracy and real-time performance. Traditional methods mostly adopt fixed-size adjustment and standard feature extraction processes, and cannot effectively process small-target snap-fastener features; at the same time, the calculation complexity of the feature extraction and enhancement process is high, and it is difficult to support the real-time monitoring requirements of high-speed assembly lines. In addition, existing methods have insufficient adaptability to changes in illumination and fluctuations in image quality, and are easily interfered in complex production environments, resulting in misjudgments. Especially when the quality of snap-fastener images is poor, the recognition accuracy drops significantly, and it cannot meet the reliability requirements of industrial production. Summary of the Invention

[0004] This application provides a real-time monitoring method and system for a snap-fastener installation assembly line. This application effectively captures snap-fastener features of different scales, meets the real-time monitoring requirements of high-speed assembly lines, and maintains the accuracy of decisions.

[0005] In a first aspect, this application provides a real-time monitoring method for a snap-fastener installation assembly line. The real-time monitoring method for a snap-fastener installation assembly line includes:

[0006] Performing image acquisition and multi-stage cascaded depthwise separable transposed convolution processing on the snap-fastener installation area to obtain a multi-scale snap-fastener feature image;

[0007] Performing spatial shuffle rearrangement and bidirectional attention calculation on the multi-scale snap-fastener feature image to obtain a snap-fastener segmentation mask and a snap-fastener type identifier;

[0008] Extracting geometric and appearance features based on the snap-fastener segmentation mask and the snap-fastener type identifier, and calculating the deviation degree from the normal interval to obtain a snap-fastener installation quality score and a defect heat map;

[0009] Perform a fixed-time grading decision mechanism analysis based on the buckle installation quality score and the defect heat map, and output a pipeline control instruction.

[0010] In a second aspect, the present application provides a real-time monitoring system for a buckle installation pipeline. The real-time monitoring system for a buckle installation pipeline includes:

[0011] An image acquisition module, configured to perform image acquisition and multi-stage cascaded depthwise separable transposed convolution processing on the buckle installation area to obtain a multi-scale buckle feature image;

[0012] An attention calculation module, configured to perform spatial shuffle rearrangement and bidirectional attention calculation on the multi-scale buckle feature image to obtain a buckle segmentation mask and a buckle type identifier;

[0013] An extraction module, configured to extract geometric and appearance features based on the buckle segmentation mask and the buckle type identifier, and calculate the deviation degree from the normal interval to obtain a buckle installation quality score and a defect heat map;

[0014] An output module, configured to perform a fixed-time grading decision mechanism analysis based on the buckle installation quality score and the defect heat map, and output a pipeline control instruction.

[0015] In the technical solution provided by the present application, the multi-angle image acquisition and original size preservation technology is adopted to avoid the loss of small target buckle features caused by image scaling. At the same time, the adaptive histogram equalization and edge enhancement processing enhance the saliency of the buckle contour features. The multi-stage cascaded depthwise separable transposed convolution structure significantly reduces the computational complexity. By constructing a receptive field pyramid through four cascaded modules with increasing dilation rates, buckle features at different scales can be effectively captured. The spatial shuffle rearrangement technology breaks the spatial limitation of traditional feature extraction, and the dynamic gain adjustment mechanism based on image quality and feature response intensity improves the stability of the system under changing lighting conditions. Few-shot prototype matching and bidirectional attention calculation solve the problem of insufficient buckle defect samples. By prototype intensity downsampling and bidirectional attention mechanism, buckle features can be accurately captured and background interference can be suppressed. The multi-dimensional normal interval deviation degree calculation based on geometric and appearance features realizes the accurate quantification of buckle installation quality. At the same time, the defect heat map generated by the gradient backpropagation technology intuitively shows the defect location and severity. The fixed-time grading decision mechanism innovatively solves the problem of unstable response time of the monitoring system. Through a three-level decision structure and time budget allocation, it ensures that the system can complete the decision within a preset fixed time regardless of the complexity of the situation, meeting the real-time monitoring requirements of high-speed pipelines while maintaining the accuracy of the decision. Description of the Drawings

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0017] Figure 1 It is a schematic diagram of an embodiment of the real-time monitoring method for the snap-fastener installation production line in the embodiments of the present application;

[0018] Figure 2 It is a schematic diagram of an embodiment of the real-time monitoring system for the snap-fastener installation production line in the embodiments of the present application. Specific embodiments

[0019] The embodiments of the present application provide a real-time monitoring method and system for a snap-fastener installation production line. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and the above drawings of the present application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the term "comprising" or "having" and any variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0020] For ease of understanding, the following describes the specific process of the embodiments of the present application. Please refer to Figure 1 , an embodiment of the real-time monitoring method for the snap-fastener installation production line in the embodiments of the present application includes:

[0021] Step S101: Perform image acquisition and multi-stage cascaded depthwise separable transposed convolution processing on the snap-fastener installation area to obtain a multi-scale snap-fastener feature image;

[0022] It can be understood that the execution subject of the present application can be the real-time monitoring system for the snap-fastener installation production line, or a terminal or a server. Specifically, it is not limited here. The embodiments of the present application will be described by taking the server as the execution subject as an example.

[0023] Specifically, a high-frame-rate industrial camera array is arranged on the snap fastener installation production line. These cameras are arranged at multiple angles to ensure a full-range monitoring of the snap fastener installation process on the decorative panel. The main camera is installed perpendicular to the surface of the decorative panel, while at least two auxiliary cameras are arranged at fixed angles from 15° to 75°, forming a multi-view observation system. The camera array adopts a high-frame-rate acquisition mode of no less than 120 frames per second, and the image resolution reaches 1920×1080 pixels to ensure that the acquired original images have sufficient clarity and details. The image acquisition system transmits the acquired image data to the image processing server at high speed through optical fibers to minimize signal delay and interference during the transmission process. To achieve synchronous control with the production line beat signal, a trigger is used to control the shooting timing of the camera, so that the camera accurately triggers the image acquisition process when the snap fastener reaches the preset position, and obtains temporally synchronized snap fastener images. And the original image size is kept unchanged during the entire preprocessing process, effectively avoiding the problem of small target snap fastener feature loss caused by size scaling. In the image preprocessing stage, the temporally synchronized snap fastener images are processed by adaptive histogram equalization. By adjusting the brightness distribution of the image pixels, the overall brightness of the image reaches an equilibrium state, thereby eliminating the brightness inconsistency problem caused by uneven light sources or reflective differences in the decorative panel material, and obtaining a brightness-corrected image. According to the noise level of the brightness-corrected image, an appropriate Gaussian filter kernel size is automatically selected for noise suppression processing. While retaining the details of the image, the Gaussian filter effectively removes high-frequency noise and obtains a noise-reduced snap fastener image. After noise elimination, color normalization is performed on the noise-reduced image to minimize the influence of material and color differences of different batches of decorative panels, and a standardized image with unified color is obtained. Edge enhancement processing is performed on the standardized image, and an improved Canny operator is used to accurately extract the edge of the decorative panel and the contour features of the snap fastener. By enhancing the contrast between the target and the background in the image, a preprocessed snap fastener image is obtained. Multi-stage cascaded depthwise separable transposed convolution processing is performed on the preprocessed snap fastener image, and a feature extraction encoder composed of multiple depthwise separable residual blocks is used for image processing. These residual blocks adopt a "dilate-compress-dilate" structure. Through 1×1 convolution operations, the image feature channels are expanded to six times the original, then spatial features are extracted through 3×3 depthwise separable convolutions, and finally the number of channels is compressed back to the original dimension through 1×1 convolutions. Compared with traditional convolution operations, depthwise separable convolution greatly reduces the amount of computation and the number of model parameters, making the entire model more lightweight while maintaining efficient feature extraction capabilities. To capture snap fastener features at different scales, the depthwise separable convolutions in each cascaded module adopt different dilation rates, which are set to 1, 2, 4, and 8 respectively, thus forming a feature pyramid structure with gradually expanding receptive fields. Finally, a multi-scale snap fastener feature image is obtained.

[0024] In this embodiment, the preprocessed buckle image is input into a 1×1 convolution operation module for channel expansion processing. Through 1×1 convolution, the number of feature channels of the input image is quickly expanded to multiple times the original dimension, for example, expanded to six times the original number of channels. The significance of channel expansion is to provide a richer feature expression space for subsequent depth convolution operations by increasing the dimension of feature channels, thereby effectively enhancing the model's learning ability for complex image features. After the expanded channel feature map is generated, the high-dimensional feature map is input into a depthwise separable inverse convolutional network composed of four cascaded feature extraction modules. In these feature extraction modules, each module uses 3×3 depthwise separable convolution operations with different dilation rates, and the dilation rates are set to 1, 2, 4, and 8 respectively. This design forms a feature pyramid structure with a gradually expanding receptive field, which can capture buckle features at different scales simultaneously. The difference between depthwise separable convolution and traditional convolution is that it decomposes the standard convolution into two steps: depth convolution and pointwise convolution, achieving a significant reduction in computational complexity, especially in the case of multi-channel input, and the computational complexity is greatly optimized. When the hierarchical features with different receptive fields are extracted through depthwise separable convolution, these feature maps are again subjected to channel compression processing through 1×1 convolution operations to obtain compressed feature maps. The high-dimensional information in the expanded feature map is condensed into a more compact representation form while retaining the most critical feature information to reduce the load of subsequent calculation modules. To capture the dynamic change information during the buckle installation process, the compressed feature map of the current frame is calculated for temporal correlation with the compressed feature map of the previous frame. This process is achieved through the joint calculation of spatial displacement and channel similarity. The displacement features of the current frame and the previous frame in the spatial dimension are calculated, and combined with the similarity of features between channels, and a custom correlation function is used to obtain the temporal correlation feature map. Temporal correlation analysis can effectively capture the change in the installation state of the buckle during the movement on the production line. Especially when the position of the buckle undergoes a slight offset or rotation, the model can still maintain a high detection accuracy. Based on the temporal correlation feature map, the compressed feature map is subjected to feature enhancement processing. By fusing temporal features and spatial features, the model uses the information in the time dimension to improve stability when judging the installation state of the buckle. The feature enhancement module combines convolution operations and non-linear activation functions to strengthen the buckle area in the feature space, highlighting the target features and suppressing background interference. During this process, the output features of the four cascaded feature extraction modules are fused in a pyramid manner to form multi-level buckle features. The pyramid fusion process unifies feature maps of different scales to the same spatial dimension through interpolation operations, and then combines detailed features and global features through a layer-by-layer fusion method to construct a multi-scale feature representation that can both focus on detailed features (such as the edges and shapes of the buckle) and maintain global information (such as the spatial position and installation angle of the buckle). Channel attention mechanism and spatial attention mechanism are used to adaptively allocate channel weights and regional weights to the multi-level buckle features.In the channel attention mechanism, global feature aggregation is performed on each channel in the feature map through global average pooling operation. Then, the importance weights of each channel are calculated through a small multi-layer perceptron to generate a channel attention map. This attention map is weighted channel by channel with the original feature map, enabling the model to focus on the feature channels that are most effective for buckle recognition. In the spatial attention mechanism, convolution operation combined with activation function is used to generate a spatial attention map. By multiplying it pixel by pixel with the feature map, the response of specific regions in the image is emphasized. The spatial attention mechanism is suitable for suppressing the interference of background regions and highlighting the feature response of the area where the buckle is located. By combining the channel and spatial attention mechanisms, the model achieves a fine feature enhancement effect, significantly improving the saliency of buckle features, and finally obtaining a multi-scale buckle feature image.

[0025] Step S102: Perform spatial shuffle rearrangement and bidirectional attention calculation on the multi-scale buckle feature image to obtain a buckle segmentation mask and a buckle type identifier;

[0026] Specifically, spatial dimension unification processing is performed on different scale feature maps in the multi-scale snap feature image. For example, bilinear interpolation is used to align feature maps of different scales to the same spatial resolution, obtaining a unified feature map with the same spatial dimension. The unified spatial dimension feature maps are concatenated in the channel dimension to obtain a mixed feature map. A spatial shuffle rearrangement operation is performed on the concatenated mixed feature map to break the spatial structure in the feature map and enhance the model's perception ability of global information. Spatial shuffling breaks the original spatial position relationship by performing rearrangement and transpose operations on the concatenated feature map, enabling the model to obtain more comprehensive global information when learning features and eliminating the influence of deviations in certain spatial positions on feature extraction. After spatial shuffling, the obtained shuffled feature map is input into a 1×1 convolution operation module for processing. 1×1 convolution is used for channel compression, and through the update of the convolution kernel parameters, useful information in the feature map is further learned. The shuffled feature map undergoes a non-linear mapping through the sigmoid activation function, making the output of the feature map meet the actual target requirements. Different weights are assigned to each channel in the feature map according to its response intensity. To select the most effective features, a feature selection gating mechanism is introduced. This mechanism generates an attention map based on the response intensity at each position in the feature map, and through the element-wise multiplication with the original feature map, enables the model to pay more attention to features important for snap detection while suppressing irrelevant or redundant features. The model can adaptively assign different attention weights to each feature channel, thereby optimizing the feature expression ability and improving the accuracy of the model in practical applications. During the feature enhancement process, the image quality score in the preprocessing stage is combined with the current feature response intensity to calculate a comprehensive quality index. This index combines the quality of the image and the current feature response intensity, providing real-time feedback on the image quality to the model. Based on the comprehensive quality index, a gain parameter is dynamically determined through a non-linear mapping function. The dynamic gain parameter is used to adjust the enhancement intensity of the feature map to meet the processing requirements of images with different qualities. When the image quality is poor, the gain parameter is increased to enhance the key information in the feature map; while when the image quality is good, the gain parameter is appropriately decreased to avoid over-enhancement causing information overload. Through the dynamic adaptive gain enhancement mechanism, the model better adapts to changes in different production environments and effectively improves the detection accuracy. The enhanced small target snap features can still accurately identify the position and type of the snap when the image quality is not high or the snap target is small. Based on the enhanced small target snap features, few-shot prototype matching and bidirectional attention calculation are performed. Feature prototypes of each type of snap are extracted from a small number of labeled snap samples, and these prototype features are matched. During the matching process, the similarity calculation between feature vectors can help the model determine the type of snap in the current image. To improve the matching accuracy, a bidirectional attention mechanism is combined to calculate the attention in the spatial dimension and the channel dimension respectively.By calculating the spatial attention map, the spatial region related to the buckle is highlighted; while by calculating the channel attention map, the most discriminative feature channels are identified. Generate the buckle segmentation mask and the buckle type identifier.

[0027] In this embodiment, the mean value of a small number of labeled buckle installation sample features is calculated to generate prototype features for matching. By collecting a small number of samples in each buckle installation state (usually 1 to 5 samples per category), then extracting the feature vectors of these samples, and obtaining a prototype feature representation by means of weighted average or directly calculating the mean value of these feature vectors. The prototype features are processed by intensity downsampling to reduce the influence of noise on the prototype features. During the intensity downsampling process, the response intensity distribution in the prototype features is analyzed, and the feature dimensions with significant responses are preferentially retained, while the feature dimensions with large noise interference are suppressed. This method extracts high-intensity signals through threshold filtering or feature selection algorithms to form the support prototype features after noise reduction. Based on the enhanced small-target buckle features, three parallel branches with different-scale receptive fields are constructed to cope with the diversity of buckle sizes and positions. The model designs three parallel branches with different-scale receptive fields, each branch corresponding to a specific feature extraction module, and realizing multi-scale feature capture through convolutional kernels of different sizes and different dilation rates (such as 1, 2, 4). These parallel branches simultaneously process the detailed features and global features in the input image, enabling the model to maintain a high recognition accuracy when facing buckles of different sizes and shapes, and generating multi-level query feature images. The cosine similarity between the multi-level query feature images and the noise-reduced support prototype features is calculated to preliminarily determine the type and position of the buckle. The cosine similarity judges the similarity degree between the two by calculating the angle size between the feature vectors. Through the calculated preliminary feature matching results, the probability distribution of the buckle area is identified. The preliminary matching results are input into the two-dimensional calculation of spatial attention and channel attention to improve the accuracy and stability of feature matching. In the spatial dimension, the spatial attention map is calculated. By multiplying the preliminary matching results with the spatial feature map and applying the softmax activation function, the spatial attention weights of each pixel point are obtained, thereby highlighting the spatial regions related to the buckle. In the channel dimension, the model calculates the channel attention vector through the channel response intensity of the feature map. Through operations such as global average pooling and multi-layer perceptrons, the importance weights of each feature channel are generated. The spatial attention map and the channel attention vector are fused to obtain the fused attention result. The fusion process is realized through pixel-by-pixel and channel-by-channel weighted processing, combining the most important spatial regions and feature channels, enabling the model to more accurately identify the features of the buckle. The fusion operation is realized through 1×1 convolution and sigmoid activation function, so that the output feature map not only contains the spatial position information of the buckle, but also fuses the feature selection results on the channel, forming a comprehensive representation of the buckle features. The fused attention result is processed by the abnormal prior constraint module. This module establishes a prior knowledge model of the buckle installation position, shape and size, maps the reasonable distribution range of the buckle to the feature map, and significantly reduces the misclassification probability of the background area through prior information by constructing a probability model, generating the buckle segmentation mask and the buckle type identification.

[0028] Step S103: Extract geometric and appearance features based on the buckle segmentation mask and buckle type identifier, and calculate the deviation degree from the normal interval to obtain the buckle installation quality score and the defect heat map;

[0029] Specifically, perform morphological operations and edge extraction operations on the buckle segmentation mask. Morphological operations, through processing steps such as erosion, dilation, opening, and closing, help to highlight the geometric shape features of the buckle, while eliminating irregular background noise or small interfering points to ensure the clarity of the buckle contour. The edge extraction operation can help the system accurately identify the morphological features of the buckle, such as geometric parameters like area, perimeter, roundness, eccentricity, etc. These geometric features are extracted as feature vectors to represent the geometric features of the buckle. Based on the corresponding regions of the buckle type identifier and the buckle segmentation mask, perform texture analysis and color statistics to extract the appearance features of the buckle. Texture analysis reveals whether there are scratches or other damages on the buckle surface by calculating texture information on the buckle surface, such as roughness and texture uniformity. Color statistics determines the appearance quality of the buckle by calculating indicators such as color uniformity and color difference in the buckle region. The appearance feature vector and the geometric feature vector form the complete feature vector of the target buckle. Calculate the difference between the buckle feature vector and the upper and lower bounds of the preset normal interval. The normal interval refers to the range within which the geometric and appearance features of the buckle should fall in the normal installation state. For example, the roundness of the buckle should be close to 1, the eccentricity should be within the normal range, and the color and texture should be uniform. By comparing the values of each dimension in the buckle feature vector with the upper and lower bounds of the normal interval, calculate the deviation distance for each feature dimension. Based on the deviation distance of each feature dimension, perform a weighted summation calculation to obtain the buckle installation quality score. Different feature dimensions have different degrees of influence on the buckle quality. By setting the importance weights of each feature dimension, adjust the calculation method of the deviation distance. Calculate the feature vector gradient of the buckle installation quality score and backpropagate it to the original image space. Through the backpropagation algorithm, according to the relationship between the buckle quality score and the deviation distance of each feature, calculate the gradient information of each pixel, reflecting the contribution degree of each region to the quality score. Through the backpropagation of the gradient, correspond the quality score to the specific regions in the image, and then determine which regions have a greater impact on the buckle quality. Perform upsampling processing on the gradient map. Upsampling enlarges the gradient map to the resolution of the original image, making the heat map aligned with the original image. The regions shown in the heat map represent the parts with a large deviation from the normal interval in the quality score, and these regions correspond to the places where defects occur during the installation process.

[0030] Step S104: Perform a fixed-time hierarchical decision-making mechanism analysis based on the buckle installation quality score and the defect heat map, and output the pipeline control instruction.

[0031] Specifically, based on the snap - fit installation quality score and the defect heat map, a hierarchical decision - making structure of a direct judgment layer, a detailed analysis layer, and an expert decision - making layer is established, and a fixed time segment is allocated to each level to ensure the real - time performance and stability of the system in a high - speed assembly - line production environment. Through this hierarchical decision - making mechanism, the system can quickly process snap - fit installation cases that are obviously qualified or unqualified, and analyze and make expert judgments on complex or ambiguous samples, providing precise assembly - line control instructions. Construct a hierarchical decision - making structure based on the snap - fit installation quality score and the defect heat map. This structure divides the decision - making process into three levels: the direct judgment layer, the detailed analysis layer, and the expert decision - making layer. Each level has a fixed time - slice allocation to ensure that the total processing time of the entire decision - making process is constant, meeting the real - time processing requirements of the high - speed assembly line. The direct judgment layer is responsible for quickly processing snap - fit installation samples that are obviously qualified or unqualified. For these samples, decisions are made by quickly comparing the snap - fit installation quality score with double thresholds. If the snap - fit installation quality score is higher than the set upper threshold, it is directly judged as qualified; if it is lower than the lower threshold, it is judged as unqualified; if the quality score falls between the two thresholds, it is marked for further analysis. The decision - making process at this level is very efficient, quickly screening out most of the simple snap - fit installation samples, obtaining preliminary decision results, and calculating the remaining amount of the first time. For snap - fit installation samples within the threshold range, they are sent to the detailed analysis layer for further analysis. At this layer, the defect heat map is used to perform region segmentation and feature clustering on the snap - fit to be analyzed, in order to more precisely identify and analyze defects. This process uses a time - aware algorithm, which dynamically adjusts the analysis granularity according to the current production rhythm and processing capacity to ensure the fine - grained analysis of defects is completed within the specified time - slice. Through the processing of this layer, complex snap - fit installation defects are analyzed, obtaining intermediate decision results, and calculating the remaining amount of the second time. For particularly complex or ambiguous samples, the intermediate decision results, the snap - fit installation quality score, and the defect heat map are input into the expert decision - making layer for the final judgment. At this level, the system relies on a historical case library to perform similarity matching, making a comprehensive judgment by combining past experience and cases to obtain high - level decision results. The historical case library contains a large number of verified snap - fit installation samples. Through similarity matching, the most likely installation state of the current sample is inferred with the help of historical data, thus making a more accurate decision. After multi - level decision - making by the direct judgment layer, the detailed analysis layer, and the expert decision - making layer, the preliminary decision results, the intermediate decision results, and the high - level decision results are fused to obtain the final decision result. The final decision result is converted into specific assembly - line control instructions through a decision - instruction mapping table, including operations such as continuing production, pausing production, adjusting production parameters, and starting maintenance.

[0032] In the embodiments of the present application, the multi-angle image acquisition and original size preservation technology avoid the loss of small target buckle features caused by image scaling. At the same time, the adaptive histogram equalization and edge enhancement processing enhance the saliency of the buckle contour features. The multi-stage cascaded depthwise separable transposed convolution structure significantly reduces the computational complexity. By constructing a receptive field pyramid through four cascaded modules with increasing dilation rates, buckle features at different scales are effectively captured. The spatial shuffle rearrangement technology breaks the spatial limitation of traditional feature extraction, and the dynamic gain adjustment mechanism based on image quality and feature response intensity improves the stability of the system under changing lighting conditions. Few-shot prototype matching and bidirectional attention calculation solve the problem of insufficient buckle defect samples. By downsampling the prototype intensity and using the bidirectional attention mechanism, buckle features are accurately captured and background interference is suppressed. The multi-dimensional normal interval deviation calculation based on geometric and appearance features realizes the accurate quantification of the buckle installation quality. At the same time, the defect heat map generated by the gradient backpropagation technology intuitively shows the defect location and severity. The fixed-time hierarchical decision-making mechanism innovatively solves the problem of unstable response time of the monitoring system. Through the three-level decision-making structure and time budget allocation, it ensures that the system can complete the decision within the preset fixed time regardless of the complexity of the situation, meeting the real-time monitoring requirements of high-speed assembly lines while maintaining the accuracy of the decision.

[0033] In a specific embodiment, the process of executing step S101 may specifically include the following steps:

[0034] (1) Perform multi-angle image acquisition on the buckle installation area to obtain the original buckle image;

[0035] (2) Based on the pipeline beat signal, perform synchronous trigger control on the original buckle image and keep the original size unchanged to obtain the time-sequence synchronous buckle image;

[0036] (3) Perform adaptive histogram equalization processing on the time-sequence synchronous buckle image to obtain the brightness-corrected image;

[0037] (4) Determine the Gaussian filter kernel based on the noise level of the brightness-corrected image and perform noise suppression processing to obtain the noise-reduced buckle image. Perform color normalization processing on the noise-reduced buckle image to obtain the color-unified standardized image;

[0038] (5) Perform edge enhancement processing on the color-unified standardized image to obtain the preprocessed buckle image;

[0039] (6) Perform multi-stage cascaded depthwise separable transposed convolution processing on the preprocessed buckle image to obtain the multi-scale buckle feature image.

[0040] Specifically, multi-angle image acquisition is performed on the buckle installation area to obtain sufficient image information and comprehensively cover all aspects of the buckle. By configuring multiple industrial cameras in the buckle installation area and arranging these cameras at different angles, various perspectives of the buckle during the pipeline assembly process are captured. To obtain high-quality original images, cameras with high frame rates are selected, and the acquisition frequency of the cameras is set above 120 frames per second to capture every detail of the buckle on the high-speed pipeline in real time. The images acquired from multiple angles can ensure that the system accurately judges the state of the buckle from all directions, thus avoiding image distortion or incompleteness caused by a single angle. Based on the pipeline beat signal, synchronous trigger control is performed on the original buckle images to ensure the timing consistency of the images and obtain buckle images with synchronized timing. The core of the synchronous control lies in using the pipeline beat signal to trigger the image acquisition of the camera. By synchronizing with the beat of the pipeline, it is ensured that the images of each buckle are accurately captured at the specified time point. During synchronous acquisition, it is ensured that the image size does not change, and each captured image maintains the same resolution and size as the original image, avoiding the loss of key information during image scaling or cropping. Brightness correction processing is performed on the buckle images with synchronized timing. Adaptive histogram equalization is an image enhancement technique that can effectively improve the brightness distribution of images. Through this technique, the brightness distribution of the image can be adaptively adjusted so that the dark and bright parts in the image can be more clearly displayed. For the processing of buckle images, adaptive histogram equalization can eliminate the brightness deviation caused by the pipeline environment or uneven lighting conditions and obtain images with corrected brightness. The Gaussian filter kernel is determined based on the noise level of the brightness-corrected images for noise suppression. The Gaussian filter effectively smooths the image and reduces the influence of noise through convolution operations on the image. The size of the filter kernel is dynamically determined according to the noise level of the image. The greater the noise, the larger the size of the filter kernel to better smooth the image. Color normalization processing is performed on the noise-reduced buckle images to effectively reduce the influence of color and material differences on image processing, enabling comparison and analysis of buckle images from different batches within the same color space. Color normalization calculates the color mean and standard deviation of the image and then adjusts the color distribution of the image using a standardization method so that all images have consistent color characteristics. This embodiment can effectively reduce the differences in color and material among different buckle images. Edge enhancement processing is performed on the standardized images with unified colors to improve the clarity of the buckle edges and make the boundaries of the buckle more prominent. Edge enhancement uses an improved Canny operator. Through this operator, the contour features of the buckle are extracted and the significance of these features in the image is enhanced. The Canny operator determines the edge positions in the image through multiple convolution operations and threshold judgments and sharpens them to enhance the edge information in the image. The enhanced buckle images have higher contrast. Depthwise separable convolution processing is performed on the preprocessed buckle images.Depthwise separable convolution is an efficient convolution operation that decomposes standard convolution into two steps: depthwise convolution followed by pointwise convolution, thereby significantly reducing the computational load. Depthwise separable convolution can efficiently extract multi-scale features of the buckle. Especially when the size of the buckle is small, it can more accurately identify the detailed features of the buckle. This process realizes the extraction of features at each scale through multiple cascaded convolution modules, and by setting different dilation rates (such as 1, 2, 4, 8, etc.), it captures feature information such as the shape, edges, and texture of the buckle at different receptive field scales. Through these operations, a multi-scale buckle feature image is obtained. In the multi-cascaded depthwise separable convolution process, considering the different scales of the receptive fields, the dilation rate of the receptive field is defined for each module. Assume the dilation rate of each convolution module is. , through convolution operations with different dilation rates, feature maps at different scales are captured. By using different convolutional kernels and dilation rates, the features of the buckle are extracted at multiple scales, ensuring that small buckle features in the image are not ignored.

[0041] In a specific embodiment, the process of performing multi-cascaded depthwise separable transposed convolution on the preprocessed buckle image to obtain a multi-scale buckle feature image may specifically include the following steps:

[0042] (1) Perform channel expansion processing on the preprocessed buckle image through 1×1 convolution to obtain an expanded channel feature map;

[0043] (2) Input the expanded channel feature map into four cascaded feature extraction modules. Each cascaded feature extraction module respectively uses 3×3 depthwise separable convolution with dilation rates of 1, 2, 4, and 8 to obtain hierarchical features with different receptive fields;

[0044] (3) Perform channel compression processing on the hierarchical features with different receptive fields through 1×1 convolution to obtain a compressed feature map;

[0045] (4) Perform spatial displacement and channel similarity calculation on the compressed feature map of the current frame and the compressed feature map of the previous frame to obtain a time-domain correlation feature map that captures the dynamic changes in the buckle installation;

[0046] (5) Perform feature enhancement on the compressed feature map based on the time-domain correlation feature map, and perform pyramid fusion on the output features of the four cascaded feature extraction modules to obtain multi-level buckle features;

[0047] (6) Use channel attention mechanism and spatial attention mechanism to perform adaptive channel weight and regional weight allocation on the multi-level buckle features to obtain a multi-scale buckle feature image.

[0048] Specifically, perform channel expansion processing on the preprocessed buckle image through 1×1 convolution. The role of convolution is to expand the number of feature channels of an image, increase the dimension of feature representation, enhance the representation ability of the network, so as to learn more diverse feature information. In the convolution operation, the 1×1 convolution does not change the spatial dimension of the image, but only changes the number of channels. This embodiment can obtain a new feature map with more channels. The expanded-channel feature map is input into four cascaded feature extraction modules. Each module uses a 3×3 depthwise separable convolution for feature extraction, and the dilation rates of each module are set to 1, 2, 4, and 8 respectively to capture features with different receptive fields. Different from traditional convolution, depthwise separable convolution decomposes the convolution operation into two steps: first, depthwise convolution is performed, and then pointwise convolution is performed. This structure can effectively reduce the amount of computation and improve the operation efficiency of the model, especially in tasks that require extracting a large number of features. By setting different dilation rates, the features extracted by each module have different receptive fields, which means that convolution operations with different dilation rates can capture snap features at different scales. The larger the receptive field, the wider the area of the image that the convolution operation can focus on, thus improving the recognition ability of large-scale features, while small receptive fields are suitable for capturing detailed features. After passing through four cascaded convolution modules, channel compression processing is performed on the obtained feature maps with different receptive fields. Channel compression is achieved through Convolution reduces the number of channels, removes redundant feature information, and preserves the core features of the image. Through the compression process, the dimension of the feature map is reduced while maintaining the most important information, reducing the computational amount and improving efficiency. To capture the dynamic changes during the snap installation process, the compressed feature map of the current frame is calculated for temporal correlation with the compressed feature map of the previous frame. Temporal correlation analysis captures the dynamic changes in the image by comparing the feature differences between adjacent frames. This embodiment can effectively capture the dynamic change information of the snap, helping to analyze the movement trajectory and state changes of the snap on the assembly line. Based on the temporally correlated feature map, feature enhancement processing is performed on the compressed feature map to enhance the saliency of important features while suppressing the interference of irrelevant information. Enhancement is achieved through convolution operations and activation functions, specifically by adjusting the weights to strengthen the features related to the snap state, obtaining the enhanced feature map. In the multi-stage cascaded convolution module, the output features of all modules are pyramidally fused by layer-by-layer fusing feature maps of different scales. Pyramidal fusion unifies feature maps of different scales to the same spatial dimension through interpolation and convolution operations, and then performs weighted averaging or concatenation to obtain a multi-level snap feature image. On the fused multi-level snap feature image, an attention mechanism is introduced to optimize the feature representation. The channel attention mechanism and the spatial attention mechanism are used to adaptively allocate the channel weights and regional weights of the features respectively. The channel attention mechanism calculates the importance of each channel through global average pooling and a multi-layer perceptron, and the spatial attention mechanism emphasizes the key regions by calculating the response intensity of each spatial position in the feature map. Through the attention weighting of channels and space, the multi-level snap feature image is transformed into a high-quality multi-scale snap feature map.

[0049] In a specific embodiment, the process of executing step S102 may specifically include the following steps:

[0050] (1) Perform spatial dimension unification processing on the feature maps of different scales in the multi-scale snap feature image to obtain a unified feature map with the same spatial dimension;

[0051] (2) Concatenate the unified feature maps with the same spatial dimension in the channel dimension to obtain a mixed feature map, and perform spatial shuffle rearrangement on the mixed feature map to obtain a shuffled feature map;

[0052] (3) Input the shuffled feature map into a 1×1 convolution and sigmoid activation function for processing, and screen out effective features through a feature selection gating mechanism to obtain a spatially shuffled feature with attention weights;

[0053] (4) Calculate a comprehensive quality index based on the image quality score in the preprocessing stage and the current feature response intensity, and determine a dynamically adaptive gain parameter through non-linear mapping;

[0054] (5) Perform gain enhancement processing on the spatially shuffled features with attention weights according to the dynamically adaptive gain parameter to obtain enhanced small target buckle features;

[0055] (6) Perform few-shot prototype matching and bidirectional attention calculation based on the enhanced small target buckle features to obtain a buckle segmentation mask and a buckle type identifier.

[0056] Specifically, perform spatial dimension unification processing on buckle feature images of different scales. Use an interpolation method, such as bilinear interpolation, to make the spatial dimensions of all scale feature maps consistent, obtaining a unified feature map with the same spatial dimension. Concatenate the unified feature maps with the same spatial dimension in the channel dimension to fuse information from multiple scales and form a mixed feature map containing more features. Perform spatial shuffle rearrangement on the mixed feature map to disrupt the spatial structure of the image, enabling the model to learn global information without being restricted by local information. Spatial shuffling helps eliminate position dependence in the image, prompting the model to better capture global features and improve sensitivity to information in different regions, obtaining a shuffled feature map. Input the shuffled feature map into a 1×1 convolution and an activation function for processing. The 1×1 convolution is used for weighted summation in the channel dimension, enabling each channel's features to be processed and refined. Through the non-linear transformation of the activation function, enhance the expression of important channels in the feature map and suppress the influence of unimportant channels. Screen the features through a feature selection gating mechanism and select the most effective features. The role of the feature selection gating mechanism is to dynamically adjust the weights of each feature channel based on the response intensity of each feature. Through this mechanism, focus on the features most relevant to the buckle state, thereby improving the accuracy and efficiency of subsequent processing. The process of feature selection helps focus the system's attention on important features and avoid interference from background noise. Calculate a comprehensive quality index based on the image quality score in the preprocessing stage and the current feature response intensity. The image quality score reflects the clarity, contrast, and signal-to-noise ratio of the image, while the feature response intensity characterizes the importance of the buckle features. The calculation of the comprehensive quality index helps the system make adjustments under different quality conditions to ensure real-time adaptation to environmental changes. Based on the comprehensive quality index, dynamically adjust the gain parameter through a non-linear mapping function. The adjustment of the gain parameter is determined based on the current image quality and feature response intensity, aiming to enhance important features in the image and reduce interference from irrelevant information. Through mapping, dynamically adjust the gain according to different environmental conditions, thereby improving the detection accuracy of the buckle features. After obtaining the dynamic gain parameter, perform gain enhancement processing on the spatially shuffled features with attention weights. Adjust the intensity of the features according to their importance, making the key features of the buckle more prominent. By applying the gain to the feature map, strengthen the features related to the buckle and reduce the influence of background noise, obtaining an enhanced feature map.

[0057] Perform few-shot prototype matching and bidirectional attention calculation based on enhanced small target buckle features. The few-shot prototype matching helps the system judge the current state of the buckle by comparing with historical samples. The bidirectional attention mechanism adaptively adjusts weights in the spatial and channel dimensions, enabling the system to more accurately identify buckle features. Through this step, an accurate buckle segmentation mask and buckle type identification are obtained.

[0058] In a specific embodiment, the process of performing few-shot prototype matching and bidirectional attention calculation based on enhanced small target buckle features to obtain a buckle segmentation mask and buckle type identification may specifically include the following steps:

[0059] (1) Calculate the mean of the features of a small number of labeled buckle installation samples to obtain prototype features, and perform intensity downsampling on the prototype features to obtain denoised support prototype features;

[0060] (2) Construct three parallel branches with different scale receptive fields based on the enhanced small target buckle features to obtain a multi-level query feature image;

[0061] (3) Calculate the cosine similarity between the multi-level query feature image and the denoised support prototype features to obtain a preliminary matching result;

[0062] (4) Input the preliminary matching result into the dual-dimensional calculation of spatial attention and channel attention, respectively obtain a spatial attention map and a channel attention vector, and fuse the spatial attention map and the channel attention vector to obtain a fused attention result;

[0063] (5) Apply an anomaly prior constraint module to process the fused attention result, and suppress the interference of the background area according to the prior knowledge of the buckle installation position, shape and size to obtain a buckle segmentation mask and buckle type identification.

[0064] Specifically, calculate the mean of the features of a small number of labeled buckle installation samples to obtain prototype features, and these sample features provide a standard reference. The goal of the denoising operation is to remove high-frequency noise and reduce unnecessary details, making the support prototype features more stable and easier to use subsequently. Based on the enhanced small target buckle features, construct three parallel branches with different scale receptive fields to obtain a multi-level query feature image. By parallel processing of different scale receptive fields, capture the diversity and detailed information of the buckle features. Through parallel calculation, obtain a multi-level feature image Calculate the cosine similarity between the multi-level feature image and the denoised support prototype features to obtain a preliminary matching result. Cosine similarity is a standard method for measuring the similarity between two vectors, and it measures their similarity by calculating the angle between the two vectors. Through cosine similarity, accurately find the similar region between the current buckle and the standard buckle. Input the preliminary matching result into the two-dimensional calculation of spatial attention and channel attention to optimize the feature map. The spatial attention mechanism focuses on the importance of each spatial position in the image, and the spatial attention map represents the weight of each spatial position, while the channel attention mechanism focuses on the importance of each channel in the feature map, and the channel attention vector represents the weight of each channel. The spatial and channel attention mechanisms are calculated through global pooling and convolution operations respectively, and finally merged into a fused attention result. Apply the anomaly prior constraint module to the fused attention result. This module helps the system suppress the interference of the background region by combining the prior knowledge of the installation position, shape, and size of the buckle. The position, shape, and size of the buckle are important features in buckle detection. By using this prior knowledge, better exclude background noise and focus on the actual features of the buckle. Through prior constraint, obtain the final buckle segmentation mask and buckle type identification.

[0065] Among them, calculate the mean value of the features of a small number of labeled buckle installation samples to obtain the prototype features:

[0066]

[0067] Among them: Represents the original prototype feature; Represents the feature representation of the k-th support sample;

[0068] Represents the total number of labeled support samples; In the buckle installation pipeline, extract and average the features of a small number of known correctly installed buckle samples to form a standard feature template as the benchmark for subsequent matching.

[0069] When the enhanced small target buckle features are obtained, calculate the cosine similarity with the prototype features:

[0070]

[0071] Among them: Represents the cosine similarity map; Represents the query feature map (i.e., the enhanced small target buckle features); It represents the support prototype features after noise reduction; (x, y) represents the spatial position coordinates; c represents the feature channel index. In practical applications, this formula is used to evaluate the similarity between the currently detected buckle features and the standard prototype. The higher the similarity, the closer the buckle installation state is to the standard state. After fusing spatial attention and channel attention, an anomaly prior constraint is applied:

[0072]

[0073] Where: It represents the final buckle segmentation mask; It represents the fused attention result; It represents the prior constraint on the buckle position, shape, and size; It represents a constraint function used to suppress background area interference. In the buckle installation monitoring system, the prior knowledge of the buckle (such as installation position, normal size range, etc.) is used to correct the attention result, filter out the misdetection areas that do not meet the expectations, and obtain the accurate buckle segmentation mask and type identification.

[0074] In a specific embodiment, the process of executing step S103 may specifically include the following steps:

[0075] (1) Perform morphological operations and edge extraction on the buckle segmentation mask to obtain the buckle geometric feature vector;

[0076] (2) Based on the corresponding regions of the buckle type identification and the buckle segmentation mask, perform texture analysis and color statistics to obtain the buckle appearance feature vector;

[0077] (3) Combine the buckle geometric feature vector and the buckle appearance feature vector to obtain the target buckle feature vector;

[0078] (4) Calculate the difference between the buckle feature vector and the upper and lower bounds of the preset normal interval to obtain the deviation distance of each feature dimension;

[0079] (5) Perform a weighted summation calculation based on the deviation distance of each feature dimension and the feature importance weight to obtain the buckle installation quality score;

[0080] (6) Calculate the feature vector gradient for the buckle installation quality score and backpropagate it to the original image space, and perform upsampling processing to obtain the defect heat map.

[0081] Specifically, morphological operations and edge extraction are performed on the buckle segmentation mask. Morphological operations, such as dilation, erosion, and opening operations on the image, help extract the structural features of the buckle. Through morphological processing, the edges and contours of the buckle become more prominent, thus better describing the geometric characteristics of the buckle. Edge extraction is performed to obtain the geometric feature vector of the buckle, including information such as the shape, size, and contour of the buckle. Extract the appearance features of the buckle. Based on the corresponding regions of the buckle type identifier and the buckle segmentation mask, texture analysis and color statistics are performed. Texture analysis captures the detailed changes on the buckle surface by calculating the distribution characteristics of the patterns on the buckle surface, and color statistics analyze the color distribution on the buckle surface. Through these analyses, the appearance features of the buckle are obtained. The geometric feature vector and the appearance feature vector of the buckle are merged to obtain the complete feature vector of the target buckle. The merging operation forms a comprehensive feature vector containing information such as the shape, size, color, and texture of the buckle by concatenating the geometric feature vector and the appearance feature vector in the feature dimension. The deviation distance of each feature dimension is calculated by comparing the target buckle feature vector with the upper and lower bounds of the preset normal interval. By comparing with the normal range, it is evaluated whether the buckle meets the standard. By calculating the deviation distances of all feature dimensions, the buckle deviation degree is obtained, thereby determining whether the buckle has defects. Based on the deviation distance of each feature dimension and the feature importance weight, a weighted summation calculation is performed to obtain the buckle installation quality score. Considering the importance of each feature dimension comprehensively, the overall quality score of the buckle is calculated. Through this score, the installation quality of the buckle is quantitatively evaluated, and the buckles on the production line are monitored in real time. Calculate the feature vector gradient of the buckle installation quality score and backpropagate the gradient to the original image space. Through gradient backpropagation, the influence of the quality score is transmitted back to each pixel of the image, thereby calibrating the key regions in the buckle image. By upsampling the feature vector gradient, a high-resolution defect heat map is obtained.

[0082] Among them, for the obtained target buckle feature vector, it is necessary to calculate its deviation distance from the upper and lower bounds of the preset normal interval. The preset normal interval is obtained based on a large amount of historical data statistics and represents the feature range of normal installed buckles. The calculation formula for the deviation distance is:

[0083]

[0084] Where: represents the deviation distance of the th dimension of the feature vector; represents the th dimension value of the target buckle feature vector; represents the th dimension value of the lower bound of the preset normal interval; represents the Dimension value. This formula calculates the degree of deviation of the snap feature from the normal range in each dimension. When the feature value falls within the normal range, the deviation distance is zero; when the feature value exceeds the upper bound, the deviation distance is the difference between the feature value and the upper bound; when the feature value is below the lower bound, the deviation distance is the difference between the lower bound and the feature value. Then, the deviation distances need to be weighted and summed using the feature importance weights to obtain the final snap installation quality score. The feature importance weights reflect the degree of influence of different features on the snap installation quality and are usually set by domain experts or automatically learned through machine learning methods. The formula for calculating the snap installation quality score is:

[0085]

[0086] Where: represents the snap installation quality score, with a full score of 100 points; represents the importance weight of the dimensional feature; represents the total number of dimensions of the feature vector; represents the scaling factor used to adjust the sensitivity of the score; represents the penalty exponent, usually set to 2, which gives a more severe penalty for larger deviations.

[0087]

[0088] In actual application scenarios, the feature dimensions may include geometric features of the snap (such as area, perimeter, aspect ratio, roundness, etc.) and appearance features (such as color mean, texture gradient, surface smoothness, etc.). For example, for the monitoring of snap installation on an automotive parts assembly line, higher weights may be given to key geometric features such as position offset and angle tilt, while lower weights may be assigned to secondary features such as color change.To more precisely express the mutual influence between different features, a feature correlation matrix can be introduced for correction to obtain an improved quality score calculation formula:

[0089]

[0090] Where: represents the correlation coefficient between feature and feature ; It represents the correlation impact factor, which controls the impact degree of correlation on the final score. In this way, the buckle installation quality score not only considers the independent deviation degree of each feature, but also takes into account the interaction between features, making the scoring result more comprehensive and accurate. The higher the score, the better the buckle installation quality, and the lower the score, the more likely there are quality problems. In the actual application of the buckle installation production line, this scoring mechanism can quickly identify the poorly installed buckles and guide the subsequent generation of defect heat maps and production line control decisions, effectively ensuring product quality and production efficiency.

[0091] In a specific embodiment, the process of executing step S104 may specifically include the following steps:

[0092] (1) Based on the buckle installation quality score and the defect heat map, establish a hierarchical decision-making structure of a direct judgment layer, a detailed analysis layer, and an expert decision-making layer, allocate fixed time slices for each layer of the hierarchical decision-making structure, and obtain a fixed-time decision-making framework with a constant total processing time;

[0093] (2) In the direct judgment layer, perform a double-threshold quick comparison on the buckle installation quality score, directly output the decision result for samples that are obviously qualified or obviously unqualified, mark the samples within the threshold range as to be further analyzed, and obtain a preliminary decision result and the remaining amount of the first time;

[0094] (3) Input the samples to be further analyzed and the defect heat map into the detailed analysis layer, perform regional segmentation and feature clustering on the defect heat map, and use a time-aware algorithm to complete the fine analysis of the defects to obtain an intermediate decision result and the remaining amount of the second time;

[0095] (4) For particularly complex samples, input the intermediate decision result, the buckle installation quality score, and the defect heat map into the expert decision-making layer together, perform similarity matching based on the historical case library, and obtain a high-level decision result;

[0096] (5) Perform decision fusion on the preliminary decision result, the intermediate decision result, and the high-level decision result to obtain a final decision result, and convert the final decision result into specific production line control instructions through a decision-instruction mapping table.

[0097] Specifically, the hierarchical decision-making mechanism based on the buckle installation quality score and the defect heat map is a time-sensitive processing framework. By establishing a three-layer structure: a direct judgment layer, a detailed analysis layer, and an expert decision-making layer, each layer is responsible for judgment tasks of different complexities. To ensure real-time performance, fixed time slices are allocated to each layer to implement a decision-making framework with a constant total processing time. For example, within a total processing time limit of 60 milliseconds, 20 milliseconds are allocated to the direct judgment layer, 30 milliseconds to the detailed analysis layer, and 10 milliseconds to the expert decision-making layer. This time allocation ensures that the buckle detection can keep up with the rhythm of the production line and avoid production stagnation caused by slow processing.

[0098] In the direct judgment layer, a double-threshold quick comparison is performed on the buckle installation quality score. The double thresholds refer to setting an upper threshold and a lower threshold, which divide the buckle quality into three categories. When the score is higher than the upper threshold (e.g., the score is 95 and the upper threshold is 90), the buckle is judged to be obviously qualified; when the score is lower than the lower threshold (e.g., the score is 60 and the lower threshold is 75), the buckle is judged to be obviously unqualified; when the score falls between the two thresholds (e.g., the score is 85), it is marked for further analysis. The processing speed of this layer is extremely fast, usually only taking a few milliseconds to complete the judgment, and the remaining time (the remaining amount at the first time) is passed to the next layer to improve the overall efficiency.

[0099] The samples to be further analyzed and the defect heat map are input into the detailed analysis layer for more in-depth processing. Region segmentation refers to pixel-level division of the heat map to extract the abnormal regions. This process uses an adaptive threshold segmentation algorithm to divide the image into foreground (defect region) and background (normal region) according to the intensity values of the heat map. Then, feature clustering is performed, that is, the extracted regions are grouped according to features such as position, shape, and intensity to identify defects with similar characteristics. The time-aware algorithm is an algorithm that automatically adjusts the analysis accuracy within a limited time, giving priority to processing important features and gradually reducing the computational complexity as the time consumption increases to ensure completion of processing within the allocated time slice. After this layer is completed, an intermediate decision result and the remaining amount of the second time are obtained. For particularly complex samples, the intermediate decision result, the buckle installation quality score, and the defect heat map are jointly input into the expert decision layer. This layer performs similarity matching based on the historical case library, which contains a large amount of known buckle status data and their corresponding processing results. The similarity matching uses the weighted Euclidean distance to calculate the similarity between the current sample and the historical cases, and selects the closest case as a reference. This experience-based decision-making method can handle complex situations with fuzzy boundaries and difficult-to-define rules, improving the system's ability to handle abnormal situations.

[0100] Decision fusion is performed on the decision results of the three layers to obtain the final decision result. The decision fusion uses the weighted voting method, assigning different weights according to the credibility of the results of each layer and comprehensively obtaining the final judgment. Through a pre-set decision-instruction mapping table, the decision result is converted into specific pipeline control instructions, such as continuing production, pausing inspection, adjusting parameters, etc., to achieve precise control of the production line.

[0101] Taking an actual buckle installation monitoring scenario as an example: A fixed buckle on a car door panel is installed through an assembly line. After the camera captures the buckle installation image, it enters the hierarchical decision-making system. First, in the direct judgment layer, the system calculates that the buckle installation quality score is 82 points, which is between the lower threshold of 75 points and the upper threshold of 90 points. Therefore, it is marked for further analysis and transmitted to the detailed analysis layer. In the detailed analysis layer, the system performs regional segmentation on the defect heat map and finds that there is a small-area high-intensity area at the left edge of the buckle in the heat map, with an area of about 3% of the total area of the buckle. Through feature clustering, this area is identified as a potential defect of the "edge not fully fitted" type. Since this situation is not clear enough, the sample enters the expert decision-making layer. The system finds 7 similar cases in the historical case library, and 5 of these cases have loose buckles during use. Based on this matching result, combined with the analysis of the previous two layers, the system finally determines that this buckle is "potentially risky and needs adjustment", and generates the corresponding assembly line control instruction to pause the current station and prompt the operator to readjust the buckle installation parameters. This judgment process takes a total of 58 milliseconds, which does not exceed the preset time limit of 60 milliseconds, ensuring the continuous operation of the assembly line.

[0102] The above describes the real-time monitoring method for the buckle installation assembly line in the embodiments of the present application. Next, the real-time monitoring system for the buckle installation assembly line in the embodiments of the present application will be described. Please refer to Figure 2 , an embodiment of the real-time monitoring system for the buckle installation assembly line in the embodiments of the present application includes:

[0103] An image acquisition module 201, configured to perform image acquisition and multi-stage cascaded depthwise separable transposed convolution processing on the buckle installation area to obtain a multi-scale buckle feature image;

[0104] An attention calculation module 202, configured to perform spatial shuffle rearrangement and bidirectional attention calculation on the multi-scale buckle feature image to obtain a buckle segmentation mask and a buckle type identifier;

[0105] An extraction module 203, configured to extract geometric and appearance features based on the buckle segmentation mask and the buckle type identifier, and calculate the deviation degree from the normal interval to obtain a buckle installation quality score and a defect heat map;

[0106] An output module 204, configured to perform fixed-time hierarchical decision-making mechanism analysis based on the buckle installation quality score and the defect heat map, and output an assembly line control instruction.

[0107] Through the collaborative cooperation of the above-mentioned various components, the adoption of multi-angle image acquisition and original size preservation technology avoids the loss of small target snap features caused by image scaling. At the same time, the significance of snap contour features is enhanced through adaptive histogram equalization and edge enhancement processing. The multi-cascaded depthwise separable transposed convolution structure significantly reduces the computational complexity. A receptive field pyramid is constructed through four cascaded modules with increasing dilation rates to effectively capture snap features at different scales. The spatial shuffle rearrangement technology breaks the spatial limitation of traditional feature extraction, and the dynamic gain adjustment mechanism based on image quality and feature response intensity improves the stability of the system under changing lighting conditions. Few-shot prototype matching and bidirectional attention calculation solve the problem of insufficient snap defect samples. The snap features are accurately captured and background interference is suppressed through prototype intensity downsampling and bidirectional attention mechanism. The multi-dimensional normal interval deviation calculation based on geometric and appearance features realizes the accurate quantification of snap installation quality. At the same time, the defect heat map generated by gradient backpropagation technology intuitively shows the defect location and severity. The fixed-time hierarchical decision-making mechanism innovatively solves the problem of unstable response time of the monitoring system. Through a three-level decision-making structure and time budget allocation, it ensures that the system can complete decision-making within a preset fixed time regardless of the complexity of the situation, meeting the real-time monitoring requirements of high-speed assembly lines while maintaining the accuracy of decision-making.

[0108] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0109] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a snap installation pipeline real-time monitoring device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0110] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A real-time monitoring method for a buckle installation line, characterized in that: The buckle installation line real-time monitoring method comprises: Step S101, performing image acquisition and multi-cascaded depth-separable inverse convolution processing on the buckle installation area to obtain a multi-scale buckle feature image; The process of executing step S101 includes: (1) Capture images of the buckle installation area from multiple angles to obtain the original buckle image; (2) Based on the pipeline beat signal, the original snap-in image is synchronously triggered and controlled, and the original size is kept unchanged to obtain a time-series synchronous snap-in image; (3) performing adaptive histogram equalization processing on the time-series synchronous snap-in image to obtain a brightness-corrected image; (4) determining a Gaussian filter kernel based on the noise level of the brightness-corrected image, performing noise suppression processing to obtain a denoised snap-in image, and performing color normalization processing on the denoised snap-in image to obtain a standardized image with uniform color; (5) performing edge enhancement processing on the standardized image with uniform color to obtain a preprocessed snap-in image; (6) performing multi-cascade depth-separable inverse convolution processing on the pre-processed buckle image to obtain a multi-scale buckle feature image; The process of executing step (6) includes: performing channel expansion processing on the pre-processed buckle image through 1×1 convolution to obtain an expanded channel feature map; inputting the expanded channel feature map into four cascade feature extraction modules, each cascade feature extraction module adopts 3×3 depth-separable convolution with expansion rates of 1, 2, 4, and 8 respectively to obtain hierarchical features of different receptive fields; performing channel compression processing on the hierarchical features of different receptive fields through 1×1 convolution to obtain a compressed feature map; performing spatial displacement and channel similarity calculation on the compressed feature map of the current frame and the compressed feature map of the previous frame to obtain a time domain related feature map that captures the dynamic changes of buckle installation; performing feature enhancement on the compressed feature map based on the time domain related feature map, and pyramidally fusing the output features of the four cascade feature extraction modules to obtain a multi-level buckle feature; using a channel attention mechanism and a spatial attention mechanism to adaptively allocate channel weights and regional weights to the multi-level buckle features to obtain a multi-scale buckle feature image; Step S102: performing spatial shuffling and rearrangement and bidirectional attention calculation on the multi-scale buckle feature image to obtain a buckle segmentation mask and a buckle type identifier; Step S103: extracting geometric and appearance features based on the buckle segmentation mask and the buckle type identifier, and calculating the normal interval deviation to obtain a buckle installation quality score and a defect heat map; Step S104: performing fixed-time grading decision mechanism analysis based on the buckle installation quality score and the defect heat map, and outputting pipeline control instructions.

2. The method for real-time monitoring of a buckle installation line according to claim 1, characterized in that: The step of performing spatial shuffling and re-arrangement and bidirectional attention calculation on the multi-scale snap feature image to obtain a snap segmentation mask and a snap type identifier includes: performing spatial dimension unification processing on feature maps of different scales in the multi-scale snap feature image to obtain a unified feature map of the same spatial dimension; The unified feature maps of the same spatial dimension are concatenated in the channel dimension to obtain a mixed feature map, and the mixed feature map is spatially shuffled and rearranged to obtain a shuffled feature map; The shuffled feature map is input into a 1×1 convolution and sigmoid activation function, and effective features are screened through a feature selection gating mechanism to obtain a spatial shuffled feature with attention weights; The comprehensive quality index is calculated based on the image quality score in the preprocessing stage and the current feature response strength, and the dynamically adaptive gain parameter is determined through nonlinear mapping; According to the dynamically adaptive gain parameter, the spatial shuffle feature with attention weight is subjected to gain enhancement processing to obtain an enhanced small target snap feature; Based on the enhanced small target buckle feature, few-shot prototype matching and bidirectional attention calculation are performed to obtain a buckle segmentation mask and a buckle type identifier.

3. The real-time monitoring method for buckle installation line according to claim 2, characterized in that: The performing of few-sample prototype matching and bidirectional attention calculation based on the enhanced small-target buckle feature to obtain a buckle segmentation mask and a buckle type identifier includes: performing mean calculation on a small number of labeled buckle installation sample features to obtain a prototype feature, and performing intensity downsampling processing on the prototype feature to obtain a noise-reduced supporting prototype feature; Based on the enhanced small target buckle feature, three parallel branches with different scales of receptive fields are constructed to obtain a multi-level query feature image; Calculating cosine similarity between the multi-level query feature image and the noise-reduced supporting prototype feature to obtain a preliminary matching result; Input the preliminary matching result into the dual-dimensional calculation of spatial attention and channel attention to obtain a spatial attention map and a channel attention vector respectively, and fuse the spatial attention map and the channel attention vector to obtain a fused attention result; The fused attention result is processed by an abnormal prior constraint module, and background area interference is suppressed according to prior knowledge of buckle installation position, shape and size to obtain a buckle segmentation mask and a buckle type identification.

4. The real-time monitoring method for buckle installation line according to claim 1, characterized in that: The extracting of geometric and appearance features based on the buckle segmentation mask and the buckle type identifier, and calculating the normal interval deviation to obtain the buckle installation quality score and defect heat map includes: performing morphological operations and edge extraction on the buckle segmentation mask to obtain a buckle geometric feature vector; Performing texture analysis and color statistics based on the buckle type identifier and the corresponding area of ​​the buckle segmentation mask to obtain a buckle appearance feature vector; Merging the buckle geometric feature vector and the buckle appearance feature vector to obtain a target buckle feature vector; Calculate the difference between the buckle feature vector and the preset upper and lower bounds of the normal interval to obtain the deviation distance of each feature dimension; A weighted sum calculation is performed based on the deviation distance and feature importance weight of each feature dimension to obtain a buckle installation quality score; The characteristic vector gradient of the buckle installation quality score is calculated and back-propagated to the original image space, and an upsampling process is performed to obtain a defect heat map.

5. The real-time monitoring method for buckle installation line according to claim 1, characterized in that: The fixed-time hierarchical decision mechanism analysis is performed based on the buckle installation quality score and the defect heat map, and the pipeline control instruction is output, including: based on the buckle installation quality score and the defect heat map, a hierarchical decision structure of a direct judgment layer, a detailed analysis layer and an expert decision layer is established, and a fixed time slice is allocated to each layer of the hierarchical decision structure to obtain a fixed-time decision framework with a constant total processing time; In the direct judgment layer, a dual-threshold quick comparison is performed on the buckle installation quality score, a decision result is directly output for obviously qualified or obviously unqualified samples, and samples within the threshold range are marked as pending for further analysis, so as to obtain a preliminary decision result and the first-time remaining quantity; wherein the dual threshold refers to setting an upper threshold and a lower threshold, when the buckle installation quality score is higher than the upper threshold, the buckle is judged to be obviously qualified; when the buckle installation quality score is lower than the lower threshold, the buckle is judged to be obviously unqualified; when the buckle installation quality score falls between the upper threshold and the lower threshold, it is marked as pending for further analysis; The samples to be further analyzed and the defect heat map are input into the detailed analysis layer, the defect heat map is segmented and feature clustered, and the time-aware algorithm is used to complete the detailed defect analysis to obtain the intermediate decision result and the second time remaining amount; For samples with unclear intermediate decision results, the intermediate decision results, the buckle installation quality score and the defect heat map are input into the expert decision layer, and similarity matching is performed based on the historical case library to obtain high-level decision results; Decision fusion is performed on the preliminary decision result, the intermediate decision result and the advanced decision result to obtain a final decision result, and the final decision result is converted into a specific pipeline control instruction through a decision-instruction mapping table.

6. A real-time monitoring system for a buckle installation line, characterized in that: Used to implement the real-time monitoring method of the buckle installation assembly line according to any one of claims 1 to 5, the real-time monitoring system of the buckle installation assembly line comprises: an image acquisition module, used to perform image acquisition and multi-cascaded deep separable inverse convolution processing on the buckle installation area to obtain a multi-scale buckle feature image; An attention calculation module, used for performing spatial shuffling and bidirectional attention calculation on the multi-scale buckle feature image to obtain a buckle segmentation mask and a buckle type identifier; An extraction module, used to extract geometric and appearance features based on the buckle segmentation mask and the buckle type identifier, and calculate the normal interval deviation to obtain a buckle installation quality score and a defect heat map; An output module is used to perform fixed-time grading decision mechanism analysis based on the buckle installation quality score and the defect heat map, and output pipeline control instructions.

Citation Information

Patent Citations

  • Golden finger quality detection method, device, equipment and medium

    CN119048418A

  • Surface detection method and system for precise fastener

    CN119338827A

  • Visual inspection system of automobile door panel assembly

    CN119555684A