Real-time monitoring method and system for buckle installation assembly line

Through multi-angle image acquisition and multi-cascade depth separation inverse convolution processing, combined with spatial shuffling and rearrangement and bidirectional attention calculation, the complexity and real-time problems of snap installation assembly line monitoring in the prior art are solved, and accurate capture and real-time monitoring of snap features are achieved.

CN119941723AActive Publication Date: 2025-05-06XIAN QIAOLUMING TECH CO LTD

Patent Information

Application Number
CN202510421759.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing automotive decorative panel clamp installation assembly line monitoring methods have problems such as high complexity, accuracy and real-time difficulty, especially when dealing with small target clamp features and adapting to light changes.

Method used

Multi-angle image acquisition and original size retention technology are adopted, combined with multi-cascade depth, reverse convolution processing, spatial shuffling and rearrangement, and bidirectional attention calculation, multi-scale features of the snaps are extracted, and the significance of the snapshot profile features is enhanced through adaptive histogram equalization and edge enhancement processing.

Benefits of technology

It effectively captures snap features of different scales to meet the real-time monitoring needs of high-speed assembly lines, while maintaining the accuracy of decisions and improving the stability of the system under light changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941723A_ABST
    Figure CN119941723A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and discloses a real-time monitoring method and system for a buckle installation assembly line. The method comprises the following steps: carrying out image acquisition and multi-cascade depth separable reverse convolution processing on a buckle installation area to obtain a multi-scale buckle feature image; performing spatial shuffling rearrangement and bidirectional attention calculation on the multi-scale buckle feature image to obtain a buckle segmentation mask and a buckle type identifier; extracting geometric and appearance characteristics based on the buckle segmentation mask and the buckle type identifier, and calculating a normal interval deviation degree to obtain a buckle installation quality score and a defect thermodynamic diagram; and executing fixed time grading decision mechanism analysis based on the buckle installation quality score and the defect thermodynamic diagram, and outputting an assembly line control instruction. According to the method and the device, the buckle characteristics of different scales are effectively captured, the real-time monitoring requirement of a high-speed assembly line is met, and meanwhile, the accuracy of decision making is kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a real-time monitoring method and system for a buckle installation assembly line. Background Art

[0002] The installation quality of the buckles of automobile decorative panels has an important impact on the assembly quality and driving safety of the whole vehicle. The traditional manual inspection method has problems such as low efficiency and poor consistency, which makes it difficult to meet the production needs of modern high-speed assembly lines. With the continuous improvement of the intelligent level of the automobile manufacturing industry, automated monitoring technology based on image processing has emerged, but the current monitoring methods face multiple technical challenges: on the one hand, the supervised learning paradigm requires a large number of defective samples, while it is difficult to obtain samples of various types of buckle installation defects in actual production; on the other hand, although the anomaly detection paradigm does not require a large number of samples, it cannot accurately locate specific buckle defect types. In addition, the buckles on the decorative panels are small in size and have unclear features, which causes the existing image processing methods to have problems of target semantic distortion and background false alarms when applied.

[0003] Existing monitoring methods for the installation of automotive decorative panel buckles generally have the problems of high complexity and difficulty in balancing accuracy and real-time performance. Traditional methods mostly use fixed size adjustment and standard feature extraction processes, which cannot effectively handle small target buckle features; at the same time, the feature extraction and enhancement process has high computational complexity and is difficult to support the real-time monitoring needs of high-speed production lines. In addition, existing methods are not adaptable enough to changes in lighting and fluctuations in image quality, and are easily disturbed and misjudged in complex production environments. In particular, when the buckle image quality is poor, the recognition accuracy drops significantly, which cannot meet the reliability requirements of industrial production. Summary of the invention

[0004] The present application provides a real-time monitoring method and system for a buckle installation assembly line, which effectively captures buckle features of different scales, meets the real-time monitoring requirements of high-speed assembly lines, and maintains the accuracy of decision-making.

[0005] In a first aspect, the present application provides a method for real-time monitoring of a buckle installation assembly line, the method comprising: Perform image acquisition and multi-cascaded depth-separable inverse convolution processing on the buckle installation area to obtain a multi-scale buckle feature image; Performing spatial shuffling and rearrangement and bidirectional attention calculation on the multi-scale buckle feature image to obtain a buckle segmentation mask and a buckle type identifier; Extracting geometric and appearance features based on the buckle segmentation mask and the buckle type identifier, and calculating the normal interval deviation, to obtain a buckle installation quality score and a defect heat map; A fixed-time hierarchical decision mechanism analysis is performed based on the buckle installation quality score and the defect heat map, and a pipeline control instruction is output.

[0006] In a second aspect, the present application provides a real-time monitoring system for a buckle installation assembly line, the real-time monitoring system for a buckle installation assembly line comprising: An image acquisition module is used to perform image acquisition and multi-cascaded depth-separable inverse convolution processing on the buckle installation area to obtain a multi-scale buckle feature image; An attention calculation module, used for performing spatial shuffling and bidirectional attention calculation on the multi-scale buckle feature image to obtain a buckle segmentation mask and a buckle type identifier; An extraction module, used to extract geometric and appearance features based on the buckle segmentation mask and the buckle type identifier, and calculate the normal interval deviation to obtain a buckle installation quality score and a defect heat map; An output module is used to perform fixed-time grading decision mechanism analysis based on the buckle installation quality score and the defect heat map, and output pipeline control instructions.

[0007] In the technical solution provided by the present application, the multi-angle image acquisition and original size retention technology are used to avoid the loss of small target buckle features caused by image scaling, and the significance of buckle contour features is enhanced by adaptive histogram equalization and edge enhancement processing. The multi-cascaded deep separable reverse convolution structure significantly reduces the computational complexity, and constructs a receptive field pyramid through four cascade modules with increasing expansion rates to effectively capture buckle features of different scales. The spatial shuffling and rearrangement technology breaks the spatial limitations of traditional feature extraction, and the dynamic gain adjustment mechanism based on image quality and feature response intensity improves the stability of the system under changing lighting conditions. Few-sample prototype matching and two-way attention calculation solve the problem of insufficient buckle defect samples, and accurately capture buckle features and suppress background interference through prototype intensity downsampling and two-way attention mechanism. The multi-dimensional normal interval deviation calculation based on geometric and appearance features realizes the accurate quantification of buckle installation quality, and the defect heat map generated by gradient back propagation technology intuitively displays the defect location and severity. The fixed-time hierarchical decision-making mechanism innovatively solves the problem of unstable response time of the monitoring system. Through the three-level decision-making structure and time budget allocation, it ensures that the system can make decisions within the preset fixed time regardless of the complexity of the situation, meeting the real-time monitoring needs of high-speed assembly lines while maintaining the accuracy of the decision. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0009] Figure 1 It is a schematic diagram of an embodiment of a method for real-time monitoring of a buckle installation assembly line in an embodiment of the present application; Figure 2 It is a schematic diagram of an embodiment of a real-time monitoring system for a buckle installation assembly line in an embodiment of the present application. DETAILED DESCRIPTION

[0010] The embodiment of the present application provides a method and system for real-time monitoring of a buckle installation assembly line. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0011] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In the embodiment of the present application, an embodiment of the real-time monitoring method of the buckle installation assembly line includes: Step S101, performing image acquisition and multi-cascaded depth-separable inverse convolution processing on the buckle installation area to obtain a multi-scale buckle feature image; It is understandable that the execution subject of the present application can be a real-time monitoring system for buckle installation assembly line, or a terminal or a server, which is not limited here. The present application embodiment is described by taking a server as the execution subject as an example.

[0012] Specifically, an array of high-frame-rate industrial cameras is arranged on the buckle installation line. These cameras are arranged at multiple angles to ensure all-round monitoring of the buckle installation process of the decorative panel. The main camera is installed perpendicular to the surface of the decorative panel, and at least two auxiliary cameras are arranged at fixed angles of 15° to 75° to form a multi-view observation system. The camera array adopts a high-frame-rate acquisition mode of not less than 120 frames per second, and the image resolution reaches 1920×1080 pixels to ensure that the acquired original image has sufficient clarity and details. The image acquisition system transmits the collected image data to the image processing server at high speed through optical fiber to minimize signal delay and interference during transmission. In order to achieve synchronous control with the beat signal of the pipeline, a trigger is used to control the shooting timing of the camera, so that the camera accurately triggers the image acquisition process when the buckle reaches the preset position, and obtains a buckle image with time synchronization. And keep the original image size unchanged during the entire preprocessing process, effectively avoiding the problem of loss of small target buckle features caused by size scaling. In the image preprocessing stage, the time-series synchronous buckle image is subjected to adaptive histogram equalization processing. By adjusting the brightness distribution of image pixels, the overall brightness of the image is balanced, thereby eliminating the brightness inconsistency caused by uneven light sources or different reflective materials of decorative panels, and obtaining a brightness-corrected image. According to the noise level of the brightness-corrected image, the appropriate Gaussian filter kernel size is automatically selected for noise suppression processing. The Gaussian filter effectively removes high-frequency noise while retaining image details, and obtains a denoised buckle image. In the noise elimination, the denoised image is subjected to color normalization processing to minimize the influence of material and color differences of decorative panels from different batches, and a standardized image with uniform color is obtained. The standardized image is subjected to edge enhancement processing, and the improved Canny operator is used to accurately extract the edges of the decorative panels and the buckle contour features. The preprocessed buckle image is obtained by enhancing the contrast between the target and the background in the image. The preprocessed buckle image is subjected to multi-cascaded deep separable inverse convolution processing, and a feature extraction encoder composed of multiple deep separable inverse residual blocks is used for image processing. These residual blocks adopt the "expansion-compression-expansion" structure, expanding the image feature channel to six times the original through 1×1 convolution operation, extracting spatial features through 3×3 depthwise separable convolution, and finally compressing the number of channels back to the original dimension through 1×1 convolution. Compared with traditional convolution operations, depthwise separable convolution greatly reduces the amount of calculation and the number of model parameters, making the entire model more lightweight while maintaining efficient feature extraction capabilities. In order to capture snap-on features of different scales, the depthwise separable convolution in each cascade module uses different expansion rates, which are set to 1, 2, 4, and 8, respectively, to form a feature pyramid structure with a layer-by-layer expansion of the receptive field. Finally, a multi-scale snap-on feature image is obtained.

[0013] In this embodiment, the preprocessed buckle image is input into a 1×1 convolution operation module for channel expansion processing, and the number of feature channels of the input image is rapidly expanded to multiple times of the original dimension through 1×1 convolution, for example, to six times the number of original channels. The significance of channel expansion is that by increasing the dimension of the feature channel, a richer feature expression space is provided for subsequent deep convolution operations, thereby effectively enhancing the model's learning ability for complex image features. After the extended channel feature map is generated, the high-dimensional feature map is input into a deep separable reverse convolution network composed of four cascaded feature extraction modules. In these feature extraction modules, each module uses a 3×3 deep separable convolution operation with different expansion rates, and its expansion rates are set to 1, 2, 4, and 8 respectively. This design forms a feature pyramid structure with a layer-by-layer expansion of the receptive field, which can simultaneously capture buckle features at different scales. The difference between deep separable convolution and traditional convolution is that it significantly reduces the amount of calculation by decomposing the standard convolution into two-step operations of deep convolution and point-by-point convolution, especially in the case of multi-channel input, the computational complexity has been greatly optimized. After the hierarchical features of different receptive fields are extracted through depthwise separable convolution, these feature maps are again processed through channel compression through 1×1 convolution operation to obtain compressed feature maps. The high-dimensional information in the extended feature map is condensed into a more compact representation while retaining the most critical feature information to reduce the load of subsequent computing modules. In order to capture the dynamic change information during the buckle installation process, the compressed feature map of the current frame is calculated in the time domain with the compressed feature map of the previous frame. This process is achieved through the joint calculation of spatial displacement and channel similarity. The displacement features of the current frame and the previous frame in the spatial dimension are calculated, and the similarity of the features between channels is combined to obtain the time domain correlation feature map through a custom correlation function. The time domain correlation analysis can effectively capture the changes in the installation state of the buckle during the movement of the assembly line, especially when the buckle position is slightly offset or rotated, the model can still maintain a high detection accuracy. Based on the time domain correlation feature map, the compressed feature map is enhanced. By fusing the time domain features with the spatial features, the model uses the information of the time dimension to improve stability when judging the buckle installation state. The feature enhancement module combines convolution operations and nonlinear activation functions to enhance the buckle area in the feature space, highlighting the target features and suppressing background interference. In this process, the output features of the four cascaded feature extraction modules are pyramidally fused to form multi-level buckle features. The pyramid fusion process unifies feature maps of different scales to the same spatial dimension through interpolation operations, and then combines detail features and global features through layer-by-layer fusion to construct a multi-scale feature representation that can focus on detail features (such as the edge and shape of the buckle) while maintaining global information (such as the spatial position and installation angle of the buckle). The channel attention mechanism and spatial attention mechanism are used to adaptively allocate channel weights and regional weights to the multi-level buckle features.In the channel attention mechanism, global features are aggregated for each channel in the feature map through a global average pooling operation, and then the importance weight of each channel is calculated through a small multi-layer perceptron to generate a channel attention map. The attention map is weighted channel by channel with the original feature map, so that the model focuses on the feature channels that are most effective for buckle recognition. In the spatial attention mechanism, the convolution operation is combined with the activation function to generate the spatial attention map, and the response of a specific area in the image is emphasized by multiplying it pixel by pixel with the feature map. The spatial attention mechanism is suitable for suppressing the interference of the background area and highlighting the feature response of the buckle area. Combining the channel and spatial attention mechanisms, the model achieves a fine feature enhancement effect, significantly improves the saliency of the buckle feature, and finally obtains a multi-scale buckle feature image.

[0014] Step S102, performing spatial shuffling and rearrangement and bidirectional attention calculation on the multi-scale buckle feature image to obtain a buckle segmentation mask and a buckle type identifier; Specifically, the spatial dimensions of feature maps of different scales in the multi-scale snap feature image are uniformly processed. For example, bilinear interpolation is used to align feature maps of different scales to the same spatial resolution to obtain a unified feature map of the same spatial dimension. The feature maps of the unified spatial dimension are spliced ​​in the channel dimension to obtain a mixed feature map. The spliced ​​mixed feature map is subjected to spatial shuffling and rearrangement operations to break the spatial structure in the feature map and enhance the model's ability to perceive global information. Spatial shuffling breaks the original spatial position relationship by rearranging and transposing the spliced ​​feature map, so that the model obtains more comprehensive global information when learning features and eliminates the influence of certain spatial position deviations on feature extraction. After spatial shuffling, the obtained shuffled feature map is input into a 1×1 convolution operation module for processing. The 1×1 convolution is used for channel compression, and the useful information in the feature map is further learned by updating the parameters of the convolution kernel. The shuffled feature map is nonlinearly mapped through the sigmoid activation function so that the output of the feature map meets the actual target requirements. Each channel in the feature map is assigned different weights according to its response strength. In order to select the most effective features, a feature selection gating mechanism is introduced. This mechanism generates an attention map based on the response strength of each position in the feature map, and through the element-by-element product with the original feature map, the model pays more attention to the features that are important for buckle detection, while suppressing irrelevant or redundant features. This enables the model to adaptively assign different attention weights to each feature channel, thereby optimizing the feature expression ability and improving the accuracy of the model in practical applications. In the process of feature enhancement, the image quality score in the preprocessing stage is combined with the current feature response strength to calculate the comprehensive quality index. This index combines the quality of the image and the response strength of the current feature to provide the model with real-time feedback on the image quality. Based on the comprehensive quality index, the gain parameter is dynamically determined through a nonlinear mapping function. The dynamic gain parameter is used to adjust the enhancement strength of the feature map to cope with the processing requirements of images of different qualities. When the image quality is poor, the gain parameter will be increased to enhance the key information in the feature map; when the image quality is good, the gain parameter will be appropriately reduced to avoid excessive enhancement and information overload. Through the dynamic adaptive gain enhancement mechanism, the model can better adapt to changes in different production environments and effectively improve the detection accuracy. Through the gain-enhanced small target buckle features, the position and type of buckles can still be accurately identified when the image quality is not high or the buckle target is small. Few-sample prototype matching and bidirectional attention calculation are performed based on the enhanced small target buckle features. The feature prototype of each buckle is extracted from a small number of labeled buckle samples, and these prototype features are matched. During the matching process, the similarity calculation between feature vectors can help the model determine the buckle type in the current image. In order to improve the matching accuracy, the bidirectional attention mechanism is combined to calculate the attention in the spatial dimension and channel dimension respectively.By calculating the spatial attention map, the spatial regions related to buckles are highlighted, while by calculating the channel attention map, the most discriminative feature channels are identified. Buckle segmentation masks and buckle type identifiers are generated.

[0015] In this embodiment, a small number of labeled buckle installation sample features are averaged to generate prototype features for matching. By collecting a small number of samples (usually 1 to 5 samples per class) under each buckle installation state, and then extracting the feature vectors of these samples, these feature vectors are weighted averaged or directly averaged to obtain a prototype feature representation. The prototype features are subjected to intensity downsampling to reduce the impact of noise on the prototype features. During the intensity downsampling process, the response intensity distribution in the prototype features is analyzed, and the feature dimensions with significant responses are preferentially retained, and the feature dimensions with large noise interference are suppressed. This method extracts high-intensity signals through threshold filtering or feature selection algorithms to form a supporting prototype feature after noise reduction. Based on the enhanced small target buckle feature, three parallel branches with different scales of receptive fields are constructed to cope with the diversity of buckle sizes and positions. The model designs three parallel branches with different scales of receptive fields, each branch corresponds to a specific feature extraction module, and multi-scale feature capture is achieved through convolution kernels of different sizes and different expansion rates (such as 1, 2, 4). These parallel branches process both the detail features and global features in the input image, so that the model can maintain a high recognition accuracy when facing buckles of different sizes and shapes, and generate a multi-level query feature image. The cosine similarity between the multi-level query feature image and the denoised supporting prototype features is calculated to preliminarily determine the type and location of the buckle. Cosine similarity determines the similarity between the two by calculating the angle between the feature vectors. The preliminary feature matching results are obtained by calculation to identify the possible distribution of the buckle area. The preliminary matching results are input into the spatial attention and channel attention dual-dimensional calculations to improve the accuracy and stability of feature matching. In the spatial dimension, the spatial attention map is calculated, and the spatial attention weight of each pixel is obtained by multiplying the preliminary matching results with the spatial feature map and applying the softmax activation function, thereby highlighting the spatial area related to the buckle. In the channel dimension, the model calculates the channel attention vector through the channel response strength of the feature map, and generates the importance weight of each feature channel through global average pooling and multi-layer perceptron operations. The spatial attention map and the channel attention vector are fused to obtain the fused attention result. The fusion process combines the most important spatial regions and feature channels through pixel-by-pixel and channel-by-channel weighted processing, so that the model can more accurately identify the features of the buckle. The fusion operation is implemented through 1×1 convolution and sigmoid activation function, so that the output feature map not only contains the spatial position information of the buckle, but also integrates the feature selection results on the channel to form a comprehensive representation of the buckle features. The fused attention result is processed by the abnormal prior constraint module. This module maps the reasonable distribution range of the buckle into the feature map by establishing a priori knowledge model of the buckle installation position, shape and size. By constructing a probability model, the probability of misclassification of the background area is significantly reduced through prior information, and the buckle segmentation mask and buckle type identification are generated.

[0016] Step S103: extracting geometric and appearance features based on the buckle segmentation mask and the buckle type identifier, and calculating the normal interval deviation to obtain the buckle installation quality score and defect heat map; Specifically, morphological operations and edge extraction operations are performed on the buckle segmentation mask. Morphological operations help highlight the geometric shape features of the buckle through processing steps such as corrosion, expansion, opening operation and closing operation, while eliminating irregular background noise or small interference points to ensure the clarity of the buckle outline. The edge extraction operation can help the system accurately identify the morphological features of the buckle, such as geometric parameters such as area, perimeter, roundness, eccentricity, etc. These geometric features are extracted into feature vectors as geometric feature representations of the buckle. Based on the buckle type identification and the corresponding area of ​​the buckle segmentation mask, texture analysis and color statistics are performed to extract the appearance features of the buckle. Texture analysis calculates the texture information of the buckle surface, such as roughness, texture uniformity, etc., to reveal whether there are scratches or other damage on the buckle surface. Color statistics judge the appearance quality of the buckle by calculating indicators such as color uniformity and color difference of the buckle area. The appearance feature vector and the geometric feature vector form a complete feature vector of the target buckle. The difference between the buckle feature vector and the preset upper and lower bounds of the normal interval is calculated. The normal interval refers to the range in which the geometric and appearance features of the buckle should fall when the buckle is installed normally. For example, the roundness of the buckle should be close to 1, the eccentricity should be within the normal range, and the color and texture should be uniform. By comparing the difference between the value of each dimension in the buckle feature vector and the upper and lower bounds of the normal interval, the deviation distance of each feature dimension is calculated. Based on the deviation distance of each feature dimension, a weighted sum calculation is performed to obtain the buckle installation quality score. Different feature dimensions have different degrees of influence on the buckle quality. By setting the importance weight of each feature dimension, the calculation method of the deviation distance is adjusted. The feature vector gradient is calculated for the buckle installation quality score and back-propagated to the original image space. Through the back-propagation algorithm, the gradient information of each pixel is calculated based on the relationship between the buckle quality score and the deviation distance of each feature, reflecting the contribution of each area to the quality score. Through the back-propagation of the gradient, the quality score is matched with the specific area in the image, and then it is determined which areas have a greater impact on the buckle quality. The gradient map is up-sampled. Up-sampling aligns the heat map with the original image by enlarging the gradient map to the resolution of the original image. The areas shown in the heat map represent the parts of the quality score that deviate greatly from the normal range. These areas correspond to where defects occurred during the installation process.

[0017] Step S104: perform fixed-time grading decision mechanism analysis based on the buckle installation quality score and the defect heat map, and output pipeline control instructions.

[0018] Specifically, based on the buckle installation quality score and defect heat map, a hierarchical decision-making structure of direct judgment layer, detailed analysis layer and expert decision layer is established, and a fixed time slice is allocated to each layer to ensure the real-time and stability of the system in a high-speed assembly line production environment. Through this hierarchical decision-making mechanism, the system can quickly handle obviously qualified or unqualified buckle installation situations, analyze and expertly judge complex or ambiguous samples, and provide accurate assembly line control instructions. A hierarchical decision-making structure based on buckle installation quality score and defect heat map is constructed. The structure divides the decision-making process into three levels: direct judgment layer, detailed analysis layer and expert decision layer. Each level has a fixed time slice allocation to ensure that the total processing time of the entire decision-making process is constant and meet the real-time processing requirements of high-speed assembly lines. The direct judgment layer is responsible for quickly processing those buckle installation samples that are obviously qualified or obviously unqualified. For these samples, decisions are made by quickly comparing the buckle installation quality score with dual thresholds. If the buckle installation quality score is higher than the set upper threshold, it is directly judged as qualified; if it is lower than the lower threshold, it is judged as unqualified; if the quality score falls between the two thresholds, it is marked for further analysis. The decision-making process at this level is very efficient, quickly screening out most of the simple buckle installation samples, obtaining preliminary decision results, and calculating the first time remaining. For the buckle installation samples within the threshold range, they are sent to the detailed analysis layer for further analysis. At this layer, the buckle to be analyzed is segmented and clustered using defect heat maps to identify and analyze defects more finely. This process uses a time-aware algorithm that dynamically adjusts the level of detail of the analysis based on the current production rhythm and processing capacity to ensure that the fine analysis of defects is completed within the specified time slice. Through this layer of processing, complex buckle installation defects are analyzed, intermediate decision results are obtained, and the second time remaining is calculated. For particularly complex or ambiguous samples, the intermediate decision results, buckle installation quality scores, and defect heat maps are input into the expert decision layer for final judgment. At this level, the system relies on the historical case library to perform similarity matching, and combines previous experience and cases for comprehensive judgment to obtain advanced decision results. The historical case library contains a large number of verified buckle installation samples. Through similarity matching, the most likely installation state of the current sample is inferred with the help of historical data, so as to make more accurate decisions. After multiple levels of decision-making at the direct judgment layer, detailed analysis layer, and expert decision layer, the preliminary decision results, intermediate decision results, and advanced decision results are integrated to obtain the final decision result. The final decision result is converted into specific assembly line control instructions through the decision-instruction mapping table, including operations such as continuing production, suspending production, adjusting production parameters, and starting maintenance.

[0019] In the embodiment of the present application, multi-angle image acquisition and original size retention technology are used to avoid the loss of small target buckle features caused by image scaling, and the significance of buckle contour features is enhanced by adaptive histogram equalization and edge enhancement processing. The multi-cascaded deep separable reverse convolution structure significantly reduces the computational complexity, and constructs a receptive field pyramid through four cascade modules with increasing expansion rates to effectively capture buckle features of different scales. The spatial shuffling and rearrangement technology breaks the spatial limitations of traditional feature extraction, and the dynamic gain adjustment mechanism based on image quality and feature response intensity improves the stability of the system under changing lighting conditions. Few-sample prototype matching and two-way attention calculation solve the problem of insufficient buckle defect samples, and accurately capture buckle features and suppress background interference through prototype intensity downsampling and two-way attention mechanism. The multi-dimensional normal interval deviation calculation based on geometric and appearance features realizes the accurate quantification of buckle installation quality, and the defect heat map generated by gradient back propagation technology intuitively displays the defect location and severity. The fixed-time hierarchical decision-making mechanism innovatively solves the problem of unstable response time of the monitoring system. Through the three-level decision-making structure and time budget allocation, it ensures that the system can make decisions within the preset fixed time regardless of the complexity of the situation, meeting the real-time monitoring needs of high-speed assembly lines while maintaining the accuracy of the decision.

[0020] In a specific embodiment, the process of executing step S101 may specifically include the following steps: (1) Capture images of the buckle installation area from multiple angles to obtain the original buckle image; (2) Based on the pipeline beat signal, the original snap-in image is synchronously triggered and controlled, and the original size is kept unchanged to obtain a time-series synchronized snap-in image; (3) Adaptive histogram equalization is performed on the time-series synchronization snap-in image to obtain a brightness-corrected image; (4) Determine the Gaussian filter kernel based on the noise level of the brightness correction image, perform noise suppression processing, obtain a denoised snap-in image, perform color normalization processing on the denoised snap-in image, and obtain a standardized image with uniform color; (5) Perform edge enhancement processing on the standardized image with uniform color to obtain a preprocessed snap-in image; (6) Perform multi-cascaded depth-wise separable inverse convolution processing on the preprocessed buckle image to obtain a multi-scale buckle feature image.

[0021] Specifically, multi-angle image acquisition is performed on the buckle installation area to obtain sufficient image information and fully cover all aspects of the buckle. By configuring multiple industrial cameras in the buckle installation area and arranging these cameras at different angles, the buckle can be captured from various perspectives during the assembly process of the assembly line. In order to obtain high-quality original images, a high-frame rate camera is selected, and the camera acquisition frequency is set to more than 120 frames per second to capture every buckle detail on the high-speed assembly line in real time. The images acquired from multiple angles can ensure that the system can accurately judge the state of the buckle from all directions, thereby avoiding image distortion or incompleteness caused by a single angle. Based on the assembly line beat signal, the original buckle image is synchronously triggered and controlled to ensure the timing consistency of the image and obtain a buckle image with timing synchronization. The core of synchronous control is to use the assembly line beat signal to trigger the camera's image acquisition, and by synchronizing with the assembly line beat, ensure that the image of each buckle is accurately captured at the specified time point. During synchronous acquisition, the image size is guaranteed not to change, and each acquired image maintains the same resolution and size as the original image, avoiding the loss of key information when the image is scaled or cropped. The time-synchronized snap-on images are subjected to brightness correction. Adaptive histogram equalization is an image enhancement technology that can effectively improve the brightness distribution of an image. With this technology, the brightness distribution of an image can be adaptively adjusted so that both the dark and bright parts of the image can be displayed more clearly. For the processing of snap-on images, adaptive histogram equalization can eliminate the brightness deviation caused by the pipeline environment or uneven lighting conditions to obtain a brightness-corrected image. The Gaussian filter kernel is determined based on the noise level of the brightness-corrected image to perform noise suppression. The Gaussian filter effectively smoothes the image and reduces the impact of noise by performing convolution operations on the image. The size of the filter kernel is dynamically determined according to the noise level of the image. The larger the noise, the larger the size of the filter kernel is, so as to better smooth the image. The denoised snap-on images are subjected to color normalization processing to effectively reduce the impact of color and material differences on image processing, so that snap-on images of different batches can be compared and analyzed in the same color space. Color normalization calculates the color mean and standard deviation of the image, and then uses a standardization method to adjust the color distribution of the image so that all images have consistent color characteristics. This embodiment can effectively reduce the differences in color and material between different buckle images. Edge enhancement processing is performed on standardized images with uniform colors to improve the clarity of the buckle edges and make the buckle boundaries more prominent. Edge enhancement uses an improved Canny operator. Through this operator, the contour features of the buckle are extracted and the significance of these features in the image is improved. The Canny operator determines the edge position in the image through multiple convolution operations and threshold judgments, and sharpens and enhances the edge information in the image. The enhanced buckle image has a higher contrast. The pre-processed buckle image is subjected to depth separable convolution processing.Depthwise separable convolution is an efficient convolution operation that greatly reduces the amount of computation by decomposing the standard convolution into two steps: depthwise convolution and then pointwise convolution. Depthwise separable convolution can efficiently extract multi-scale features of buckles, especially when the buckles are small in size, it can more accurately identify the detailed features of buckles. This process extracts features of each scale through multiple cascaded convolution modules, and captures feature information such as the shape, edge, and texture of the buckles at different receptive field scales by setting different expansion rates (such as 1, 2, 4, 8, etc.). Through these operations, multi-scale buckle feature images are obtained. In the multi-cascaded depthwise separable convolution processing, considering the different scales of the receptive field, the expansion rate of the receptive field is defined for each module. Assume that the expansion rate of each convolution module is. , through convolution operations with different dilation rates, feature maps of different scales are captured.,Through different convolution kernels and dilation rates, the buckle features are extracted at multiple scales, ensuring that small buckle features in the image are not ignored.

[0022] In a specific embodiment, the execution step performs multi-cascade depth-separable inverse convolution processing on the pre-processed snap-in image to obtain a multi-scale snap-in feature image may specifically include the following steps: (1) Perform channel expansion on the preprocessed snap-in image through 1×1 convolution to obtain an extended channel feature map; (2) The extended channel feature map is input into four cascade feature extraction modules. Each cascade feature extraction module uses a 3×3 depthwise separable convolution with dilation rates of 1, 2, 4, and 8 to obtain hierarchical features with different receptive fields. (3) Perform channel compression processing on the hierarchical features of different receptive fields through 1×1 convolution to obtain a compressed feature map; (4) Perform spatial displacement and channel similarity calculation on the compressed feature map of the current frame and the compressed feature map of the previous frame to obtain a time domain related feature map that captures the dynamic changes of the buckle installation; (5) Based on the time-domain correlation feature map, the compressed feature map is enhanced, and the output features of the four cascade feature extraction modules are fused in a pyramidal manner to obtain multi-level buckle features; (6) The channel attention mechanism and spatial attention mechanism are used to adaptively allocate channel weights and regional weights to the multi-level snap features to obtain a multi-scale snap feature image.

[0023] Specifically, the pre-processed snap-in image is subjected to channel expansion processing through 1×1 convolution. The role of convolution is to expand the number of feature channels of the image, increase the dimension of feature expression, and enhance the representation ability of the network, so that more diverse feature information can be learned. In the convolution operation, 1×1 convolution does not change the spatial dimension of the image, but only changes the number of channels. This embodiment can obtain a new feature map with more channels. The extended channel feature map is input into four cascaded feature extraction modules. Each module uses 3×3 depth separable convolution for feature extraction, and the expansion rate of each module is set to 1, 2, 4 and 8 respectively to capture the features of different receptive fields. Depth separable convolution is different from traditional convolution, and the convolution operation is decomposed into two steps: first, depth convolution is performed, and then point-by-point convolution is performed. This structure can effectively reduce the amount of calculation and improve the operational efficiency of the model, especially in tasks that require the extraction of a large number of features. By setting different expansion rates, the features extracted by each module have different receptive fields, which means that convolution operations with different expansion rates can capture snap-on features of different scales. The larger the receptive field, the convolution operation can focus on a wider area in the image, thereby improving the recognition ability of large-scale features, while a small receptive field is suitable for capturing detail features. After four cascaded convolution modules, the feature maps of different receptive fields are subjected to channel compression. Convolution reduces the number of channels, removes redundant feature information and maintains the core features of the image. Through the compression process, the dimension of the feature map is reduced while maintaining the most important information, reducing the amount of calculation and improving efficiency. In order to capture the dynamic changes in the buckle installation process, the compressed feature map of the current frame is calculated with the compressed feature map of the previous frame for time domain correlation. Time domain correlation analysis captures dynamic changes in the image by comparing the feature differences between adjacent frames. This embodiment can effectively capture the dynamic change information of the buckle and help analyze the motion trajectory and state changes of the buckle on the assembly line. Based on the time domain correlation feature map, the compressed feature map is enhanced to enhance the significance of important features while suppressing the interference of irrelevant information. Enhancement is achieved through convolution operations and activation functions, specifically by adjusting the weights to strengthen the features related to the buckle state to obtain an enhanced feature map. In the multi-cascade convolution module, the output features of all modules are pyramidally fused by fusing feature maps of different scales layer by layer. Pyramid fusion unifies feature maps of different scales to the same spatial dimension through interpolation and convolution operations, and then performs weighted averaging or splicing to obtain a multi-level buckle feature image. On the fused multi-level snap feature image, an attention mechanism is introduced to optimize the expression of features. The channel attention mechanism and the spatial attention mechanism are used to adaptively allocate the channel weights and regional weights of the features respectively. The channel attention mechanism calculates the importance of each channel through global average pooling and multi-layer perceptron, while the spatial attention mechanism emphasizes the key areas by calculating the response intensity of each spatial position in the feature map. By weighting the attention of channels and spaces, the multi-level snap feature image is converted into a high-quality multi-scale snap feature map.

[0024] In a specific embodiment, the process of executing step S102 may specifically include the following steps: (1) Performing unified spatial dimension processing on feature maps of different scales in the multi-scale buckle feature image to obtain a unified feature map of the same spatial dimension; (2) The unified feature maps of the same spatial dimension are concatenated in the channel dimension to obtain a mixed feature map, and the mixed feature map is spatially shuffled and rearranged to obtain a shuffled feature map; (3) The shuffled feature map is input into 1×1 convolution and sigmoid activation function, and the effective features are screened through the feature selection gating mechanism to obtain the spatial shuffled features with attention weights; (4) Calculate the comprehensive quality index based on the image quality score in the preprocessing stage and the current feature response strength, and determine the dynamically adaptive gain parameter through nonlinear mapping; (5) According to the dynamically adaptive gain parameter, the spatial shuffled features with attention weights are subjected to gain enhancement processing to obtain enhanced small target snap-in features; (6) Based on the enhanced small target buckle features, few-shot prototype matching and bidirectional attention calculation are performed to obtain the buckle segmentation mask and buckle type identification.

[0025] Specifically, the snap feature images of different scales are processed uniformly in terms of spatial dimensions. Interpolation methods, such as bilinear interpolation, are used to make the spatial dimensions of feature maps of all scales consistent, so as to obtain a unified feature map of the same spatial dimension. The unified feature maps of the same spatial dimension are spliced ​​in the channel dimension to fuse information from multiple scales to form a mixed feature map containing more features. The mixed feature map is spatially shuffled and rearranged to disrupt the spatial structure of the image, so that the model can learn global information without being restricted by local information. Spatial shuffling helps to eliminate position dependence in the image, prompting the model to better capture global features and improve sensitivity to information in different regions, thereby obtaining a shuffled feature map. The shuffled feature map is input into a 1×1 convolution and activation function for processing. The 1×1 convolution is used to perform weighted summation in the channel dimension, so that the features of each channel can be processed and refined. Through the nonlinear transformation of the activation function, the expression of important channels in the feature map is enhanced and the influence of unimportant channels is suppressed. The features are screened through the feature selection gating mechanism to select the most effective features. The function of the feature selection gating mechanism is to dynamically adjust the weight of each feature channel based on the response strength of each feature. Through this mechanism, the features most relevant to the buckle state are focused on, thereby improving the accuracy and efficiency of subsequent processing. The process of feature selection helps to focus the system's attention on important features and avoid interference from background noise. Based on the image quality score in the preprocessing stage and the current feature response strength, a comprehensive quality index is calculated. The image quality score reflects the clarity, contrast and signal-to-noise ratio of the image, while the feature response strength characterizes the importance of the buckle feature. The calculation of the comprehensive quality index helps the system adjust under different quality conditions to ensure that it can adapt to environmental changes in real time. Based on the comprehensive quality index, the gain parameter is dynamically adjusted through a nonlinear mapping function. The adjustment of the gain parameter is determined based on the current image quality and feature response strength, with the purpose of enhancing the important features in the image and reducing the interference of irrelevant information. Through mapping, the gain is dynamically adjusted according to different environmental conditions, thereby improving the detection accuracy of the buckle feature. After obtaining the dynamic gain parameter, the spatial shuffled features with attention weights are subjected to gain enhancement. The strength of the feature is adjusted according to its importance to make the key features of the buckle more prominent. By applying the gain to the feature map, the features related to the buckle are strengthened and the influence of background noise is reduced, thereby obtaining an enhanced feature map.

[0026] Based on the enhanced small target buckle features, few-shot prototype matching and bidirectional attention calculation are performed. Few-shot prototype matching helps the system determine the current buckle status by comparing with historical samples. The bidirectional attention mechanism adaptively adjusts the weights in the spatial and channel dimensions, allowing the system to more accurately identify buckle features. Through this step, accurate buckle segmentation masks and buckle type identifications are obtained.

[0027] In a specific embodiment, the process of performing few-shot prototype matching and bidirectional attention calculation based on enhanced small-target buckle features to obtain buckle segmentation masks and buckle type identifiers may specifically include the following steps: (1) Calculate the mean of a small number of labeled buckle installation sample features to obtain prototype features, and perform intensity downsampling on the prototype features to obtain noise-reduced supporting prototype features; (2) Based on the enhanced small target buckle features, three parallel branches with different scales of receptive fields are constructed to obtain a multi-level query feature image; (3) Calculate the cosine similarity between the multi-level query feature image and the denoised support prototype feature to obtain a preliminary matching result; (4) Input the preliminary matching results into the dual-dimensional calculation of spatial attention and channel attention to obtain the spatial attention map and channel attention vector respectively, and fuse the spatial attention map and channel attention vector to obtain the fused attention result; (5) The fused attention results are processed by the abnormal prior constraint module, and the background area interference is suppressed according to the prior knowledge of the buckle installation position, shape and size to obtain the buckle segmentation mask and buckle type identification.

[0028] Specifically, prototype features are obtained by averaging a small number of labeled buckle installation sample features, which provide a standard reference. The goal of the denoising operation is to remove high-frequency noise and reduce unnecessary details, making the supporting prototype features more stable and easy to use in the future. Based on the enhanced small target buckle features, three parallel branches with receptive fields of different scales are constructed to obtain multi-level query feature images. By processing receptive fields of different scales in parallel, the diversity and detail information of buckle features are captured. Through parallel calculation, a multi-level feature image is obtained. The cosine similarity between the multi-level feature image and the denoised supporting prototype features is calculated to obtain a preliminary matching result. Cosine similarity is a standard method to measure the similarity of two vectors. The degree of similarity between the two vectors is measured by calculating the angle between the two vectors. Through cosine similarity, the similar areas between the current buckle and the standard buckle are accurately found. The preliminary matching results are input into the dual-dimensional calculation of spatial attention and channel attention to optimize the feature map. The spatial attention mechanism focuses on the importance of each spatial position in the image, and the spatial attention map represents the weight of each spatial position, while the channel attention mechanism focuses on the importance of each channel in the feature map, and the channel attention vector represents the weight of each channel. The spatial and channel attention mechanisms are calculated through global pooling and convolution operations respectively, and finally merged into a fused attention result. The fused attention result is processed by the abnormal prior constraint module. This module helps the system suppress the interference of the background area by combining the prior knowledge of the installation position, shape and size of the buckle. The position, shape and size of the buckle are important features in buckle detection. Using this prior knowledge, background noise can be better eliminated and the actual features of the buckle can be focused. Through the prior constraints, the final buckle segmentation mask and buckle type identification are obtained.

[0029] Among them, the mean value of a small number of labeled buckle installation sample features is calculated to obtain the prototype features:

[0030] in: Represents the original prototype features; represents the feature representation of the k-th support sample; Indicates the total number of annotated support samples; in the buckle installation pipeline, the features of a small number of buckle samples that are known to be correctly installed are extracted and averaged to form a standard feature template as a benchmark for subsequent matching.

[0031] After obtaining the enhanced small target buckle feature, the cosine similarity calculation is performed with the prototype feature:

[0032] in: represents the cosine similarity graph; represents the query feature map (i.e., the enhanced small target snap feature); represents the supporting prototype feature after denoising; (x, y) represents the spatial position coordinates; c represents the feature channel index. In practical applications, this formula is used to evaluate the similarity between the currently detected buckle feature and the standard prototype. The higher the similarity, the closer the buckle installation state is to the standard state. After fusing spatial attention and channel attention, apply the abnormal prior constraint:

[0033] in: represents the final snap segmentation mask; represents the attention result after fusion; Representing prior constraints on the location, shape, and size of the buckle; Represents a constraint function, which is used to suppress background area interference. In the buckle installation monitoring system, the prior knowledge of the buckle (such as installation position, normal size range, etc.) is used to correct the attention result, filter out the false detection area that does not meet the expectations, and obtain accurate buckle segmentation mask and type identification.

[0034] In a specific embodiment, the process of executing step S103 may specifically include the following steps: (1) Perform morphological operations and edge extraction on the buckle segmentation mask to obtain the buckle geometric feature vector; (2) Perform texture analysis and color statistics based on the corresponding areas of the buckle type identification and the buckle segmentation mask to obtain the buckle appearance feature vector; (3) Merge the buckle geometry feature vector and the buckle appearance feature vector to obtain the target buckle feature vector; (4) Calculate the difference between the buckle feature vector and the preset upper and lower bounds of the normal interval to obtain the deviation distance of each feature dimension; (5) Perform a weighted sum calculation based on the deviation distance and feature importance weight of each feature dimension to obtain the buckle installation quality score; (6) Calculate the feature vector gradient of the buckle installation quality score and back-propagate it to the original image space, and perform upsampling to obtain the defect heat map.

[0035] Specifically, morphological operations and edge extraction are performed on the buckle segmentation mask. Morphological operations help extract the structural features of the buckle by performing operations such as dilation, erosion, and opening operations on the image. Through morphological processing, the edges and contours of the buckle are more prominent, thereby better describing the geometric characteristics of the buckle. Edge extraction is performed to obtain the buckle geometric feature vector, including information such as the shape, size, and contour of the buckle. The appearance features of the buckle are extracted. Texture analysis and color statistics are performed based on the corresponding areas of the buckle type identification and the buckle segmentation mask. Texture analysis captures the detailed changes on the buckle surface by calculating the distribution characteristics of the buckle surface pattern, and color statistics analyzes the color distribution on the buckle surface. Through these analyses, the appearance features of the buckle are obtained. The geometric feature vector and the appearance feature vector of the buckle are merged to obtain the complete feature vector of the target buckle. The merging operation forms a comprehensive feature vector containing multiple aspects of information such as buckle shape, size, color, and texture by splicing the geometric feature vector and the appearance feature vector in the feature dimension. The difference between the target buckle feature vector and the preset upper and lower bounds of the normal interval is calculated to obtain the deviation distance of each feature dimension. By comparing with the normal range, evaluate whether the buckle meets the standard. By calculating the deviation distance of all feature dimensions, the buckle deviation is obtained, so as to determine whether the buckle has defects. Based on the deviation distance of each feature dimension and the feature importance weight, a weighted sum calculation is performed to obtain the buckle installation quality score. Taking into account the importance of each feature dimension, the overall quality score of the buckle is calculated. Through this score, the installation quality of the buckle is quantitatively evaluated, and the buckles on the assembly line are monitored in real time. Calculate the feature vector gradient of the buckle installation quality score and back-propagate the gradient to the original image space. Through gradient back-propagation, the influence of the quality score is transferred back to each pixel of the image, so as to calibrate the key areas in the buckle image. By upsampling the feature vector gradient, a high-resolution defect heat map is obtained.

[0036] Among them, for the obtained target buckle feature vector, it is necessary to calculate its deviation distance from the upper and lower limits of the preset normal interval. The preset normal interval is obtained based on a large amount of historical data statistics and represents the characteristic range of the normal installation buckle. The calculation formula for the deviation distance is:

[0037] in: Represents the eigenvector The deviation distance of the dimension; The first character vector of the target buckle Dimension value; The first one represents the lower limit of the preset normal interval. Dimension value; The upper limit of the preset normal interval Dimension value. This formula calculates the degree of deviation of the buckle feature from the normal interval in each dimension. When the eigenvalue falls within the normal interval, the deviation distance is zero; when the eigenvalue exceeds the upper bound, the deviation distance is the difference between the eigenvalue and the upper bound; when the eigenvalue is lower than the lower bound, the deviation distance is the difference between the lower bound and the eigenvalue. Then, the feature importance weight is needed to weight the sum of the deviation distances to obtain the final buckle installation quality score. The feature importance weight reflects the degree of influence of different features on the buckle installation quality, which is usually set by domain experts or automatically learned through machine learning methods. The calculation formula for the buckle installation quality score is:

[0038] in: Indicates the buckle installation quality score, with a full score of 100 points; Indicates Importance weight of dimension features; Represents the total number of dimensions of the feature vector; represents the scaling factor used to adjust the sensitivity of the score; Represents the penalty exponent, usually set to 2, giving more severe penalties for larger deviations.

[0039] In actual application scenarios, feature dimensions may include geometric features of buckles (such as area, perimeter, aspect ratio, roundness, etc.) and appearance features (such as color mean, texture gradient, surface smoothness, etc.). For example, for buckle installation monitoring on an automotive parts assembly line, the importance weight may give higher weights to key geometric features such as position offset and angle tilt, while giving lower weights to minor features such as color change.

[0040] In order to more accurately express the mutual influence between different features, the feature correlation matrix can be introduced for correction to obtain an improved quality score calculation formula:

[0041] in: Representation characteristics and Features The correlation coefficient between ; Represents the correlation influencing factor, which controls the degree of influence of the correlation on the final score. In this way, the buckle installation quality score not only considers the independent deviation degree of each feature, but also considers the interaction between the features, making the scoring result more comprehensive and accurate. The higher the score, the better the buckle installation quality, and the lower the score, the possible quality problem. In the actual application of the buckle installation assembly line, this scoring mechanism can quickly identify poorly installed buckles, and guide the subsequent defect heat map generation and assembly line control decision, effectively ensuring product quality and production efficiency.

[0042] In a specific embodiment, the process of executing step S104 may specifically include the following steps: (1) Based on the buckle installation quality score and defect heat map, a hierarchical decision structure consisting of direct judgment layer, detailed analysis layer, and expert decision layer is established. A fixed time slice is allocated to each layer of the hierarchical decision structure to obtain a fixed-time decision framework with a constant total processing time. (2) In the direct judgment layer, the buckle installation quality score is quickly compared with the double threshold value, and the decision result is directly output for the obviously qualified or obviously unqualified samples. The samples within the threshold range are marked as waiting for further analysis, and the preliminary decision result and the first-time remaining quantity are obtained; (3) Input the samples and defect heat maps to be further analyzed into the detailed analysis layer, perform regional segmentation and feature clustering on the defect heat maps, use the time-aware algorithm to complete the detailed defect analysis, and obtain the intermediate decision results and the second time residual; (4) For particularly complex samples, the intermediate decision results, buckle installation quality scores, and defect heat maps are input into the expert decision layer, and similarity matching is performed based on the historical case library to obtain high-level decision results; (5) Perform decision fusion on the preliminary decision results, intermediate decision results, and advanced decision results to obtain the final decision result, and convert the final decision result into a specific pipeline control instruction through a decision-instruction mapping table.

[0043] Specifically, the hierarchical decision-making mechanism based on the buckle installation quality score and defect heat map is a time-sensitive processing framework. By establishing a three-layer structure: direct judgment layer, detailed analysis layer, and expert decision layer, each layer is responsible for judgment tasks of different complexity. To ensure real-time performance, each layer is allocated a fixed time slice to implement a decision framework with a constant total processing time. For example, within the total processing time limit of 60 milliseconds, the direct judgment layer is allocated 20 milliseconds, the detailed analysis layer is allocated 30 milliseconds, and the expert decision layer is allocated 10 milliseconds. This time allocation ensures that buckle detection can keep up with the rhythm of the assembly line and avoid production stagnation due to slow processing.

[0044] In the direct judgment layer, a dual-threshold quick comparison is performed on the buckle installation quality score. The dual threshold refers to setting an upper threshold and a lower threshold to divide the buckle quality into three categories. When the score is higher than the upper threshold (such as a score of 95 points and an upper threshold of 90 points), the buckle is judged to be obviously qualified; when the score is lower than the lower threshold (such as a score of 60 points and a lower threshold of 75 points), the buckle is judged to be obviously unqualified; when the score falls between the two thresholds (such as a score of 85 points), it is marked for further analysis. The processing speed of this layer is extremely fast, and it usually takes only a few milliseconds to complete the judgment. The remaining time (the remaining amount of the first time) is passed to the next layer to improve overall efficiency.

[0045] The samples to be further analyzed and the defect heat map are input to the detailed analysis layer for further processing. Region segmentation refers to the pixel-level division of the heat map to extract abnormal areas. This process uses an adaptive threshold segmentation algorithm to divide the image into foreground (defective area) and background (normal area) according to the intensity value of the heat map. Then feature clustering is performed, that is, the extracted areas are grouped according to characteristics such as position, shape, and intensity to identify defects with similar characteristics. The time-aware algorithm refers to an algorithm that automatically adjusts the analysis accuracy within a limited time, prioritizes important features, and gradually reduces the computational complexity as time consumption increases to ensure that the processing is completed within the allocated time slice. After this layer is completed, the intermediate decision results and the second time remainder are obtained. For particularly complex samples, the intermediate decision results, the buckle installation quality score, and the defect heat map are jointly input into the expert decision layer. This layer performs similarity matching based on the historical case library, which contains a large amount of known buckle state data and its corresponding processing results. Similarity matching uses weighted Euclidean distance to calculate the similarity between the current sample and the historical case, and selects the closest case as a reference. This experience-based decision-making method can cope with complex situations where boundaries are fuzzy and rules are difficult to define, and improve the system's ability to handle abnormal situations.

[0046] The decision fusion is performed on the three-layer decision results to obtain the final decision result. The decision fusion adopts the weighted voting method, assigns different weights according to the credibility of each layer of results, and comprehensively obtains the final judgment. Through the pre-set decision-instruction mapping table, the decision results are converted into specific assembly line control instructions, such as continuing production, pausing inspection, adjusting parameters, etc., to achieve precise control of the production line.

[0047] Take an actual buckle installation monitoring scenario as an example: a fixed buckle on a car door panel is installed through an assembly line. After the camera collects the buckle installation image, it enters the hierarchical decision system. First, in the direct judgment layer, the system calculates the buckle installation quality score as 82 points, which is between the lower threshold of 75 points and the upper threshold of 90 points. Therefore, it is marked for further analysis and passed to the detailed analysis layer. In the detailed analysis layer, the system performs regional segmentation on the defect heat map and finds that there is a small area of ​​high intensity on the left edge of the buckle in the heat map, which is about 3% of the total area of ​​the buckle. This area is identified as a potential defect of the "edge is not fully embedded" type through feature clustering. Because this situation is not clear enough, the sample enters the expert decision layer. The system finds 7 similar cases in the historical case library, of which 5 cases have loose buckles during use. Based on this matching result, combined with the analysis of the first two layers, the system finally determines that the buckle is "potentially risky and needs to be adjusted", and generates the corresponding assembly line control instructions, suspends the current station and prompts the operator to readjust the buckle installation parameters. This judgment process took a total of 58 milliseconds, which did not exceed the preset time limit of 60 milliseconds, ensuring the continuous operation of the assembly line.

[0048] The above describes the real-time monitoring method of the buckle installation assembly line in the embodiment of the present application. The following describes the real-time monitoring system of the buckle installation assembly line in the embodiment of the present application. Figure 2 In the embodiment of the present application, an embodiment of the real-time monitoring system of the buckle installation assembly line includes: An image acquisition module 201 is used to perform image acquisition and multi-cascade depth-separable inverse convolution processing on the buckle installation area to obtain a multi-scale buckle feature image; An attention calculation module 202 is used to perform spatial shuffling and re-arrangement and bidirectional attention calculation on the multi-scale buckle feature image to obtain a buckle segmentation mask and a buckle type identifier; An extraction module 203 is used to extract geometric and appearance features based on the buckle segmentation mask and the buckle type identifier, and calculate the normal interval deviation to obtain the buckle installation quality score and defect heat map; The output module 204 is used to perform fixed-time grading decision mechanism analysis based on the buckle installation quality score and the defect heat map, and output the pipeline control instructions.

[0049] Through the synergy of the above components, the multi-angle image acquisition and original size retention technology are used to avoid the loss of small target buckle features caused by image scaling, and the buckle contour features are enhanced by adaptive histogram equalization and edge enhancement processing. The multi-cascaded deep separable inverse convolution structure significantly reduces the computational complexity. The receptive field pyramid is constructed through four cascade modules with increasing expansion rates to effectively capture buckle features of different scales. The spatial shuffle permutation technology breaks the spatial limitations of traditional feature extraction, and the dynamic gain adjustment mechanism based on image quality and feature response intensity improves the stability of the system under changing lighting conditions. Few-sample prototype matching and bidirectional attention calculation solve the problem of insufficient buckle defect samples. The buckle features are accurately captured and background interference is suppressed through prototype intensity downsampling and bidirectional attention mechanism. The multi-dimensional normal interval deviation calculation based on geometric and appearance features realizes the accurate quantification of buckle installation quality. At the same time, the defect heat map generated by gradient back propagation technology intuitively displays the defect location and severity. The fixed-time hierarchical decision-making mechanism innovatively solves the problem of unstable response time of the monitoring system. Through the three-level decision-making structure and time budget allocation, it ensures that the system can make decisions within the preset fixed time regardless of the complexity of the situation, meeting the real-time monitoring needs of high-speed assembly lines while maintaining the accuracy of the decision.

[0050] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0051] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a real-time monitoring device for a buckle installation assembly line (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.

[0052] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A real-time monitoring method for a buckle installation line, characterized in that: The buckle installation line real-time monitoring method comprises: Perform image acquisition and multi-cascaded depth-separable inverse convolution processing on the buckle installation area to obtain a multi-scale buckle feature image; Performing spatial shuffling and rearrangement and bidirectional attention calculation on the multi-scale buckle feature image to obtain a buckle segmentation mask and a buckle type identifier; Extracting geometric and appearance features based on the buckle segmentation mask and the buckle type identifier, and calculating the normal interval deviation, to obtain a buckle installation quality score and a defect heat map; A fixed-time hierarchical decision mechanism analysis is performed based on the buckle installation quality score and the defect heat map, and a pipeline control instruction is output.

2. The method for real-time monitoring of a buckle installation line according to claim 1, characterized in that: The method of performing image acquisition and multi-cascaded depth-separable inverse convolution processing on the buckle installation area to obtain a multi-scale buckle feature image includes: Capture images of the buckle installation area from multiple angles to obtain an original buckle image; Based on the pipeline beat signal, the original snap image is synchronously triggered and controlled, and the original size is kept unchanged to obtain a time-series synchronous snap image; Performing adaptive histogram equalization processing on the time-series synchronous snap-in image to obtain a brightness-corrected image; Determining a Gaussian filter kernel based on the noise level of the brightness-corrected image, performing noise suppression processing to obtain a denoised snap image, and performing color normalization processing on the denoised snap image to obtain a standardized image with uniform color; Performing edge enhancement processing on the standardized image with uniform color to obtain a preprocessed snap-in image; The pre-processed buckle image is subjected to multi-cascade depth-separable inverse convolution processing to obtain a multi-scale buckle feature image.

3. The real-time monitoring method for buckle installation line according to claim 2, characterized in that: The step of performing multi-cascade depth-separable inverse convolution processing on the pre-processed buckle image to obtain a multi-scale buckle feature image includes: Performing channel expansion processing on the preprocessed snap-in image by 1×1 convolution to obtain an expanded channel feature map; The extended channel feature map is input into four cascade feature extraction modules, each of which uses 3×3 depth-separable convolution with expansion rates of 1, 2, 4, and 8 to obtain hierarchical features with different receptive fields; Perform channel compression processing on the hierarchical features of the different receptive fields through 1×1 convolution to obtain a compressed feature map; Performing spatial displacement and channel similarity calculation on the compressed feature map of the current frame and the compressed feature map of the previous frame to obtain a time domain related feature map capturing dynamic changes in buckle installation; Based on the time-domain correlation feature map, the compressed feature map is enhanced, and the output features of the four cascade feature extraction modules are fused in a pyramidal manner to obtain a multi-level buckle feature; A channel attention mechanism and a spatial attention mechanism are used to perform adaptive channel weights and regional weight allocation on the multi-level snap features to obtain a multi-scale snap feature image.

4. The real-time monitoring method for buckle installation line according to claim 3 is characterized in that: The performing spatial shuffling and rearrangement and bidirectional attention calculation on the multi-scale buckle feature image to obtain a buckle segmentation mask and a buckle type identifier includes: Performing spatial dimension uniform processing on feature maps of different scales in the multi-scale buckle feature image to obtain a unified feature map of the same spatial dimension; The unified feature maps of the same spatial dimension are concatenated in the channel dimension to obtain a mixed feature map, and the mixed feature map is spatially shuffled and rearranged to obtain a shuffled feature map; The shuffled feature map is input into a 1×1 convolution and sigmoid activation function, and effective features are screened through a feature selection gating mechanism to obtain a spatial shuffled feature with attention weights; The comprehensive quality index is calculated based on the image quality score in the preprocessing stage and the current feature response strength, and the dynamically adaptive gain parameter is determined through nonlinear mapping; According to the dynamically adaptive gain parameter, the spatial shuffle feature with attention weight is subjected to gain enhancement processing to obtain an enhanced small target snap feature; Based on the enhanced small target buckle feature, few-shot prototype matching and bidirectional attention calculation are performed to obtain a buckle segmentation mask and a buckle type identifier.

5. The real-time monitoring method for buckle installation line according to claim 4, characterized in that: The performing of few-sample prototype matching and bidirectional attention calculation based on the enhanced small target buckle feature to obtain a buckle segmentation mask and a buckle type identifier includes: Performing mean calculation on a small number of labeled buckle installation sample features to obtain prototype features, and performing intensity downsampling processing on the prototype features to obtain noise-reduced supporting prototype features; Based on the enhanced small target buckle feature, three parallel branches with different scales of receptive fields are constructed to obtain a multi-level query feature image; Calculating cosine similarity between the multi-level query feature image and the noise-reduced supporting prototype feature to obtain a preliminary matching result; Input the preliminary matching result into the dual-dimensional calculation of spatial attention and channel attention to obtain a spatial attention map and a channel attention vector respectively, and fuse the spatial attention map and the channel attention vector to obtain a fused attention result; The fused attention result is processed by an abnormal prior constraint module, and background area interference is suppressed according to prior knowledge of buckle installation position, shape and size to obtain a buckle segmentation mask and a buckle type identification.

6. The real-time monitoring method for buckle installation line according to claim 1, characterized in that: The extracting of geometric and appearance features based on the buckle segmentation mask and the buckle type identifier, and calculating the normal interval deviation to obtain the buckle installation quality score and defect heat map include: Performing morphological operations and edge extraction on the buckle segmentation mask to obtain a buckle geometric feature vector; Performing texture analysis and color statistics based on the buckle type identifier and the corresponding area of ​​the buckle segmentation mask to obtain a buckle appearance feature vector; Merging the buckle geometric feature vector and the buckle appearance feature vector to obtain a target buckle feature vector; Calculate the difference between the buckle feature vector and the preset upper and lower bounds of the normal interval to obtain the deviation distance of each feature dimension; A weighted sum calculation is performed based on the deviation distance and feature importance weight of each feature dimension to obtain a buckle installation quality score; The characteristic vector gradient of the buckle installation quality score is calculated and back-propagated to the original image space, and an upsampling process is performed to obtain a defect heat map.

7. The method for real-time monitoring of a buckle installation line according to claim 1, characterized in that: The fixed time classification decision mechanism analysis is performed based on the buckle installation quality score and the defect heat map, and the pipeline control instruction is output, including: Based on the buckle installation quality score and the defect heat map, a hierarchical decision structure of a direct judgment layer, a detailed analysis layer and an expert decision layer is established, and a fixed time slice is allocated to each layer of the hierarchical decision structure to obtain a fixed time decision framework with a constant total processing time; In the direct judgment layer, a dual-threshold rapid comparison is performed on the buckle installation quality score, and a decision result is directly output for obviously qualified or obviously unqualified samples, and samples within the threshold range are marked as awaiting further analysis, thereby obtaining a preliminary decision result and the first-time remaining quantity; The samples to be further analyzed and the defect heat map are input into the detailed analysis layer, the defect heat map is segmented and feature clustered, and the time-aware algorithm is used to complete the detailed defect analysis to obtain the intermediate decision result and the second time remaining amount; For particularly complex samples, the intermediate decision results, the buckle installation quality score and the defect heat map are input into the expert decision layer, and similarity matching is performed based on the historical case library to obtain high-level decision results; Decision fusion is performed on the preliminary decision result, the intermediate decision result and the advanced decision result to obtain a final decision result, and the final decision result is converted into a specific pipeline control instruction through a decision-instruction mapping table.

8. A real-time monitoring system for a buckle installation line, characterized in that: Used to implement the real-time monitoring method of the buckle installation assembly line according to any one of claims 1 to 7, the real-time monitoring system of the buckle installation assembly line comprises: An image acquisition module is used to perform image acquisition and multi-cascaded depth-separable inverse convolution processing on the buckle installation area to obtain a multi-scale buckle feature image; An attention calculation module, used for performing spatial shuffling and bidirectional attention calculation on the multi-scale buckle feature image to obtain a buckle segmentation mask and a buckle type identifier; An extraction module, used to extract geometric and appearance features based on the buckle segmentation mask and the buckle type identifier, and calculate the normal interval deviation to obtain a buckle installation quality score and a defect heat map; An output module is used to perform fixed-time grading decision mechanism analysis based on the buckle installation quality score and the defect heat map, and output pipeline control instructions.

Citation Information

Patent Citations

  • Intelligent control method and system for automatic optical detection of apparent defects and medium

    CN116188475A

  • Method for detecting sealing performance of aluminum foil seal based on unsupervised learning

    CN116934725A

  • Golden finger quality detection method, device, equipment and medium

    CN119048418A

  • Surface detection method and system for precise fastener

    CN119338827A

  • Visual inspection system of automobile door panel assembly

    CN119555684A

Cited By

  • Wall body seam beautifying defect monitoring method and system based on image recognition and medium

    CN120997161A

  • Intelligent circulation system for stainless steel continuous casting slabs and control method of intelligent circulation system

    CN121962123A