Multi-type split part splicing error detection and position correction method and system based on two-stage cascade network and application

By combining a two-stage cascade network with industrial camera calibration technology, the problems of insufficient flexibility and noise interference in the detection of split parts are solved, high-precision, low-cost stitching error detection and calibration are achieved, and detection efficiency and robustness are improved.

CN120634997APending Publication Date: 2025-09-12EAST CHINA NORMAL UNIV

Patent Information

Application Number
CN202510719510.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

The existing technology for split parts inspection has problems such as insufficient production line flexibility, detection accuracy affected by subjective factors, insufficient generalization ability for multiple types of split parts, and severe environmental noise interference, resulting in high production costs, low efficiency and reduced detection accuracy.

Method used

A multi-type split parts assembly error detection method based on a two-stage cascade network is adopted. Through MoE-RT-DETR and HRNet network cascade training, combined with industrial camera calibration technology, closed-loop detection from pixel coordinates to physical coordinates is achieved, assembly deviations are dynamically adjusted, noise interference is suppressed, and detection accuracy and robustness are improved.

Benefits of technology

It achieves high-precision, low-cost error detection for the splicing of split parts, reduces key point positioning errors by 80%, improves robustness by 50%-70%, increases inference speed by 50%, reduces memory usage by 26%, and reduces calibration errors by 30%, making it suitable for intelligent manufacturing in complex industrial scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005429124600000031
    Figure BDA0005429124600000031
  • Figure BDA0005429124600000032
    Figure BDA0005429124600000032
  • Figure BDA0005429124600000033
    Figure BDA0005429124600000033
Patent Text Reader

Abstract

The invention discloses a multi-type split part splicing error detection and position correction method based on a two-stage cascade network, and the method comprises the steps: 1, carrying out the data collection and processing of multi-type split parts, and enabling the data to serve as the input of two-stage network cascade training of a MoE-RT-DETR network and an HRNet network; 2, performing model training based on a two-stage cascade network; step 3, reasoning based on the trained two-stage network model, extracting pixel coordinates of key points, and converting the pixel coordinates into physical world coordinates through preloaded industrial camera calibration parameters; and step 4, carrying out splicing error detection based on the converted physical coordinates, calculating a key point difference value and guiding position correction adjustment. The invention further discloses a detection and position correction system for implementing the method, and the detection and position correction system has a wide application scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of industrial automation detection and deep learning technology, and relates to a method, system and application for detecting and calibrating the splicing errors of multiple types of split parts based on a two-stage cascade network. Background Art

[0002] Traditional inspection of the assembly of split industrial parts relies primarily on manual visual alignment or traditional rule-based visual algorithms (such as the YOLO series of convolutional networks). This model has three core pain points: First, the production line lacks flexibility, requiring dedicated inspection tooling and assembly lines for different split parts, which is costly. Second, manual visual alignment is inefficient, requiring operators to repeatedly adjust the position of parts, and accuracy is significantly affected by subjective factors. Third, existing automated systems use the traditional YOLO inspection algorithm, which has certain limitations, insufficient generalization capabilities for multiple types of split parts, and the local receptive field limitations of YOLO's CNN architecture. These issues directly increase production costs and hinder efficiency improvements in industrial manufacturing.

[0003] Deep learning technology is gradually being introduced into the field of industrial split-part inspection, but existing solutions still have certain shortcomings. For example, the improved YOLO method proposed in Chinese invention patent application 202311595225.1 optimizes local detection accuracy through the Transformer prediction head and the CIOU loss function. However, due to the limited local receptive field, the CNN architecture of its network cannot effectively model the two-dimensional XY spatial coordinate matching relationship between split parts. It is also unable to eliminate the coupling interference of translation and rotation errors through dynamic weight allocation, which leads to the accumulation of correlation errors of key features and thus reduces detection accuracy. Existing design solutions attempt to improve detection accuracy by improving the network structure, but in actual industrial scenarios, they still face the severe challenge of environmental noise interference. For example, factors such as uneven lighting caused by aging bulbs and reflective distortion caused by notches, oil stains on the part surface covering the fiducial marks, and image blur caused by lens dust can all have a negative impact on detection, further exacerbating production delays and resource waste. In addition, the existing solutions have not conducted targeted research on the two-dimensional splicing errors of multiple types of split parts. They are unable to correct small coordinate offsets with high precision, and there is no mechanism to dynamically adjust assembly deviations, making it difficult to meet the assembly needs in complex industrial scenarios. Summary of the Invention

[0004] In order to address the shortcomings of the existing technology, the purpose of the present invention is to provide a method, system and application for detecting and calibrating the splicing errors of multiple types of split parts based on a two-stage cascade network. The method and system in the present invention have low detection cost, high detection accuracy and good real-time performance, and are particularly suitable for solving the precision calibration problem caused by the splicing deviation of split notches during the operation of split industrial parts on the assembly line; the network used by the method and system in the present invention is highly innovative and has strong processing capabilities, and has greater advantages than the current mainstream detection model (YOLO).

[0005] The specific technical solution for achieving the purpose of the present invention is:

[0006] The present invention provides a method for detecting and calibrating the splicing errors of multiple types of split parts based on a two-stage cascade network, the method comprising:

[0007] Step 1: Collect and process data of multiple types of split parts as input for the two-stage cascade training of the MoE-RT-DETR network and the HRNet network;

[0008] Step 2: Model training based on a two-stage cascade network;

[0009] Step 3: Perform inference based on the trained two-stage network model, extract the pixel coordinates of key points, and convert the pixel coordinates into physical world coordinates using pre-loaded industrial camera calibration parameters;

[0010] Step 4: Perform stitching error detection based on the converted physical coordinates, calculate key point differences and guide calibration adjustments.

[0011] In step 1, the data collection and processing includes: data standardization collection, data preprocessing, data rough labeling, and data fine labeling;

[0012] The data standardization acquisition includes using an automated detection device to collect images of split industrial parts under multiple conditions to generate an original split part image dataset; the dataset includes images covering typical interference factors in actual industrial scenes;

[0013] The data preprocessing refers to performing operations such as denoising, enhancement, illumination correction, and image standardization on the collected original split part image dataset, so that the images taken from different locations have stronger contrast, no reflections or shadows, and have the same resolution;

[0014] The data rough labeling refers to rough labeling of the scratch area on the split parts to determine the label category;

[0015] The data fine marking refers to the pixel-level fine marking of the geometric feature points of the separated parts in the rough marking area;

[0016] and / or,

[0017] The coarsely annotated files and preprocessed normalized images are used as input units for MoE-RT-DETR network training;

[0018] The precisely annotated files and preprocessed standardized images are used as input units for HRNet network training.

[0019] In step 2, the MoE-RT-DETR network is a coarse detection network, and the HRNet network is a fine detection network;

[0020] The coarse detection network outputs the bounding box of the split part area, and the fine detection network performs key point detection and coordinate regression based on the ROI area cropped by the bounding box.

[0021] The MoE-RT-DETR network includes: a backbone network, a feature reshaping module, a hybrid expert module and a decoder;

[0022] The backbone network extracts detailed features and semantic information at different levels based on the residual structure;

[0023] The feature reshaping module reshapes the shallow feature map output by the backbone network, converts it into a serialized representation to obtain reshaped features, and performs feature enhancement through the Transformer encoder;

[0024] The hybrid expert module consists of K fully connected independent expert networks, which distribute and fuse the enhanced features of the Transformer encoder through a dynamic routing mechanism;

[0025] The decoder associates encoder features via a deformable criss-cross attention mechanism.

[0026] The HRNet network includes: a multi-resolution feature extraction module, a dynamic sparse feature selection module and a cross-resolution feature fusion module;

[0027] The multi-resolution feature extraction module achieves efficient fusion of local details and global semantic information by working in parallel with the high-resolution branch and the low-resolution branch;

[0028] The dynamic sparse feature selection module uses a dynamic gating network to generate a sparse adjacency matrix, retaining only the strongly correlated features related to the annotated key points and reducing background noise;

[0029] The cross-resolution feature fusion module receives the output from the dynamic sparse feature selection module and the high- and low-resolution branch features generated by HRNet itself, adjusts the branch feature fusion weights, and outputs the fused multi-resolution features.

[0030] In step 3, after obtaining the industrial camera calibration parameters, manually implement radial and / or tangential distortion correction. Convert the pixel coordinates output by the inference during the two-stage model detection process into physical world coordinates in real time. Refer to the following formula to implement the conversion correction:

[0031]

[0032] Among them, r 2 =x 2 +y 2 , x and y represent the normalized coordinates of the original image coordinate system, and x corrected and y corrected is the coordinate after distortion correction, k1 and k2 are radial distortion coefficients, and p1 and p2 are tangential distortion coefficients; the formula for normalized coordinates (x, y) is:

[0033]

[0034] Then, based on the external parameters R and T, the normalized coordinates are converted to physical world coordinates using the following formula:

[0035]

[0036] In step 4, the target key point detection of the split parts appearing in the detection process is performed by calling the two-stage cascade network training model D2 in the main function program and adding the pixel coordinate conversion module into the physical world coordinate operation module in the main program function to obtain the detection key point 1 (x actual1 ,y actual1 ) and key point 2(x actual2 ,y actual2 );

[0037] The center point coordinates (x mid ,y mid ), and the error detection of split industrial parts splicing is realized by the difference between the physical world coordinates of the two key points and the physical world coordinates of the center point, and the difference x is obtained. res1 , x res2 ,y res1 ,y res2 , the mathematical formula is as follows:

[0038]

[0039] and / or,

[0040] During the operation, after running the main program function, the difference x displayed on the page is obtained res1 , x res2 ,y res1 ,y res2To guide manual operation to control the machine to assemble split industrial parts, where a positive value of x indicates adjustment to the left; a negative value of x indicates adjustment to the right; a positive value of y indicates adjustment downward; and a negative value of y indicates adjustment upward.

[0041] The present invention also provides a detection and calibration system for implementing the above method, the system comprising: a data acquisition module, a data preprocessing module, a model cascade training module, a model deployment module, and a dynamic feedback module;

[0042] The data acquisition module is used to acquire images under multiple conditions;

[0043] The data preprocessing module is used for performing denoising, illumination correction and image standardization;

[0044] The model cascade training module is used to perform two-stage training based on the MoE-RT-DETR and HRNet networks;

[0045] The model deployment module is used to lightweight the trained model and integrate it into the inference engine;

[0046] The dynamic feedback module is used to complete the conversion from pixels to physical coordinates and the detection of stitching errors.

[0047] The present invention also provides the application of the above method or system in the detection of assembled and split parts in industrial parts production plants.

[0048] The beneficial effects of the present invention include: the present invention proposes a method and system for detecting and calibrating the splicing errors of multi-type split industrial parts based on two-stage network cascade training. Through the innovatively designed cascade network architecture and full-process optimization strategy, a technological breakthrough in the detection of splicing errors of industrial parts has been achieved. The system adopts the MoE-RT-DETR and HRNet cascade architecture to construct a "coarse positioning-fine detection" collaborative mechanism, which improves the detection accuracy of small targets (such as bounding box areas) and complex shapes (multiple types of split parts), and reduces the key point positioning error by 80% (from ±2.5 to ±0.5), thanks to the cross-resolution fusion and ROI cropping strategy of HRNet. The hybrid expert module of MoE-RT-DETR and the dynamic sparse feature selection of HRNet effectively suppress noise interference, and the robustness is improved by 50%-70%. , while also lightweighting the Transformer encoder and dynamic routing mechanism, increasing inference speed by 50% and reducing memory usage by 26%. The system deeply integrates industrial camera calibration technology to build a closed-loop detection system from pixel coordinates to physical coordinates. Combined with an automated error calculation module, it implements calibration adjustments, reducing calibration errors by 30% without the need for repeated manual adjustments. Standardized acquisition processes and a two-level labeling strategy significantly improve production efficiency, enabling real-time processing to adapt to high-speed production lines. The dynamic weight adjustment mechanism supports the rapid expansion of multiple types of split parts, providing a high-precision, low-cost, and highly real-time full-process solution for intelligent manufacturing. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.

[0050] Figure 1 Flowchart of the training inference method in an embodiment of the present invention.

[0051] Figure 2 This is a diagram of the first-stage MoE-RT-DETR network framework constructed in an embodiment of the present invention.

[0052] Figure 3 This is a diagram of the two-stage HRNet network framework constructed in an embodiment of the present invention.

[0053] Figure 4 This is a flow chart of a multi-type split industrial parts assembly error detection and alignment system based on two-stage network cascade training in an embodiment of the present invention.

[0054] Figure 5 2 is a comparison chart of model network performance parameters in an embodiment of the present invention.

[0055] Figure 6 This is an example diagram of the HTML web page interface of the system in an embodiment of the present invention. DETAILED DESCRIPTION

[0056] The present invention is further described in detail with reference to the following specific examples and accompanying drawings. The processes, conditions, experimental methods, etc. for implementing the present invention, except for those specifically mentioned below, are common knowledge and common common sense in the art and are not particularly limited by the present invention.

[0057] The present invention provides a method for detecting and calibrating the splicing errors of multiple types of split parts based on a two-stage cascade network, comprising the following steps:

[0058] Step 1: Standardize data collection, preprocess data, and perform rough and fine labeling on multiple types of split parts to serve as input for the two-stage cascade training of the Mixture of Experts Real-Time Detection Transformer (MoE-RT-DETR) and the High-Resolution Network (HRNet);

[0059] The standardized data collection for multiple types of split parts refers to:

[0060] Using industrial automated inspection equipment (such as high-resolution industrial cameras), we capture images of split industrial parts under various conditions, including different angles and lighting intensities. This ensures that typical interference factors in actual industrial scenarios (such as oil, reflections, and dust) are covered, generating an original split part image dataset.

[0061] The split parts data standardization and preprocessing method includes:

[0062] 1) Input the original split part image data:

[0063] Input original images of various types of split parts acquired by an industrial camera, including but not limited to images of cross-shaped split parts, V-shaped split parts, X-shaped split parts, and parts in different split states;

[0064] 2) Image data preprocessing:

[0065] Denoising and enhancement: Gaussian filtering is applied to remove noise, and histogram equalization is used to improve image contrast;

[0066] Lighting correction: Eliminate local reflections or shadows caused by uneven lighting based on the Retinex algorithm;

[0067] Image normalization: All images are scaled to a fixed resolution (e.g., 512 × 512 pixels) and converted to grayscale to reduce computational complexity.

[0068] The data rough labeling and fine labeling include:

[0069] 1) Rough annotation stage:

[0070] Use the rectangular box annotation tool to roughly label the notched area on the split part with Label1 and define the label category (such as cross-bbox (cross area bounding box), v-bbox (V area bounding box), x-bbox (X area bounding box)).

[0071] The coarse annotation JSON file in the output COCO format, which contains the bounding box coordinates and category information of the region, is used together with the preprocessed image as the input unit for the first stage MoE-RT-DETR network training;

[0072] 2) Fine marking stage:

[0073] Based on the rough labeling results, use the key point labeling tool to perform pixel-level precise labeling of the geometric feature points of the split parts in the rough labeling area (Label2);

[0074] Generate a JSON-formatted annotation file containing keypoint coordinates (X / Y axis), spatial relationships, and label categories (e.g., cross-point, v-point, x-point). Together with the preprocessed image, this file serves as the input unit for the second-stage HRNet network training.

[0075] Step 2: Model training based on a two-stage cascade network (MoE-RT-DETR and HRNet);

[0076] After the data standardization collection, preprocessing and coarse and fine annotation are completed, the preprocessed standardized image and coarse and fine annotation files are input into the two-stage cascade network training; specifically, the preprocessed standardized image and coarse annotation file are input into the MoE-RT-DETR network, and the preprocessed standardized image and fine annotation file are input into the HRNet network;

[0077] The two-stage cascade network consists of a coarse detection network (based on MoE-RT-DETR) and a fine detection network (based on HRNet). The cascade training process includes:

[0078] The MoE-RT-DETR network and the HRNet network are cascaded in sequence. The coarse detection network inputs the coarse annotation file containing the coordinates of the scratch area bounding box and category information, and can output the bounding box of the separated part area. The fine detection network inputs the fine annotation file that performs pixel-level fine annotation of the geometric feature points of the separated parts in the coarse annotation area, and performs key point detection and coordinate regression based on the ROI area cropped by the bounding box (that is, the bounding box area detected by MoE-RT-DETR);

[0079] After training, the MoE-RT-DETR network obtains an object detection model for a coarse detection task, including a network structure and trained model weight parameters, and can output a coarse bounding box for the input image;

[0080] The MoE-RT-DETR network is based on the existing DETR network and includes a backbone network, a feature reshaping module (including a Transformer encoder), a mixture of experts (MoE) module, and a decoder.

[0081] The backbone network extracts image features based on the residual architecture and inputs the standardized split part image I∈R H×W (H and W are the image height and width, i.e., spatial dimensions), and four-level feature maps (C1-C4) are extracted through multi-level convolution. [1] Among them, the C1 layer retains high-resolution details (resolution ), used to capture the fine structure of the edge of the mark; the C4 layer is compressed to Extract global semantic features; the middle C2 is C3 is In order to enhance the adaptability to multiple types of notches, the C2-C4 layers introduce deformable convolution, dynamically adjust the convolution kernel offset, and output shallow feature maps. (C is the number of feature channels), providing a basis for local details for subsequent processing.

[0082] The feature reshaping module is used to reshape the shallow feature map output by the deformed convolution in the backbone network, convert it into a serialized representation, obtain the reshaped features, and further enhance the feature expression capability through the Transformer encoder.

[0083] The feature reshaping module transforms F ResNet Divided into Indicates that the feature map is divided into H′×W′ image blocks, and the feature dimension of each image block is C (i.e., the number of feature channels). Through the reshaping operation, the feature map is converted into an N×C matrix form (N=H′×W′, the number of image blocks), where the reshaped feature of the i-th image block is represented by X i ∈R1×C All reshaped features are X = [X1, X2, ..., X N ] That is, X = Reshape (F ResNet )∈R C×N , and in order to preserve spatial information, add a learnable position code P pos ∈R N×C , get the input features: X input =X+P pos ;

[0084] The serialization representation refers to converting spatial information into a sequence format to adapt to the Transformer, and the position encoding P pos By encoding spatial coordinate information through learnable parameters, the Transformer encoder's position-sensitivity requirement is addressed.

[0085] The Transformer encoder is used to enhance the feature expression capability by converting the serialized input feature X input Input Transformer encoder to get the encoded feature F enc ∈R N×C , through the multi-head self-attention mechanism (Multi-HeadSelf-Attention), local details are integrated with global context information to model global dependencies. The calculation formula for single-head attention is:

[0086]

[0087] in, (h is the number of attention heads, usually h = 8);

[0088] The Q h is the input feature X input The linearly transformed query matrix represents the attention weight of the current feature on the target position; K h is the input feature X input The linearly transformed key matrix represents the set of eigenvectors at all positions; is the bond matrix K h The transposed matrix of V h is the input feature X input The linearly transformed value matrix represents the actual content vector set of each position in the input sequence; d k For each head dimension; is the learnable projection matrix (LearnableProjection Matrices);

[0089] The multi-head self-attention mechanism splices the multi-head outputs and passes through the linear layer W O ∈R C×CFusion features, finally get the encoding feature Z∈R N×C , the formula is as follows:

[0090] Z=MultiHead(Q,K,V)=Concat((head1,...,head h )W O

[0091] The Mixture of Experts (MoE) module is the core module of MoE-RT-DETR and consists of K independent expert networks, all of which are fully connected networks (FCNs). Its function is to distribute the features Z enhanced by the Transformer encoder to different expert networks through a dynamic routing mechanism, and output a more robust feature representation after fusion. Specifically:

[0092] The dynamic routing mechanism receives the input feature Z through the gating network and generates each image block Z i Routing weight G∈R N×K =Softmax(Top2(W g ×Z+b g ), where K is the number of experts, W g ∈R C×K 、b g ∈R K For learnable parameters, Top2 selects the two experts with the highest weights (e.g., split cross-type k1, split V-type k2), ignores low-correlation experts, reduces the activation intensity of noise blocks, achieves sparse routing, reduces computational redundancy, and suppresses the impact of noise blocks;

[0093] For each image block Z in the input feature Z i ∈R 1×C According to the above routing weight Input two selected expert networks have to

[0094] The expert network optimizes the feature expression of different split part types through specific parameter configuration, and weightedly fuses the expert output to obtain the fused output:

[0095]

[0096] The weighted fusion expert output retains the diverse modeling capabilities of the expert network, dynamically adapts to multiple types of split parts, and can suppress noise interference. Finally, the fused feature Z is generated. MoE ∈R N×C ;

[0097] The decoder inputs the fused feature Z MoE, associate encoder features through the deformable cross-attention mechanism. Initialize the query matrix Q∈R with 300 learnable anchors (Anchor Queries) 300×C And the key-value matrix K=V=Z MoE Calculate attention weights:

[0098]

[0099] where d k = C is the feature dimension. And through the two-layer feedforward network (FFN), the final result of MoE-RT-DETR network training is returned: the coarse positioning bounding box

[0100] In the first stage of network training, the final output network model D1 (including weights and loss function parameters) is the total loss function: loU stage1 =L GloU +L Focal +L Sparse The formulas and meanings of each loss function are as follows:

[0101] Bounding Box Regression Loss (GIoU): Optimize the intersection-over-union ratio and shape matching between the bounding box and the real box; represents the coarse localization bounding box predicted by the model; B coarse represents the bounding box of the ground truth annotated by the label;

[0102] Focal Loss: Suppress the interference of simple negative samples (noise, background) on training, where P cls is the regional confidence;

[0103] Expert Equilibrium Loss (Sparsity Loss): Constrain expert load balancing to avoid overfitting and prevent the model from relying on a single expert network;

[0104] In the second stage training, the model D1 trained by the first stage network, the original image I and the fine annotation results are used as the input of the second stage network HRNet. Identify and crop the ROI region I from the original image ROI ∈R H′×W′×3 (i.e. original image I + coarse label Label1), where H′=y max -y min , W'=X max -X min ;

[0105] The two-stage HRNet network includes: a multi-resolution feature extraction module, a dynamic sparse feature selection (DSFS) module, and a cross-resolution feature fusion module;

[0106] The multi-resolution feature extraction module achieves efficient fusion of local details and global semantic information by working in parallel with the high-resolution branch (HRB) and the low-resolution branch (LRB);

[0107] The dynamic sparse feature selection (DSFS) module uses a dynamic gating network to generate a sparse adjacency matrix, retaining only the strongly correlated features related to the labeled key points, thereby effectively reducing the impact of background noise. The asymmetric convolution kernel is used to extract local features in the row and column directions, further capturing the texture details of the edge, so that the system can more accurately identify the key features of the annotation. Specifically, in the high-resolution branch, a sparse adjacency matrix A is generated through a gating network to suppress the interference of background noise nodes. The sparse adjacency matrix is ​​defined as follows:

[0108]

[0109] Among them, S ij Represents the similarity score between feature nodes i and j, ξ is the set threshold used to filter strong correlation features;

[0110] The cross-resolution feature fusion module receives the output from the dynamic sparse feature selection (DSFS) module and the high (HR) and low (LR) resolution branch features generated by HRNet itself. In this process, the LR branch features are upsampled to the same HR resolution H×W by bilinear interpolation to obtain F LR ∈R C×H×W , enhance the semantic information of detail features; while the high-resolution branch compresses the HR branch features to the same resolution h×w of LR by convolution (such as convolution with a step size of 2) to obtain F HR ∈R C ×h×w , supplement the local details of the LR branch. In this way, the fusion weights of HR and LR features are dynamically adjusted to ensure the dominance of key point features and output the fused multi-resolution feature F fuse (Having both high-resolution details and low-resolution semantics), the specific fusion formula is:

[0111] F fuse =αF HR +(1-α)Downsample(F LR );

[0112] Among them, F HR and FLR Represent high-resolution and low-resolution features respectively, are dynamic weights, and α is dynamically adjusted by the channel attention mechanism;

[0113] Finally, the feature F fuse It is fed into the keypoint regression head, which predicts each annotation point through a lightweight convolution layer (1×1 convolution) and outputs its pixel coordinates (x i ,y i ), i represents the i-th key point, and outputs the confidence score. Finally, the network model D2 (including weights and loss function parameters) that combines the weights and parameters of the D1 network model is obtained;

[0114] Step 3: Calibrate the parameters of the industrial camera (intrinsic parameter matrix K, distortion coefficient D, external parameter R, external parameter T). In the process of using the two-stage network model inference, the calibrated parameters are pre-loaded into the system and the detected pixel coordinates (x i ,y i ) is converted to physical world (real) coordinates (x actual ,y actual );

[0115] The industrial camera calibration parameters refer to the calibration parameters obtained by Zhang's calibration method using a calibration plate (such as a chessboard or circular array) [2] Get the camera's intrinsic parameter matrix K, distortion coefficient D, and extrinsic parameter matrix (rotation matrix R and translation vector T).

[0116] After obtaining the relevant calibration parameters, manually implement radial and / or tangential distortion correction. Convert the pixel coordinates output by the inference during the two-stage model detection process into physical world (real) coordinates in real time. Refer to the following formula to implement the conversion correction:

[0117]

[0118] Among them, r 2 =x 2 +y 2 , x and y represent the normalized coordinates in the original image coordinate system (i.e., the coordinates after considering the camera intrinsic parameter matrix K), and x corrected and y corrected is the coordinate after distortion correction, k1, k2 are radial distortion coefficients, p1, p2 are tangential distortion coefficients. The formula for normalized coordinates (x, y) is:

[0119]

[0120] Then, based on the external parameters R and T, the normalized coordinates are converted to physical world coordinates using the following formula:

[0121]

[0122] Step 4: Implement the detection of assembly errors of split industrial parts and guide the alignment in the main function program. The specific steps are as follows:

[0123] The process of detecting the splicing error of split industrial parts is as follows: by calling the model D2 trained by the two-stage cascade network in the main function program, the target key point detection of the split parts appearing in the detection process is performed, and the pixel coordinate conversion into the physical world coordinate operation module is added in the main program function to obtain the detection key point 1 (x actual1 ,y actual1 ) and key point 2(x actual2 ,y actual2 );

[0124] The center point coordinates (x mid ,y mid ) (Add a mathematical calculation module for calculating the center point coordinates in the main program function) and realize the error detection of split industrial parts splicing by the difference between the physical world coordinates of the two key points and the physical world coordinates of the center point, and get the difference x res1 , x res2 ,y res1 ,y res2 , the mathematical formula is as follows:

[0125]

[0126] The guidance calibration means: during the operation, after running the main program function, the difference x displayed on the page is obtained. res1 , x res2 ,y res1 ,y res2 To guide manual operation and control of the machine to assemble split industrial parts, where a positive value of x means adjustment to the left; a negative value of x means adjustment to the right; a positive value of y means adjustment downward; a negative value of y means adjustment upward;

[0127] The present invention also provides a multi-type split parts splicing error detection and calibration system based on a two-stage cascade network, such as Figure 4 As shown, the system includes: data acquisition module, data preprocessing module, model cascade training module, model deployment module, and dynamic feedback module;

[0128] The data acquisition module includes an industrial camera array unit and an environmental parameter acquisition unit;

[0129] The industrial camera array unit uses a multi-viewing angle high-resolution industrial camera, eliminates shadow interference through a ring fill light, and collects multi-angle images of split parts in real time;

[0130] The environmental parameter acquisition unit integrates a light sensor and a posture sensor to monitor and record the light intensity, camera posture angle and lens focal length parameters during shooting in real time;

[0131] The data preprocessing module includes an image standardization processing unit and a labeling processing unit;

[0132] The image normalization processing unit uses Gaussian filtering for denoising and Retinex algorithm for illumination correction, and uniformly scales the image to 512×512 pixels (grayscale image);

[0133] The labeling processing unit implements coarse and fine labeling through Label Studio and Supervisely tools. The coarse labeling outputs COCO format bounding boxes, and the fine labeling generates a JSON file containing key point coordinates.

[0134] The model cascade training module refers to taking the data preprocessing module as input and performing two-stage cascade training through the MoE-RT-DETR network and the HRNet network;

[0135] The model deployment module includes a model lightweight unit and an inference engine integration unit;

[0136] The model lightweight unit uses TensorRT to perform FP16 quantization and layer fusion optimization on the trained MoE-RT-DETR and HRNet models;

[0137] The inference engine integration unit deploys the optimized model to an edge computing device (such as NVIDIA Jetson AGX Xavier) based on the TensorRT framework;

[0138] The dynamic feedback module includes a coordinate conversion and error calculation unit;

[0139] The coordinate conversion unit calls Zhang's calibration parameters (intrinsic parameter matrix K, distortion coefficient D, external parameter R / T) to convert the detected pixel coordinates (x, y) into physical world coordinates (X, Y) through a formula;

[0140] The error calculation unit outputs the result through the key point center coordinate difference formula to guide manual calibration adjustment.

[0141] like Figure 5 The figure shows the comparison results between the proposed method and the existing methods. A larger number of parameters means higher computing requirements. FPS = represents the number of image frames that can be processed per second. Precision (mAP = 0.5) is an important indicator for measuring the accuracy of target detection. A higher value indicates more accurate detection.

[0142] Example

[0143] like Figure 6 The figure shows an embodiment of the method of the present invention for realizing the splicing detection of split parts in an industrial parts production factory, and realizing the alignment by deploying it to the edge device and then obtaining the results on the software system display interface on the computer.

[0144] Step 1: Edge device deployment:

[0145] A high-resolution camera is used to capture images of split parts in real time, and a pre-trained image processing algorithm is deployed for preliminary detection (two-stage cascade training network).

[0146] Step 2: Use JavaScript to build the user interface and integrate the following functions:

[0147] Image visualization: Real-time display of part images and inspection results transmitted by edge devices;

[0148] Key point coordinate display: Dynamically update the physical coordinates of the detected key points of the part (such as key point 1: X: 150.06mm, Y: 200.68mm);

[0149] Calibration operation panel: Generates adjustment suggestions (such as "adjust 0.24mm to the right") based on the error calculation results, and assists workers in performing calibration actions through buttons or sliders.

[0150] In this embodiment, the calibration process is as follows:

[0151] Detection phase:

[0152] The edge device identifies key point 1 (150.06mm, 200.68mm) and key point 2 (150.54mm, 200.72mm), and calculates the center point as (150.30mm, 200.70mm).

[0153] Error calculation:

[0154] Key point 1 needs to be moved 0.24mm to the right and 0.02mm up; key point 2 needs to be moved 0.24mm to the left and 0.02mm down.

[0155] During the calibration process of the embodiment, the edge device collects image data of the split parts in real time through a high-resolution camera, completes the detection using a two-stage cascade training model, identifies the coordinates of key point 1 (150.06mm, 200.68mm), key point 2 coordinates (150.54mm, 200.72mm), and calculates the coordinates of the center point of the part (150.30mm, 200.70mm) based on this. The data is transmitted to the HTML interface of the computer via the TCP / IP protocol. The system dynamically renders the image and the key point position in the interface, and compares the preset standard coordinates (center point X: 150.30mm, Y: 200.70mm) for error analysis. It is found that key point 1 has ΔX = -0.24mm (needs to move right 0.24mm), ΔY = -0.02mm (needs to move up 0.02mm), while key point 2 has ΔX = +0.24mm (needs to move left 0.24mm), ΔY = +0.02mm (needs to move down 0.02mm). The HTML interface highlights key points where errors exceed the limit and displays adjustment suggestions (e.g., "Adjust key point 1 0.24mm to the right, key point 2 0.24mm to the left"). Workers can then manually adjust the alignment. The system then calculates the new coordinates in real time and transmits feedback to the edge device, simultaneously updating the image of the calibrated part and its error status. Throughout the entire process, the real-time processing capabilities of the edge device reduce data transmission latency (improving efficiency), the interactive visualization of the HTML interface reduces the complexity of manual operations (enhancing scalability), and precise alignment based on physical coordinate error calculations ensures high efficiency and reliability in industrial production.

[0156] References

[0157] [1]He K, Zhang X, Ren S, et al.Deep Residual Learning for ImageRecognition[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition.2016:770-778.

[0158] [2]Zhang, Z.(2000).A flexible new technique for cameracalibration.IEEE Transactions on Pattern Analysis and Machine Intelligence, 22(11), 1330-1334.

[0159] The protection content of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the present invention, changes and advantages that can be thought of by those skilled in the art are included in the present invention and are protected by the appended claims.

Claims

1. A method for detecting and calibrating the splicing errors of multiple types of split parts based on a two-stage cascade network, characterized in that: The method comprises: Step 1: Collect and process data of multiple types of split parts as input for the two-stage cascade training of the MoE-RT-DETR network and the HRNet network; Step 2: Model training based on a two-stage cascade network; Step 3: Perform inference based on the trained two-stage network model, extract the pixel coordinates of key points, and convert the pixel coordinates into physical world coordinates using pre-loaded industrial camera calibration parameters; Step 4: Perform stitching error detection based on the converted physical coordinates, calculate key point differences and guide calibration adjustments.

2. The method according to claim 1, wherein In step 1, the data collection and processing includes: data standardization collection, data preprocessing, data rough labeling, and data fine labeling; The data standardization acquisition includes collecting images of split-type industrial parts under multiple conditions to generate an original split-part image dataset; The data preprocessing refers to performing denoising, enhancement, illumination correction and image standardization on the collected original split parts image data set; The data rough marking refers to rough marking of the scratch area on the split parts; The data fine marking refers to the pixel-level fine marking of the geometric feature points of the separated parts in the rough marking area; and / or, The coarsely annotated files and preprocessed normalized images are used as input units for MoE-RT-DETR network training; The precisely annotated files and preprocessed standardized images are used as input units for HRNet network training.

3. The method according to claim 1, wherein In step 2, the MoE-RT-DETR network is a coarse detection network, and the HRNet network is a fine detection network; The coarse detection network outputs the bounding box of the split part area, and the fine detection network performs key point detection and coordinate regression based on the ROI area cropped by the bounding box.

4. The method according to claim 1, wherein The MoE-RT-DETR network includes: a backbone network, a feature reshaping module, a hybrid expert module and a decoder; The backbone network extracts detailed features and semantic information at different levels based on the residual structure; The feature reshaping module reshapes the shallow feature map output by the backbone network, converts it into a serialized representation to obtain reshaped features, and performs feature enhancement through the Transformer encoder; The hybrid expert module consists of K fully connected independent expert networks, which distribute and fuse the enhanced features of the Transformer encoder through a dynamic routing mechanism; The decoder associates encoder features via a deformable criss-cross attention mechanism.

5. The method according to claim 1, wherein The HRNet network includes: a multi-resolution feature extraction module, a dynamic sparse feature selection module and a cross-resolution feature fusion module; The multi-resolution feature extraction module achieves efficient fusion of local details and global semantic information by working in parallel with the high-resolution branch and the low-resolution branch; The dynamic sparse feature selection module uses a dynamic gating network to generate a sparse adjacency matrix, retaining only the strongly correlated features related to the annotated key points and reducing background noise; The cross-resolution feature fusion module receives the output from the dynamic sparse feature selection module and the high- and low-resolution branch features generated by HRNet itself, adjusts the branch feature fusion weights, and outputs the fused multi-resolution features.

6. The method according to claim 1, wherein In step 3, after obtaining the industrial camera calibration parameters, manually implement radial and / or tangential distortion correction. Convert the pixel coordinates output by the inference during the two-stage model detection process into physical world coordinates in real time. Refer to the following formula to implement the conversion correction: Among them, r 2 =x 2 +y 2 , x and y represent the normalized coordinates of the original image coordinate system, and x corrected and y corrected is the coordinate after distortion correction, k1 and k2 are radial distortion coefficients, and p1 and p2 are tangential distortion coefficients; the formula for normalized coordinates (x, y) is: Then, based on the external parameters R and T, the normalized coordinates are converted to physical world coordinates using the following formula:

7. The method according to claim 1, wherein In step 4, the target key points of the split parts appearing in the detection process are detected by calling the model trained by the two-stage cascade network in the main function program and adding the pixel coordinate conversion operation module into the physical world coordinate operation module in the main program function to obtain the detection key point 1 (x actual1 ,y actual1 ) and key point 2(x actual2 ,y actual2 ); The center point coordinates (x mid ,y mid ), and the error detection of split industrial parts splicing is realized by the difference between the physical world coordinates of the two key points and the physical world coordinates of the center point, and the difference x is obtained. res1 , x res2 ,y res1 ,y res2 , the mathematical formula is as follows: and / or, During the operation, after running the main program function, the difference x displayed on the page is obtained res1 , x res2 ,y res1 ,y res2 To guide manual operation to control the machine to assemble split industrial parts, where a positive value of x indicates adjustment to the left; a negative value of x indicates adjustment to the right; a positive value of y indicates adjustment downward; and a negative value of y indicates adjustment upward.

8. A detection and calibration system for implementing the method according to any one of claims 1 to 7, characterized in that: The system includes: a data acquisition module, a data preprocessing module, a model cascade training module, a model deployment module, and a dynamic feedback module; The data acquisition module is used to acquire images under multiple conditions; The data preprocessing module is used for performing denoising, illumination correction and image standardization; The model cascade training module is used to perform two-stage training based on the MoE-RT-DETR and HRNet networks; The model deployment module is used to lightweight the trained model and integrate it into the inference engine; The dynamic feedback module is used to complete the conversion from pixels to physical coordinates and the detection of stitching errors.

9. Use of the method according to any one of claims 1 to 7, or the system according to claim 8, in the detection of assembled and separated parts in an industrial parts production plant.

Citation Information

Patent Citations

  • Nut welding detection identification method based on deep learning

    CN117576056A

Cited By

  • Dense overlapping target detection method based on wavelet enhancement sparse hybrid expert model

    CN121353953A