Identification method and system for automobile welding spot stud, electronic equipment and storage medium
By acquiring multi-view image data and combining epipolar geometric constraints and triangulation algorithms, the problem of misjudgment in weld stud identification under complex backgrounds was solved, achieving high-precision and robust automated detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU DONGHAN INTELLIGENT EQUIP CO LTD
- Filing Date
- 2026-01-28
- Publication Date
- 2026-05-08
AI Technical Summary
Existing automotive weld stud identification methods are prone to misjudgment in complex backgrounds, especially when there is obstruction or dirt, and the accuracy of machine vision detection is insufficient.
This method combines multi-view image data acquisition, specified model detection, epipolar geometric constraint screening, and triangulation algorithm. The detection results from each viewpoint are screened based on the principle of epipolar geometric constraint, the candidate three-dimensional spatial coordinates of the target are calculated using the triangulation algorithm, and the final recognition result is determined by reprojection error verification.
It improves the robustness and reliability of solder joint stud recognition, effectively avoids occlusion and false detection problems, and realizes high-precision automated detection in complex backgrounds.
Smart Images

Figure CN121998950A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of automotive welding stud identification, and more particularly to an automotive welding stud identification method, system, electronic device, and storage medium. Background Technology
[0002] In the automotive manufacturing and after-sales quality inspection stages, the quality assessment of automotive parts (such as welds, studs, and other critical connecting components) is a core aspect of ensuring vehicle structural safety and reliability. Currently, the industry primarily relies on two methods for quality assessment of common targets such as automotive welds and studs: one is manual visual inspection, where inspectors visually observe the presence and appearance of parts and combine this with their experience to determine whether they meet quality standards; the other is the use of traditional machine vision inspection solutions, which analyze part images using pre-set image feature extraction algorithms to determine the presence of the target.
[0003] In the existing process of identifying automotive weld studs, machine vision detection solutions are used for analysis and identification. Although this method is highly efficient, it is prone to misjudgment when there is partial obstruction or dirt on the parts. Summary of the Invention
[0004] This application provides a method, system, electronic device, and storage medium for identifying automotive welding studs to solve the problems existing in related technologies. The technical solution is as follows: In a first aspect, embodiments of this application provide a method for identifying automotive weld studs, including: Acquire first-view image data of automotive parts from multiple perspectives; Based on the specified model, weld point targets and stud targets are detected in the first image data, and detection results are generated from each viewpoint; Based on the principle of epipolar geometry constraint, the detection results from each perspective are screened to obtain pairs of target points from the same source; The candidate three-dimensional spatial coordinates of the target are obtained by calculating the point pairs of the same source using triangulation algorithm; Reprojection error verification is performed based on the candidate 3D spatial coordinates of the target to determine the 3D detection result; The 3D detection results are compared with the preset standard process parameters to obtain the identification results.
[0005] Secondly, embodiments of this application provide a system for identifying automotive welding studs, comprising: The first acquisition module acquires first image data of automotive parts from multiple perspectives. The first generation module performs weld point target and stud target detection on the first image data based on a specified model, and generates detection results from various perspectives. The detection results include the target's position coordinates, target category, and confidence information. The first module obtains pairs of target points from the same source by filtering the detection results from each viewpoint based on the principle of epipolar geometry constraint. The second module obtains the candidate three-dimensional spatial coordinates of the target by using a triangulation algorithm to calculate the points of the same source target; The first determination module is used to verify the reprojection error based on the candidate 3D spatial coordinates of the target and determine the 3D detection result. The third module is used to compare the 3D detection results with preset standard process parameters to obtain the identification results.
[0006] Thirdly, embodiments of this application provide an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the above-described method for identifying automotive welding studs.
[0007] Fourthly, embodiments of this application provide a computer-readable storage medium that stores computer instructions, wherein when the computer instructions are executed on a computer, the methods in any of the above-described embodiments are performed.
[0008] The advantages or beneficial effects of the above technical solutions include at least the following: In this embodiment, the method for identifying automotive weld studs includes: acquiring first image data of automotive parts from multiple perspectives; detecting weld targets and stud targets in the first image data based on a specified model, generating detection results for each perspective; filtering the detection results for each perspective based on the epipolar geometric constraint principle to obtain pairs of target points from the same source; calculating the candidate three-dimensional spatial coordinates of the target using a triangulation algorithm; verifying the reprojection error based on the candidate three-dimensional spatial coordinates of the target to determine the three-dimensional detection result; and comparing the three-dimensional detection result with preset standard process parameters to obtain the identification result. In this embodiment, by detecting weld targets and stud targets in the first image data based on a specified model and generating detection results for each perspective, the identification of weld targets and stud targets on automotive parts is achieved, solving the problem of difficult identification of small targets in complex backgrounds. Furthermore, the reprojection error verification of the detection results for each perspective is performed using the epipolar geometric constraint principle and triangulation algorithm, avoiding some occlusion and false detection problems through redundancy, solving the occlusion or false detection problems that may exist in a single perspective, and further improving the robustness and reliability of the detection.
[0009] The above overview is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of this application will become readily apparent from the accompanying drawings and the following detailed description. Attached Figure Description
[0010] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments disclosed in this application and should not be construed as limiting the scope of this application.
[0011] Figure 1 This is a flowchart of a method for identifying automotive welding studs according to an embodiment of this application.
[0012] Figure 2 An operating system interface for a method of identifying automotive weld studs according to an embodiment of this application. Figure 1 .
[0013] Figure 3 A structural diagram of a model specified for an automotive weld stud identification method according to an embodiment of this application.
[0014] Figure 4 An operating system interface for a method of identifying automotive weld studs according to an embodiment of this application. Figure 2 .
[0015] Figure 5 This is an operating system framework diagram of an automotive welding stud identification method according to an embodiment of this application.
[0016] Figure 6 This is an operating system communication signal interface for a method of identifying automotive welding studs according to an embodiment of this application.
[0017] Figure 7 This is an operating system identification interface for an automotive welding stud identification method according to an embodiment of this application.
[0018] Figure 8 This is the second operating system recognition interface for a method of recognizing automotive welding studs according to an embodiment of this application.
[0019] Figure 9 This is a block diagram of an electronic device according to an embodiment of the present application. Detailed Implementation
[0020] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of this application. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.
[0021] Figure 1 A flowchart illustrating a method for identifying automotive weld studs according to an embodiment of this application is shown. Figures 1-8 As shown, a method for identifying automotive weld studs may include: S110: Acquire first image data of automotive parts from multiple perspectives; S120: Detect weld point targets and stud targets on the first image data based on the specified model, and generate detection results from each viewpoint; S130: Based on the principle of epipolar geometry constraint, the detection results from each perspective are screened to obtain pairs of target points from the same source; S140: Use triangulation algorithms to calculate the candidate three-dimensional spatial coordinates of the target by calculating the points of the same source target; S150: Verify the reprojection error based on the candidate 3D spatial coordinates of the target to determine the 3D detection result; S160: Compare the three-dimensional detection results with the preset standard process parameters to obtain the identification results.
[0022] In the embodiments of this application, welding point targets and stud targets are detected on the first image data based on a specified model, and detection results from various perspectives are generated. This enables the identification of welding point targets and stud targets on automotive parts, solving the problem of difficult identification of small targets in complex backgrounds. Furthermore, the reprojection error of the detection results from various perspectives is verified by using the epipolar geometric constraint principle and triangulation algorithm. By using redundancy, some occlusion and false detection problems are avoided, solving the occlusion or false detection problems that may exist in a single perspective, and further improving the robustness and reliability of the detection.
[0023] In the embodiments of this application, the overall software architecture consists of four core functional modules, which are connected sequentially through a data transmission interface. The module numbers, names, and connection relationships are as follows: Module 1: Communication module, used to acquire data from the PLC and trigger image capture, connected to the image acquisition module.
[0024] Module 2: Image acquisition module (e.g.) Figure 2 As shown in the image, this module is used to acquire images and send them to the recognition module. This module is connected to the recognition module. Module 3: Recognition Module. The recognition module can use the YOLOv11 object detection algorithm as the specified model. Its network structure diagram is as follows: Figure 3 As shown.
[0025] Module 4: Display module (e.g.) Figure 4 As shown in the image, the system displays the identification results, production logs, and operational status information for each module. It employs the Modbus TCP industrial communication protocol, with a signal transmission delay of ≤10ms, meeting the real-time requirements of the production line.
[0026] The communication module is the core control and interaction unit of the software architecture. Its core function is to achieve bidirectional signal interaction with the PLC, image acquisition module, and recognition module, ensuring the orderly linkage of the detection process. This module uses the Modbus TCP communication protocol to establish a stable connection with the PLC. The specific interaction logic is as follows: When receiving a photo-taking command from the PLC, the communication module immediately forwards the photo-taking signal to the image acquisition module. After the image acquisition module completes image capture, the communication module receives its feedback "photo-taking complete" status signal, verifies it, and forwards it to the PLC. When receiving a recognition start command from the PLC, the communication module triggers the recognition module to start the image analysis task. After the recognition module completes target detection, it receives its feedback recognition result, processes it, and uploads it to the PLC. Specific signal interaction configuration can be completed through the accompanying configuration software, the interface of which is shown below. Figure 5 As shown.
[0027] In step S110, first image data of the automotive parts from multiple perspectives is acquired.
[0028] In the embodiments of this application, the image acquisition module provides a high-quality image data source for target detection. Its core functions include camera image acquisition and adaptive parameter configuration, and it possesses multi-camera collaborative shooting capabilities and compatibility with multiple camera models. At the parameter configuration level, shooting parameter thresholds can be manually preset to ensure the acquisition of optimal image quality that meets recognition requirements. Regarding multi-device adaptation, the image acquisition module supports parallel shooting by multiple cameras, is compatible with mainstream industrial camera models, and is plug-and-play, requiring no separate adaptation driver development. It can flexibly adapt to the equipment configuration requirements of different production lines.
[0029] Automotive parts can be captured separately using industrial cameras arranged from multiple perspectives. Alternatively, an image acquisition module can control the industrial cameras arranged from multiple perspectives to capture images of automotive parts simultaneously, resulting in first image data of automotive parts from multiple perspectives with timestamps.
[0030] In step S120, the first image data is used to detect weld point targets and stud targets based on the specified model, and detection results are generated from each viewpoint.
[0031] In the embodiments of this application, the specified model is a pre-trained convolutional neural network model, which can be a YOLOv11 object detection algorithm model or other high-performance object detection algorithm models (such as Faster R-CNN, SSD, Rtdetr). It is trained on the same custom-labeled training set (10,000 images of multiple car parts and multiple state samples). By adjusting the model input size and training hyperparameters, the presence or absence of solder joints and studs can also be detected. The only differences are in detection speed, accuracy and hardware resource consumption.
[0032] In this embodiment, the model employs a target detection algorithm to achieve automated and high-precision identification of the presence or absence of key automotive parts such as weld points and studs in images, and outputs standardized detection results. To adapt to the needs of different vehicle models and different detection scenarios, the module has a built-in model selection mechanism, providing a variety of preset optimized models for users to choose from as needed, such as a dedicated detection model for small targets, an anti-interference model for weld points in complex backgrounds, and a general model that balances detection speed and accuracy. Training uses 10,000 sample images covering different working conditions, comprehensively covering common automotive part scenarios such as weld points and studs with various styles and states (parts present normally, missing / missing welds, surface stains, partial occlusion, etc.). During the dataset construction process, precise annotation provides high-quality supervision data for model training. In the model training phase, based on the basic network architecture, the training efficiency is optimized through transfer learning strategies, and reasonable hyperparameters (number of iterations, learning rate, batch size, etc.) are set to ensure model convergence. The finally trained model has fast inference capabilities, with a single image detection time of ≤0.1 seconds, which can meet the real-time detection requirements of the production line.
[0033] During the detection process, after receiving the first image data transmitted by the image acquisition module, the recognition module automatically matches it with the specified model selected by the user. The user can also set the detection area and obtain the detection results, such as the presence or absence of targets in the set area, the position coordinates of the targets, the number of targets, and the confidence level, through the specified model inference. The targets can be weld points or studs. After the detection is completed, the detection results from various perspectives are obtained. Simultaneously, the detection completion status is fed back to the communication module, and the detection results are pushed to the display module. The parameters of the recognition module can be set through the accompanying software, the interface of which is shown below. Figure 6-8 As shown.
[0034] In step S130, the detection results from each viewpoint are filtered based on the principle of epipolar geometry constraint to obtain pairs of target points from the same source. In the embodiments of this application, in order to verify the consistency of detection results from multiple perspectives and calculate the three-dimensional position of the target, it is necessary to first determine whether the targets detected in images from different perspectives belong to the same object in physical space (i.e., common source targets). Common source target point pairs refer to a set of pixel coordinates detected in images from different cameras that correspond to the same object in real physical space (such as a solder joint).
[0035] In this embodiment, the epipolar geometric constraint principle based on the fundamental matrix can be used to filter and match the detection results from the first and second perspectives.
[0036] Specifically, the fundamental matrix is obtained based on the relative positional relationship of the image acquisition units (i.e., industrial cameras). The fundamental matrix describes the epipolar geometry between the two views in a binocular stereo vision system. In this embodiment, it is assumed that the intrinsic parameter matrix of the master camera corresponding to the first viewpoint is K1, and the intrinsic parameter matrix of the slave camera corresponding to the second viewpoint is K2. The relative positional relationship between the two cameras is described by the rotation matrix R and the translation vector t (i.e., the pose of the slave camera relative to the master camera). This fundamental matrix F can map the pixels of the first viewpoint on the image plane to a straight line (i.e., the epipolar line) of the second viewpoint on the image plane.
[0037] Using the center point pixel coordinates of the specified target in the first-view detection result as input, matrix multiplication is performed using the fundamental matrix to obtain the epipolar equation of the point on the second-view image plane.
[0038] The center pixel coordinates of all candidate targets in the second-view detection results that belong to the same category group as the specified target in the first view (e.g., both are solder joints or both are studs) are traversed to obtain a set of candidate target center pixel coordinates. For each candidate target in this set, its center pixel coordinates in the second view are set. To verify whether the candidate target matches the target in the first view, the geometric distance from this point to the epipolar line needs to be calculated. That is, the vertical pixel distance from the center point of each candidate target to the corresponding epipolar line is calculated.
[0039] If the calculated vertical pixel distance is less than the preset tolerance threshold, it means that the candidate target point is located near the epipolar line and satisfies the epipolar geometric constraint condition.
[0040] At this point, the candidate target in the second viewpoint is determined to be a common target point pair with the specified target in the first viewpoint. This means that the two detection boxes have a high probability of geometric correspondence and can be regarded as observations of the same weld point or stud in physical space, which can then be used for subsequent triangulation and 3D reconstruction steps.
[0041] In step S140, the triangulation algorithm is used to calculate the candidate three-dimensional spatial coordinates of the target points.
[0042] In one embodiment of this application, after determining the same source target point pair corresponding to the same physical object in the first and second viewpoints through the aforementioned steps, triangulation logic is executed to recover the depth information of the target. The process of calculating the candidate three-dimensional spatial coordinates of the target using a triangulation algorithm on the same source target point pair, where triangulation, also known as multi-view reconstruction, refers to the technique in computer vision that, by using the positional relationship of corresponding pixels in images taken from two or more different viewpoints, combined with the camera's imaging geometric parameters, the specific coordinates of the object corresponding to that pixel in three-dimensional space are deduced in reverse. Its geometrical meaning is to determine the target position using the intersection point in space of two or more rays from the optical centers of different cameras.
[0043] Specifically, to construct a system of linear equations containing projection matrices for both the first and second views, it is first necessary to obtain projection matrices that describe the mapping relationship between points in 3D space and pixels in 2D images. Based on the calibration parameters of the binocular vision system, a first-view projection matrix corresponding to the first image acquisition unit and a second-view projection matrix corresponding to the second image acquisition unit are constructed. Both projection matrices integrate the camera's intrinsic parameters (such as focal length and principal point coordinates) and extrinsic parameters (such as rotation matrices and translation vectors), enabling the transformation of points in the 3D world coordinate system onto their respective 2D image planes.
[0044] Based on the pinhole imaging model, the relationship between any 3D spatial point and its projection point on the image can be linearly constrained using a projection matrix. To solve for the unknown 3D spatial coordinates, the first-view projection matrix and the second-view projection matrix are combined. For example, using the relationship between the row vectors of the first-view projection matrix and the pixel coordinates of the target point in the first view, two linear constraint equations can be constructed for that 3D point; similarly, using the relationship between the row vectors of the second-view projection matrix and the pixel coordinates of the target point in the second view, two more linear constraint equations can be constructed. Combining these four equations constructs an overdetermined system of linear equations containing parameters of both the first-view and second-view projection matrices. The unknowns in this system of equations are the target 3D spatial coordinates to be determined.
[0045] After constructing the structure of the linear equation system, the specific values of the same source target point pairs, namely the horizontal and vertical pixel coordinates of the target in the first view and the horizontal and vertical pixel coordinates of the target in the second view, are substituted into the above linear equation system as known quantities.
[0046] Due to pixel discretization errors and system noise in actual imaging processes, the lines of sight from the first and second perspectives often cannot perfectly intersect at a single point in three-dimensional space, resulting in the linear equations typically lacking an exact zero solution. Therefore, this embodiment employs a numerical optimization method to obtain the optimal approximate solution. Specifically, the linear equations are solved using the least squares method or singular value decomposition (SVD).
[0047] If singular value decomposition is used, the constructed coefficient matrix will be decomposed, and the right singular vector corresponding to the smallest singular value will be taken as the least squares solution of the linear equation system. This solution corresponds to the point in 3D space that minimizes the sum of the squared distances from the two line-of-sight points (which are emitted by the optical centers of the two cameras through their corresponding pixels).
[0048] The solution vector obtained through the above solution process, after normalization (converting homogeneous coordinates to non-homogeneous coordinates), yields the specific location data of the original target point in the world coordinate system, i.e., the candidate three-dimensional spatial coordinates of the target. These coordinates include the target's lateral position, longitudinal position, and depth information.
[0049] In step S150, reprojection error verification is performed based on the candidate three-dimensional spatial coordinates of the target to determine the three-dimensional detection result.
[0050] In one embodiment of this application, since mismatches may occur during multi-view matching, the candidate 3D spatial coordinates calculated solely by triangulation may not be accurate, and may even be false targets generated by background noise. Therefore, it is necessary to perform reprojection error verification based on the candidate 3D spatial coordinates of the target to determine the final valid 3D detection result. Specifically, as follows: Using the first-view projection matrix and the second-view projection matrix, the calculated candidate 3D spatial coordinates of the target are reverse-mapped. For example, simulating the imaging process of a camera, the coordinates of the candidate point in 3D space are substituted into the projection models of the first and second views, respectively. Through this mathematical replay, the pixel position where the 3D point should theoretically appear in the first-view and second-view images is calculated. The result of this step is called the reprojected pixel coordinates. It represents the ideal projection point on the image plane if the 3D point actually exists at the calculated position.
[0051] The theoretically calculated reprojected pixel coordinates are compared with the target position coordinates (i.e., the center point coordinates of the original detection box) identified in the actual image processing steps. The Euclidean distance between these two two-dimensional coordinate points on the image plane is calculated. This distance value intuitively reflects the degree of deviation between the theoretically derived position and the actual observed position, which is the reprojection error. This error is calculated separately for the first and second viewpoints, or the combined average error of the two is calculated as the final metric.
[0052] The calculated reprojection error is compared with a pre-set empirical value—a preset pixel threshold—to determine the authenticity of the candidate 3D spatial coordinates. When the calculated reprojection error is less than the preset pixel threshold, it means that the theoretically derived position and the actual observed position highly coincide, indicating strong geometric consistency. At this point, the candidate 3D spatial coordinates are determined to be the correct solution result of a real physical target, and are confirmed as a 3D detection result and retained for subsequent size analysis.
[0053] When the calculated reprojection error is greater than or equal to the preset pixel threshold, it means that although the 3D point is mathematically the intersection of two lines of sight, it deviates from the actual observation in terms of geometric projection. This is most likely due to ghosting or noise caused by mismatch at the intersection of lines of sight. In this case, the candidate 3D spatial coordinates are determined to be false targets, and the corresponding candidate 3D spatial coordinates are directly discarded and not output as the final detection result.
[0054] This verification process using proof by contradiction can effectively filter out erroneous intersections in space, significantly improving the robustness of the detection system.
[0055] In S160, the three-dimensional detection results are compared with preset standard process parameters to obtain the identification results.
[0056] In one embodiment of this application, the three-dimensional detection result only represents the actual position coordinates of the weld point or stud in physical space. To determine whether this position meets the process requirements of automobile manufacturing, a further quality assessment process is needed. Specifically, the process of comparing the three-dimensional detection result with preset standard process parameters to obtain the identification result is mainly accomplished based on the calculation of spatial distance deviation and the determination of tolerance range, as follows: The system invokes the standard design coordinates contained in the preset standard process parameters. These standard design coordinates are determined based on the theoretical digital model of the car body (e.g., a CAD model) and represent the ideal absolute position that the weld point or stud should be in. The actual detected 3D inspection results (i.e., the actual coordinates of the target) are compared with the standard design coordinates. Specifically, the straight-line distance between these two 3D coordinate points in 3D Euclidean space is calculated. This distance objectively quantifies the degree of deviation of the actual production position from the theoretical design position and is technically defined as the target space distance deviation. The smaller this deviation value, the higher the actual production accuracy.
[0057] After obtaining the deviation value, a threshold judgment is made on the spatial distance deviation of the target based on a preset allowable tolerance range, thereby generating the final recognition result. The recognition result specifically includes two cases: a first recognition result and a second recognition result. Generate the first identification result (pass / fail judgment): When the calculated target spatial distance deviation is within the allowable tolerance range (e.g., the deviation value is less than or equal to the tolerance threshold), the installation or welding position of the target is determined to meet the process quality requirements. At this time, the first identification result is generated. In practical applications, this result indicates that the target detection is qualified and the position is accurate. Based on this, a Pass or normal signal can be output, allowing the production line to continue to the next process.
[0058] Generating a second identification result (anomaly determination): When the calculated target spatial distance deviation is outside the allowable tolerance range (e.g., the deviation value is greater than the tolerance threshold), it is determined that the actual position of the target has seriously deviated from the design position, which is a process defect (such as weld point drift or stud misalignment). At this time, a second identification result is generated. In practical applications, this result indicates that the target detection is abnormal or unqualified, and can output Fail, alarm, or NG signals to prompt manual re-inspection or trigger an automatic rejection mechanism.
[0059] The method in this embodiment not only achieves target location but also enables automated quality control.
[0060] Existing technologies for binocular or multi-view visual matching often require searching the entire image area, resulting in high computational costs and a high risk of false matches. The method in this application, based on the epipolar geometric constraint principle of the fundamental matrix, compresses the search range for homologous targets in the second view from the entire two-dimensional image plane to the vicinity of an epipolar line. This not only significantly reduces computational costs and improves real-time response speed but also geometrically eliminates interference from most non-homogeneous points, greatly reducing the probability of false target matching in both left and right views.
[0061] To address the problem of traditional triangulation algorithms being susceptible to image noise and generating false 3D points (i.e., points that are mathematically solvable but physically nonexistent), the method in this application utilizes a reprojection error verification step. By mapping the calculated candidate 3D spatial coordinates back to the original image plane and calculating the Euclidean distance between the theoretical projection point and the actual observation point, the confidence level of the 3D solution can be quantitatively evaluated. By eliminating candidate points with excessive reprojection errors, false targets caused by line-of-sight intersection errors or background noise are effectively filtered out, ensuring that the final output 3D detection result has extremely high physical authenticity and geometric accuracy. Moreover, it goes beyond simply obtaining the 3D coordinates of the target; it further compares the 3D detection result with preset standard process parameters (standard design coordinates and allowable tolerance range). By calculating the target spatial distance deviation and combining it with tolerance thresholds for automatic discrimination, this application can directly output a quality diagnosis conclusion of qualified (first identification result) or abnormal (second identification result). This allows the method to directly replace manual quality control of automotive body welds and studs, achieving full automation, digitization, and objectification of the inspection process, avoiding the subjective errors and inefficiencies of manual visual inspection.
[0062] In one embodiment of this application, acquiring first image data of automotive parts from multiple perspectives includes: Multiple image acquisition units are used to capture images of automotive parts from different preset angles, obtaining multiple raw images; Geometric distortion correction is performed on multiple original images based on pre-calibrated camera intrinsic parameters and distortion coefficients to obtain a multi-view corrected image sequence. The corrected image sequence is divided into multiple non-overlapping rectangular sub-blocks, and the histogram of each sub-block is obtained. The histogram is cropped and the number of pixels exceeding a preset threshold is evenly distributed to other gray levels of the histogram to obtain the processed sub-blocks; The processed sub-blocks are fused using bilinear interpolation to obtain first image data of automotive parts from multiple perspectives.
[0063] In one embodiment of this application, to ensure that subsequent detection algorithms can obtain high-fidelity, high-contrast input data, acquiring first image data of automotive parts from multiple perspectives includes image acquisition, geometric distortion correction, and adaptive histogram equalization processing, specifically: Multiple image acquisition units (such as high-resolution industrial cameras) deployed on the automotive production line are used to simultaneously or asynchronously capture images of target automotive parts (such as body side panels, floor panels, etc.) from different preset angles. This multi-angle imaging acquires multiple raw images containing features such as weld points or studs. These raw images capture the shape of the target object from different perspectives, but also include geometric distortion introduced by the lens and noise from ambient lighting.
[0064] Due to the physical characteristics of the lens group in industrial cameras, the original images often exhibit radial distortion (barrel or pincushion distortion) and tangential distortion. To eliminate these distortions, geometric distortion correction is performed on multiple original images using a pre-calibrated camera intrinsic parameter matrix (including focal length and principal point coordinates) and distortion coefficients.
[0065] The internal parameter calibration determines the camera's internal optical parameters. The calibration process uses the Zhang Zhengyou calibration method, and the calibration steps are as follows: First, prepare the calibration board. Use a checkerboard calibration board with 9×7 internal corner points, each square measuring 30 mm. The calibration board needs to be flat and have a non-reflective surface. Use matte paper when printing the calibration board to avoid glare affecting corner point detection.
[0066] Then, acquire calibration images. Place the calibration board at different positions and angles, and take 20-30 calibration images with each camera. The images should cover all areas of the camera's field of view, and the tilt angle of the calibration board should vary between 0 and 45 degrees. Ensure that the calibration board is fully visible in each image, occupying 20%–80% of the image area.
[0067] Next, corner detection is performed. For each calibration image, a sub-pixel corner detection algorithm is used to extract the checkerboard corners. Corner detection involves two steps: first, the Harris corner detector is used to find the initial corner positions, and then a sub-pixel optimization algorithm is used to improve the corner localization accuracy to within 0.1 pixels. The image coordinates of the corners are denoted as (u... i v i) The corresponding world coordinates are (Xi, Y). i ,0)(The calibration plate plane is defined as Z=0).
[0068] Calculate the camera intrinsic parameter matrix. Based on the pinhole camera model, the projection relationship from world coordinates to image coordinates is as follows:
[0069] Where s is the scale factor (depth value of the object point). and It is the equivalent focal length (in pixels) of the camera along the horizontal and vertical axes of the image, and is composed of the physical focal length f and the pixel size d. x d y Decide: =f / d x , =f / d y . ( , The coordinates are the principal point coordinates, i.e., the pixel coordinates of the intersection of the optical axis and the imaging plane. Ideally, these should be located at the center of the image, but in reality, they may be offset. ] are the three column vectors of the rotation matrix, and t is the translation vector.
[0070] Solving for the intrinsic parameter matrix by minimizing the reprojection error:
[0071] Where n is the number of images, and m is the number of corner points in each image. These are the coordinates of the detected corner point. These are the projected coordinates calculated based on the parameters. The Levenberg-Marquardt algorithm is used for nonlinear optimization.
[0072] Distortion parameter calibration. Actual lenses exhibit distortion, primarily radial and tangential. Radial distortion is caused by the lens shape, manifesting as barrel or pincushion distortion. Tangential distortion is caused by the non-parallelism between the lens and the image plane. The distortion model is as follows:
[0073]
[0074] in, It is the theoretical distortion-free normalized coordinate. These are the actual distortion-normalized coordinates. The distance to the principal point is given by k1, k2, and k3, which are radial distortion coefficients, and p1 and p2 are tangential distortion coefficients. These five parameters are solved together with the intrinsic parameter matrix using an optimization algorithm. Extrinsic parameter calibration determines the camera's position and attitude in the world coordinate system. The extrinsic parameters include the rotation matrix R and the translation vector T. The rotation matrix is a 3×3 orthogonal matrix that satisfies... The rotation matrix can be represented by a rotation vector (Rodrigues vector), where the direction of the rotation vector is the rotation axis and the magnitude is the rotation angle. The translation vector represents the position of the camera's optical center in the world coordinate system.
[0075] Multi-camera systems need to be unified to the same world coordinate system. A fixed reference point on the production line is chosen as the origin of the world coordinate system. The positional relationships between the cameras are established by placing calibration plates in the fields of view of multiple cameras. Let camera 1 be the reference camera, and the positional relationships of the other cameras relative to camera 1 are as follows:
[0076] Among them, P i R is the coordinate of the point in the camera's i-coordinate system. 1i and T 1i It is the rotation and translation of camera i relative to camera 1.
[0077] Establishing a complete coordinate system is fundamental to achieving precise positioning and multi-camera collaboration. In this embodiment, four coordinate systems are defined, and the transformation relationships between them are established.
[0078] World coordinate system (X) w ,Y w Z w) This is the reference coordinate system for all measurements. The origin is set at a fixed reference point on the production line, such as a locating pin at a testing station. X w The axis is along the conveying direction of the production line, pointing downstream; Y w The axis is vertically upward; Z w The axis is determined by the right-hand rule and points towards the operating side. The coordinate unit is millimeters. This coordinate system is consistent with the machine coordinate system of the production line, facilitating coordination with other equipment. Camera coordinate system (X... c ,Y c Z c (Z) The origin is the optical center of the camera. c The axis is along the optical axis and points towards the object being photographed; X c Axis to the right; Y c The axis is downward. This is a three-dimensional coordinate system, with units in millimeters. The camera coordinate system is defined according to computer vision conventions.
[0079] The image coordinate system (x, y) is a two-dimensional coordinate system on the imaging plane. The origin is located at the center of the imaging plane (principal point), the x-axis points to the right, and the y-axis points downwards. The coordinate unit is millimeters. This coordinate system describes the position on the physical imaging plane.
[0080] The pixel coordinate system (u, v) is the coordinate system for a digital image. The origin is located at the top left corner of the image, the u-axis points to the right (corresponding to the column), and the v-axis points downwards (corresponding to the row). The coordinate unit is pixels. This is the coordinate system used in the final digital image.
[0081] The coordinate transformation process is as follows: The transformation from the world coordinate system to the camera coordinate system is a rigid body transformation:
[0082] Where R is a 3×3 rotation matrix, describing the rotation of the camera coordinate system relative to the world coordinate system; T is a 3×1 translation vector, describing the position of the camera optical center in the world coordinate system.
[0083] The transformation from the camera coordinate system to the image coordinate system is perspective projection.
[0084] Where f is the camera's physical focal length (in millimeters). This projection projects a point in three-dimensional space onto a two-dimensional imaging plane.
[0085] The transformation from image coordinates to pixel coordinates involves scaling and translation:
[0086] Where dx and dy are the physical dimensions of a pixel in the x and y directions (millimeters per pixel), and (c x ,c y) These are the pixel coordinates of the principal point.
[0087] In summary, the complete transformation from world coordinates to pixel coordinates can be represented by homogeneous coordinates as follows:
[0088] Where K is the intrinsic parameter matrix, [R|T] is a 3×4 extrinsic parameter matrix, and s is the scale factor (equal to Z). c ).
[0089] Through the above calculations, each pixel in the original image is mapped to a new position, resulting in a corrected image sequence with realistic geometric shapes from multiple perspectives.
[0090] Quantitative quality assessment is performed on the corrected image to provide a basis for subsequent processing decisions. The quality assessment is conducted from multiple dimensions, each with a clear physical meaning and calculation method. Sharpness assessment reflects the richness of detail in the image. Sharp images contain more high-frequency information and have sharp edges, which is beneficial for object detection. This application adopts a sharpness evaluation method based on the Laplacian operator. The Laplacian operator is a second-order differential operator that is sensitive to image edges and texture. The discrete form of the Laplacian operator is defined as:
[0091] In digital images, approximate calculations are performed using convolution kernels:
[0092] This kernel detects the second derivative in both the horizontal and vertical directions. For more precise detection, an extended kernel can be used:
[0093] Where R(i,j) is the response value at position (i,j).
[0094] The sharpness score is defined as the variance of the response map:
[0095] Where M and N are the height and width of the image. It is the mean of the response map. The larger the variance, the richer the edges and details contained in the image, and the higher the clarity.
[0096] To normalize, define the clarity score:
[0097] in, This is an empirical maximum value, obtained through statistical analysis of a large number of high-quality images; a typical value is 1000. There are various methods for calculating contrast; this invention uses a method based on standard deviation.
[0098] in, It is the standard deviation of the image grayscale values. This is the average gray value. This ratio is called the coefficient of variation, which reflects the degree of dispersion of the gray value distribution.
[0099] More precise contrast calculations consider local contrast. Divide the image into k×k blocks (e.g., 32×32), calculate the contrast of each block, and then average them:
[0100] in, It is a small constant (such as 0.01) to avoid division by zero errors.
[0101] Contrast Rating:
[0102] in This is the target contrast value, which is set according to the requirements of the detection task; a typical value is 0.3.
[0103] Brightness assessment checks the overall brightness level of the image. Images that are too bright or too dark will affect detection results. Ideal brightness should fully utilize the dynamic range while avoiding underexposure or overexposure.
[0104] Average brightness:
[0105] For 8-bit images, the ideal average brightness is between 100 and 150.
[0106] Histogram analysis was used to assess luminance distribution. The 10th and 90th percentiles were calculated.
[0107]
[0108] Ideally, B10 > 20 (to avoid underexposure) and B90 < 235 (to avoid overexposure).
[0109] Brightness scoring takes into account both average brightness and distribution:
[0110] The first term is the Gaussian function, which reaches a maximum value of 1 when the average brightness is 128; the second term penalizes underexposure and overexposure.
[0111] The overall quality score is obtained by weighted averaging:
[0112] The weighting coefficients are set considering the degree of influence of each indicator on the detection effect. After extensive experimental verification, the weighting coefficients are set... =0.4 (Clarity is the most important factor). =0.3, =0.3. When When the value is less than 0.6, the image quality is deemed unacceptable and enhancement processing is required.
[0113] To address the issue of unclear target features caused by reflections or shadows on the surface of metal workpieces, this application employs the Limiting Contrast Adaptive Histogram Equalization (CLAHE) algorithm to enhance the image.
[0114] Each frame in the corrected image sequence is spatially divided into multiple non-overlapping rectangular sub-blocks. For each rectangular sub-block, its grayscale distribution is statistically analyzed to obtain a histogram for each sub-block.
[0115] To prevent excessive noise amplification by conventional histogram equalization, the histogram needs to be limited. A preset threshold is set. For each gray level of the histogram, pixels exceeding the preset threshold are cropped, and the total number N of cropped pixels is calculated. clip .
[0116] Then, these Nclip pixels are evenly distributed across all gray levels of the histogram. Let the total number of gray levels be L (usually 256), then the number of pixels added to each gray level is N. avg =N clip / L. This yields the corrected histogram, i.e., the processed sub-block. This step both improves contrast and suppresses the enhancement of background noise.
[0117] If the processed sub-blocks are directly stitched together, the image will exhibit obvious blocky boundary effects. Therefore, this application employs a bilinear interpolation algorithm to fuse the sub-blocks. For any point in the image, a weighted calculation is performed using the transformation functions of the center points of its four surrounding sub-blocks, based on its relative position within the sub-block.
[0118] The formula for calculating bilinear interpolation is as follows:
[0119] Where f(x,y) is the final gray value at the pixel (x,y) to be determined; Q11 Q 21 Q 12 Q 22 These are the mapping values for the center points of the four adjacent sub-blocks surrounding this pixel; (x1, y) 1) and (x2,y) 2) The coordinates of the top left and bottom right reference points are (x2-x) 1) and (y2-y 1) This represents the width and height of the interpolation region.
[0120] The bilinear interpolation described above smooths the grayscale transitions between sub-blocks, ultimately generating clear, artifact-free, and feature-rich first-view image data of automotive parts from multiple perspectives, which is then input into a subsequent designated model. Contrast evaluation reflects the degree of brightness difference in an image; in high-contrast images, the distinction between the target and the background is better, which is beneficial for detection.
[0121] In one embodiment of this application, the method for constructing the specified model includes: Obtain the initial model and training set; The first model is obtained by embedding a channel attention module and a spatial attention module into the backbone feature extraction network of the initial model. The channel attention module is used to weight each channel of the feature map, and the spatial attention module is used to weight the spatial position of the feature map. In the feature fusion network of the first model, a shallow high-resolution feature branch is set to detect small targets, resulting in the second model. The shallow high-resolution feature branch integrates deep semantic information and shallow texture information. The shallow high-resolution feature branch is used to detect solder joints or studs with pixel sizes smaller than a preset threshold. Based on the second model and the training set, the predicted bounding boxes and class prediction probabilities are obtained; Determine the ground truth bounding boxes based on the training set; The second model is trained using a multi-task joint loss function, which includes a localization loss component. The intersection-union ratio (IUU) is obtained by calculating the ratio of the overlap area between the predicted bounding box and the true bounding box to the union area. Calculate the Euclidean distance between the center point of the predicted bounding box and the center point of the true bounding box, and obtain the diagonal distance of the smallest closure region that can simultaneously contain both bounding boxes. Divide the Euclidean distance by the diagonal distance for normalization and obtain the center point distance penalty term. Calculate the difference between the aspect ratio of the predicted bounding box and the aspect ratio of the true bounding box to obtain the aspect ratio penalty term; The localization loss value of the sample is obtained by subtracting the center point distance penalty term and the aspect ratio penalty term from the intersection-union ratio. A dynamic modulation factor is constructed based on the category prediction probability, and the dynamic modulation factor is multiplied with the standard binary cross-entropy loss function term to obtain the classification loss part. The second model is optimized based on the localization loss value, the classification loss component, and the training set to obtain the specified model.
[0122] In one embodiment of this application, in order to achieve high-precision detection of automotive weld points and studs, and in particular to solve the problems of missed detection of small targets and imbalance of positive and negative samples, an improved convolutional neural network (i.e., a specified model) is used for processing.
[0123] Obtain a basic initial model (e.g., a backbone network based on ResNet or Darknet architecture) and a labeled training set. To enhance the model's ability to extract key features, a dual attention mechanism is embedded in the backbone feature extraction network of the initial model, thus obtaining the first model. For example: The standard YOLOv11 network is designed for general object detection, and its performance is unsatisfactory when directly applied as an initial model for solder joint and stud detection. In this embodiment, the following targeted optimizations were performed: The standard YOLOv11 network architecture consists of three main parts: the Backbone feature extraction network, the Neck feature fusion network, and the Head detection output network. The Backbone network is responsible for extracting multi-scale feature representations from the input image. The Backbone uses an improved version of the CSPDarknet architecture, reducing computational cost while maintaining feature extraction capabilities through cross-stage partial connections.
[0124] The network receives a 640×640×3 RGB image as input. This size is carefully chosen: large enough to preserve detail in small objects without causing excessive computation. If the original image size is different, it is adaptively scaled and padded to 640×640. The padded image uses gray (RGB value 114), which is the average color of the COCO dataset.
[0125] The Focus module is a unique design feature of YOLOv11, achieving efficient downsampling through the conversion of spatial information to channel information. Specifically, the 640×640×3 input image is divided into four 320×320×3 sub-images according to odd and even rows and columns. The segmentation method is as follows: son Figure 1 : Pixels in even-numbered rows and even-numbered columns; sub Figure 9 : Pixels in even rows and odd columns; sub Figure 3 : Pixels in odd-numbered rows and even-numbered columns; sub Figure 4The feature map consists of pixels in odd-numbered rows and columns. These four sub-images are then concatenated along the channel dimension to obtain a 320×320×12 feature map. Finally, it is compressed to 320×320×32 using a 3×3 convolution (32 filters). This design achieves 2x downsampling without information loss while reducing computational cost.
[0126] The CSP (Cross Stage Partial) module is a core component of Backbone. Its design philosophy is to divide the feature map into two parts: one part is processed through densely connected residual blocks, and the other part is directly forward-propagated. Finally, the two parts are fused. This design reduces redundant gradient information and improves learning efficiency.
[0127] Mathematical representation of the CSP module: Let the input feature map be... First, it is divided into two parts using a 1×1 convolution:
[0128]
[0129] X2 is processed through n residual blocks:
[0130] The structure of each residual block is as follows:
[0131] Where BN is batch normalization and ReLU is the activation function.
[0132] Finally, merge the two branches:
[0133] The Spatial Pyramid Pooling (SPP) module expands the receptive field through multi-scale pooling, capturing contextual information across different ranges. Implementation of the SPP module:
[0134] Pooling operations at different scales capture different ranges of contextual information: 5×5 captures local features, 9×9 captures medium-range features, and 13×13 captures large-range features. Pooling uses the same padding to keep the feature map size constant.
[0135] The Neck network employs a PANet (Path Aggregation Network) structure to achieve bidirectional feature transfer and fusion, ensuring that each layer contains rich semantic information and precise localization information.
[0136] The top-down path (FPN) propagates high-level semantic features downwards. Let the feature layers output by the backbone be {C3, C4, C5}, with corresponding feature map sizes of 80×80, 40×40, and 20×20. The calculation process of FPN is as follows:
[0137]
[0138]
[0139] Upsample2× is a 2x upsampling, implemented using nearest neighbor interpolation or transposed convolution.
[0140] Bottom-up path (PAN) propagates low-level localization features upwards:
[0141]
[0142]
[0143] It is a 2x downsampling, implemented using convolution with a stride of 2.
[0144] This bidirectional transmission mechanism enables each layer to integrate multi-scale information, thereby improving detection performance.
[0145] The Head network performs detection predictions at three different scales, each responsible for detecting targets within a specific size range. The division of labor among the three detection scales is as follows: P3 (80×80): Detects small targets (8-32 pixels), suitable for detecting small solder joints; P4 (40×40): Detects medium-sized targets (32-96 pixels), suitable for detecting standard solder joints and studs; P5 (20×20): Detects large targets (96 pixels or more), suitable for detecting large studs or solder joint groups; The output of each detector head is a K×K×(B×(5+C)) tensor, where: K is the spatial size of the feature map (80, 40, or 20). B is the number of anchors per grid cell (default is 3); 5 includes 4 bounding box regression parameters (tx, ty, tw, th) and 1 target confidence level to; C is the number of categories (in this task, C=2, namely solder joints and studs); Bounding box decoding process: Let the top-left corner coordinates of the grid cell be (cx, cy), and the width and height of the anchor be (pw, ph). Then the center coordinates and width and height of the predicted box are:
[0146]
[0147]
[0148]
[0149] Where σ is the Sigmoid function, which restricts the output to the range [0,1].
[0150] Confidence level and class probability:
[0151]
[0152] Add a CBAM (Convolutional Block Attention Module) attention module to the key layers of the backbone (after the outputs of C3, C4, and C5) to enable the network to adaptively focus on important feature regions. CBAM consists of two sub-modules: channel attention and spatial attention.
[0153] The channel attention module learns the importance weights of different channels. For the input feature map... First, global pooling is performed to obtain channel statistics:
[0154]
[0155] Then, the weights of the shared MLP learning channel are used:
[0156] The MLP consists of two fully connected layers, with the number of neurons in the middle layer being C / r (where r is the compression ratio, typically 16):
[0157]
[0158] The spatial attention module learns the importance weights of different spatial locations. First, pooling is performed along the channel dimension:
[0159]
[0160] Then, spatial weights are learned through 7×7 convolutions:
[0161] Final attention weighting:
[0162] Small target detection optimization. Solder joints are typically only 10-30 pixels in size, posing a challenge for small target detection. This invention improves small target detection performance through the following strategies: A P2 inspection layer (160×160) is added specifically for inspecting tiny solder joints. The P2 layer is constructed as follows:
[0163]
[0164]
[0165] Design anchors tailored to solder joint sizes. Optimize anchor sizes by using K-means clustering analysis to train a dataset on the size distribution of solder joints. Layer P2: [6×6, 10×10, 14×14]; Layer P3: [18×18, 25×25, 35×35]; Layer P4: [45×45, 65×65, 90×90]; Layer P5: [120×120, 180×180, 250×250] Deformable convolution is employed to enhance the modeling capability for deformable targets. While standard convolution uses a fixed grid for sampling, deformable convolution allows the sampling positions to adaptively adjust based on the input features.
[0166] Where p0 is the output position, R is the standard grid (e.g., 3×3), Δpn is the learned offset, and w(pn) is the convolution weight.
[0167] The offset is learned through additional convolutional layers:
[0168] This allows the convolution kernel to adapt to the actual shape of the target, making it particularly effective for irregularly shaped solder joints.
[0169] Multi-scale training strategy. During training, the input image size is randomly changed (within the range of 448-832 pixels, with a step size of 32), allowing the model to adapt to inputs of different scales. This improves the model's scale invariance.
[0170] The channel attention module is used to establish the interdependencies between feature channels. By weighting each channel of the feature map, it automatically strengthens channels containing solder joint texture information and suppresses background noise channels. At the same time, the embedded spatial attention module is used to focus on the spatial position of the feature map. It weights the spatial position of the feature map, making the network pay more attention to areas where the target may exist (such as metal surface joints).
[0171] Building upon this foundation, to address the issue of feature loss in small solder joints after multiple samplings, the system incorporates a shallow high-resolution feature branch within the feature fusion network (Neck part) of the first model, resulting in the second model. This branch avoids depth downsampling, preserving the shallow texture information of the image (high resolution, clear geometric contours) and fusing it with the deep semantic information (low resolution, strong category features) recovered through upsampling. This shallow high-resolution feature branch is specifically designed to detect tiny solder joints or studs with pixel sizes smaller than a preset threshold, ensuring that no small targets are missed.
[0172] The second model is used to perform forward inference on the training set, outputting predicted bounding boxes and class prediction probabilities, and the loss is calculated by combining the ground truth bounding boxes in the training set. The localization loss adopts CIoU (Complete Intersection over Union) loss logic that considers overlap, center distance, and aspect ratio.
[0173] Calculate the predicted bounding box (B pred ) and true bounding box (B gt The intersection-union ratio (IoU) is the ratio of the area of overlap (intersection) to the area of union (union) of two sets of objects. The formula is as follows:
[0174] To accelerate convergence and address the gradient vanishing problem when two bounding boxes do not intersect, a center point distance constraint is introduced. The center point (x, y) of the predicted bounding box and the center point (x, y) of the ground truth bounding box are calculated. gt ,y gt The Euclidean distance between the two bounding boxes is calculated. Simultaneously, the smallest closure region (i.e., the smallest bounding rectangle) that can simultaneously contain both bounding boxes is obtained, and the diagonal distance c of this closure region is calculated.
[0175] The square of the Euclidean distance is divided by the square of the diagonal distance for normalization, yielding the center point distance penalty term R. DIoU :
[0176] Where ρ(⋅) represents the Euclidean distance calculation, and b and b gtThese represent the center points of the predicted bounding box and the ground truth bounding box, respectively.
[0177] To make the shape of the predicted bounding box closer to the ground truth bounding box, an aspect ratio consistency constraint is introduced. The difference between the aspect ratio of the predicted bounding box and the ground truth bounding box is calculated, resulting in an aspect ratio penalty term. This term is determined by parameters α and v.
[0178] Where v is used to measure the consistency of aspect ratio:
[0179] w gt ,h gt represents the width and height of the ground truth bounding box, and w and h represent the width and height of the predicted bounding box.
[0180] Combining the above three factors, the CIoU value is obtained by subtracting the center point distance penalty and aspect ratio penalty from the intersection-union ratio, and then the final localization loss value Lloc is constructed:
[0181] The classification loss mechanism employs the concept of Focal Loss to address the imbalance between easy and difficult samples. The system constructs a dynamic modulation factor based on the class prediction probability p output by the model. Let p... t This represents the probability value predicted by the model for the true class. The standard binary cross-entropy loss function term (BCE) is -log(p). t ).
[0182] The dynamic modulation factor introduced in this application is (1-pt)γ, where γ is a focusing parameter (e.g., a value of 2). The classification loss component L is obtained by multiplying the dynamic modulation factor with the standard binary cross-entropy loss function. cls : The physical meaning of this formula is: when samples are easily distinguishable (p) t When the sample is close to 1, the modulation factor tends to 0, thus reducing the contribution of that sample to the total loss; when the samples are difficult to distinguish (p t (Smaller), the modulation factor is close to 1, maintaining its high weight, forcing the model to focus on difficult-to-distinguish samples.
[0183] The total loss is obtained by weighted summing of the localization loss and classification loss. The backpropagation algorithm is then used to update and optimize the weight parameters of the second model based on this total loss until the model converges, resulting in the final model used for actual detection.
[0184] In one embodiment of this application, the detection results from each viewpoint include the detection results from a first viewpoint and the detection results from a second viewpoint. Based on the epipolar geometry constraint principle, the detection results from each viewpoint are filtered to obtain pairs of target points from the same origin, including: Based on the relative positions of the acquisition units, the fundamental matrix is obtained; The center point pixel coordinates of the specified target in the first-view detection result are used as input, and matrix multiplication is performed using the fundamental matrix to obtain the epipolar equation of the center point of the specified target in the first-view detection result. Traverse the detection results from the first-person perspective to obtain the set of center point pixel coordinates of all candidate targets in the same category group, and obtain the set of center point pixel coordinates of the candidate targets; For each candidate target's center point in the candidate target's center point pixel coordinate set, the vertical pixel distance from the corresponding epipolar line is obtained; If the specified vertical pixel distance is less than the preset tolerance threshold, then the candidate target in the second viewpoint is determined to form a candidate homogeneous point pair with the target in the first viewpoint, and a homogeneous target point pair is obtained.
[0185] In one embodiment of this application, after obtaining the single-view detection results of the first and second perspectives respectively through a deep learning model, in order to determine whether the targets detected in the two perspectives are the same object in physical space (i.e., whether they are common source points), the principle of epipolar geometry is used for screening.
[0186] Based on the relative positions of the acquisition units (i.e., binocular or multi-view cameras), the fundamental matrix (F) describing the geometric relationship between the two views is calculated. The fundamental matrix F contains the intrinsic parameters of the cameras and the rotation and translation relationships between the two cameras. The calculation of the fundamental matrix F: Given the intrinsic parameter matrices K1, K2 and the relative pose (R, t) of the two cameras, the fundamental matrix F is:
[0187] Among them, K -1 Represents the inverse of a matrix. -T R represents the transpose of the matrix inverse; R is the rotation matrix from the first-view camera coordinate system to the second-view camera coordinate system; t is the translation vector from the first-view camera coordinate system to the second-view camera coordinate system; [t]× is the antisymmetric matrix of the translation vector, used to convert the vector cross product into matrix multiplication.
[0188] This formula transforms the relative positional relationship between the two cameras in physical space into a mathematical algebraic constraint matrix F.
[0189] Take the center pixel coordinates of a specified target from the first-view detection results, and denote them as p=(u,v,1). T (Homogeneous coordinate form). Using the fundamental matrix F, matrix multiplication is performed on this coordinate to obtain the epipolar equation corresponding to the point on the second-view image plane.
[0190] According to the epipolar geometric constraint formula:
[0191] The calculated vector That is, the epipolar parameters on the second-view image, let... =(A,B,C) T This vector defines a straight line (i.e., an epipolar line) on the second-view image plane, and its general equation is:
[0192] Where x and y are the variable coordinates on the second-view image plane, and A, B, and C are the coefficients of the line. The physical meaning of this epipolar line is that the physical point corresponding to the target in the first view must be projected onto this line in the second view.
[0193] After obtaining the epipolar equation, it is necessary to determine whether the target detected in the second-view perspective is located on this line.
[0194] The system iterates through the detection results from the first-view perspective, collecting the center point pixel coordinates of all candidate targets grouped within the same category (e.g., all solder joints or all studs), to obtain the set of center point pixel coordinates for the candidate targets. Let the center point coordinates of a candidate target in the second-view perspective be p′=(u′, v′). The system calculates the distance from this point to the corresponding epipolar line mentioned above. The vertical pixel distance d. The distance calculation formula is as follows:
[0195] Where: |·| represents absolute value operation, and the numerator represents algebraic distance; is the normalization factor used to convert algebraic distance to geometric Euclidean distance; d is the specified vertical pixel distance, in pixels.
[0196] The calculated distance dd is compared with a preset tolerance threshold (e.g., 1 to 3 pixels): If the specified vertical pixel distance d is less than the preset tolerance threshold, it indicates that the candidate point is adjacent to the epipolar line and meets the epipolar geometric constraint. At this time, it is determined that the candidate target in the second viewpoint and the target in the first viewpoint constitute a candidate homologous point pair, thus obtaining a homologous target point pair. This means that these two points are very likely projections of the same solder joint in physical space at different angles, and can proceed to subsequent triangulation for three-dimensional coordinate calculation.
[0197] In one embodiment of this application, the candidate three-dimensional spatial coordinates of the target are obtained by calculating the points of the same source target using a triangulation algorithm, including: Construct a system of linear equations containing the first-view projection matrix and the second-view projection matrix; Substitute the pixel coordinates of the target point pairs that satisfy the same source into the linear equation system, and solve the linear equation system using the least squares method or singular value decomposition method to obtain the candidate three-dimensional spatial coordinates of the target.
[0198] In one embodiment of this application, the detection result includes the target's position coordinates. Reprojection error verification is performed based on the target's candidate three-dimensional spatial coordinates to determine the three-dimensional detection result, which includes: The candidate 3D spatial coordinates of the target are projected back onto their respective image planes through the first-view projection matrix and the second-view projection matrix to obtain the reprojected pixel coordinates. The reprojection error is obtained by calculating the Euclidean distance between the reprojected pixel coordinates and the target position coordinates. If the reprojection error is less than the preset pixel threshold, the candidate 3D spatial coordinates of the corresponding target are determined to be true, and the 3D detection result is obtained. If the reprojection error is greater than or equal to a preset pixel threshold, the candidate 3D spatial coordinates of the corresponding target are determined to be false, and the candidate 3D spatial coordinates of the corresponding target are removed.
[0199] In one embodiment of this application, the three-dimensional position of the target in the world coordinate system is recovered by using a pair of target points from the same source, and the accuracy of the position is verified.
[0200] Obtain the first-view projection matrix P1 and the second-view projection matrix P2. The projection matrix describes the mathematical transformation relationship between 3D world coordinates and the 2D image plane, and is usually determined by both camera intrinsic and extrinsic parameters. Let the 3D spatial coordinates of the target be X. w =[X,Y,Z,1] T (Homogeneous coordinates). Let the pixel coordinates of the target in the first-person view be (u1, v). 1) In the second-view perspective, the pixel coordinates of the target from the same source are (u2, v2). According to the pinhole camera imaging model, the following projection relationship exists:
[0201] Where s is the scale factor, and xx is the pixel coordinate vector [u,v,1] T .
[0202] Expanding this formula, for each point in every viewpoint, the scale factor s can be eliminated, thus constructing a linear constraint equation regarding X, Y, and Z. The specific derivation is as follows:
[0203]
[0204] Where p ij This represents the element in the i-th row and j-th column of the projection matrix P.
[0205] Rearranging the above equations, for a pair of points of common origin (i.e., two equations from each of the two viewpoints, for a total of four equations), we construct a system of linear equations containing the projection matrices of the first and second viewpoints, in the form A·Xw=0:
[0206] Subsequently, the pixel coordinates of the corresponding source target point pairs are substituted into a system of linear equations, and the system is solved using the least squares method or singular value decomposition (SVD). Due to noise, the system of equations typically does not have an exact zero solution; SVD decomposition is used to find the solution that satisfies the condition. The smallest unit vector X is the candidate three-dimensional spatial coordinate of the target.
[0207] Since the calculated candidate 3D spatial coordinates may be invalid solutions generated by mismatched points (mathematically false intersections), they must be verified. The calculated candidate 3D spatial coordinates X of the target... cand =[X,Y,Z,1] T Substitute the first-view projection matrix P1 and the second-view projection matrix P2 into the matrix multiplication operation respectively:
[0208] The obtained homogeneous coordinates Normalization (divided by the third component) yields the theoretical reprojection pixel coordinates. and( .
[0209] The Euclidean distance between the reprojected pixel coordinates and the original detected target position coordinates is calculated. Taking the first-person view as an example, the reprojection error E1 is calculated using the following formula:
[0210] Similarly, calculate the error E2 of the second perspective, and take the average or maximum value of the two as the final reprojection error.
[0211] If the reprojection error is less than a preset pixel threshold (e.g., 1.0 pixel), it indicates that the calculated 3D point cloud geometrically matches the observation results of the two cameras. At this point, the candidate 3D spatial coordinates of the corresponding target are determined to be true, and these coordinates are retained as the final 3D detection result.
[0212] If the reprojection error is greater than or equal to a preset pixel threshold, it indicates that although the 3D point satisfies the least squares solution, it deviates from the actual observation point in geometric projection, and is very likely a false intersection caused by skewed lines of sight. In this case, the candidate 3D spatial coordinates of the corresponding target are determined to be false, and the candidate 3D spatial coordinates of the corresponding target are directly removed to prevent erroneous data from entering the subsequent quality assessment stage.
[0213] In one embodiment of this application, the identification result includes a first identification result and a second identification result. The three-dimensional detection result is compared with preset standard process parameters to obtain the identification result, which includes: The spatial distance deviation between the three-dimensional detection results and the standard design coordinates is calculated to obtain the target spatial distance deviation; The first identification result is obtained when the target spatial distance deviation is within the allowable tolerance range; When the target spatial distance deviation is outside the allowable tolerance range, a second identification result is obtained.
[0214] In one embodiment of this application, a pre-stored digital model (CAD model) of the automotive parts is invoked to obtain the standard design coordinates corresponding to the current inspection point. Let the calculated 3D inspection result (actual coordinates) be... Let the standard design coordinates (theoretical coordinates) of the target specified in the process document be... .
[0215] The spatial distance deviation between the 3D inspection results and the standard design coordinates is calculated. This deviation characterizes the degree of drift of the actual production position relative to the theoretical design position. The calculation formula uses the 3D Euclidean distance formula:
[0216] ΔD is the target spatial distance deviation, a non-negative scalar quantity, usually measured in millimeters (mm). These represent the component deviations along the X, Y, and Z axes, respectively.
[0217] After obtaining the deviation value ΔD, the system compares it with the preset allowable tolerance range (denoted as the threshold τ): If the spatial distance deviation of the target is within the allowable tolerance range (i.e., ΔD≤τ), it indicates that the installation accuracy of the weld or stud meets the process requirements. At this point, the target is determined to be a qualified product, and the first identification result is obtained. This result corresponds to the Pass or OK status.
[0218] If the spatial distance deviation of the target is outside the allowable tolerance range (i.e., ΔD>τ), it indicates that the actual position of the target exceeds the design allowable error limit, constituting a positional defect. In this case, the system determines the target as a non-conforming product and obtains a second identification result. This result corresponds to a Fail or NG status and can trigger subsequent audible and visual alarms.
[0219] Secondly, embodiments of this application provide a system for identifying automotive welding studs, comprising: The first acquisition module acquires first image data of automotive parts from multiple perspectives. The first generation module performs weld point target and stud target detection on the first image data based on a specified model, and generates detection results from various perspectives. The detection results include the target's position coordinates, target category, and confidence information. The first module obtains pairs of target points from the same source by filtering the detection results from each viewpoint based on the principle of epipolar geometry constraint. The second module obtains the candidate three-dimensional spatial coordinates of the target by using a triangulation algorithm to calculate the points of the same source target; The first determination module is used to verify the reprojection error based on the candidate 3D spatial coordinates of the target and determine the 3D detection result. The third module is used to compare the 3D detection results with preset standard process parameters to obtain the identification results.
[0220] In this embodiment, the method for identifying automotive weld studs includes: acquiring first image data of automotive parts from multiple perspectives; detecting weld targets and stud targets in the first image data based on a specified model, generating detection results for each perspective; filtering the detection results for each perspective based on the epipolar geometric constraint principle to obtain pairs of target points from the same source; calculating the candidate three-dimensional spatial coordinates of the target using a triangulation algorithm; verifying the reprojection error based on the candidate three-dimensional spatial coordinates of the target to determine the three-dimensional detection result; and comparing the three-dimensional detection result with preset standard process parameters to obtain the identification result. In this embodiment, by detecting weld targets and stud targets in the first image data based on a specified model and generating detection results for each perspective, the identification of weld targets and stud targets on automotive parts is achieved, solving the problem of difficult identification of small targets in complex backgrounds. Furthermore, the reprojection error verification of the detection results for each perspective is performed using the epipolar geometric constraint principle and triangulation algorithm, avoiding some occlusion and false detection problems through redundancy, solving the occlusion or false detection problems that may exist in a single perspective, and further improving the robustness and reliability of the detection.
[0221] In one embodiment of this application, acquiring first image data of automotive parts from multiple perspectives includes: Multiple image acquisition units are used to capture images of automotive parts from different preset angles, obtaining multiple raw images; Geometric distortion correction is performed on multiple original images based on pre-calibrated camera intrinsic parameters and distortion coefficients to obtain a multi-view corrected image sequence. The corrected image sequence is divided into multiple non-overlapping rectangular sub-blocks, and the histogram of each sub-block is obtained. The histogram is cropped and the number of pixels exceeding a preset threshold is evenly distributed to other gray levels of the histogram to obtain the processed sub-blocks; The processed sub-blocks are fused using bilinear interpolation to obtain first image data of automotive parts from multiple perspectives.
[0222] In one embodiment of this application, the method for constructing the specified model includes: Obtain the initial model and training set; The first model is obtained by embedding a channel attention module and a spatial attention module into the backbone feature extraction network of the initial model. The channel attention module is used to weight each channel of the feature map, and the spatial attention module is used to weight the spatial position of the feature map. In the feature fusion network of the first model, a shallow high-resolution feature branch is set to detect small targets, resulting in the second model. The shallow high-resolution feature branch integrates deep semantic information and shallow texture information. The shallow high-resolution feature branch is used to detect solder joints or studs with pixel sizes smaller than a preset threshold. Based on the second model and the training set, the predicted bounding boxes and class prediction probabilities are obtained; Determine the ground truth bounding boxes based on the training set; The second model is trained using a multi-task joint loss function, which includes a localization loss component. The intersection-union ratio (IUU) is obtained by calculating the ratio of the overlap area between the predicted bounding box and the true bounding box to the union area. Calculate the Euclidean distance between the center point of the predicted bounding box and the center point of the true bounding box, and obtain the diagonal distance of the smallest closure region that can simultaneously contain both bounding boxes. Divide the Euclidean distance by the diagonal distance for normalization and obtain the center point distance penalty term. Calculate the difference between the aspect ratio of the predicted bounding box and the aspect ratio of the true bounding box to obtain the aspect ratio penalty term; The localization loss value of the sample is obtained by subtracting the center point distance penalty term and the aspect ratio penalty term from the intersection-union ratio. A dynamic modulation factor is constructed based on the category prediction probability, and the dynamic modulation factor is multiplied with the standard binary cross-entropy loss function term to obtain the classification loss part. The second model is optimized based on the localization loss value, the classification loss component, and the training set to obtain the specified model.
[0223] In one embodiment of this application, the detection results from each viewpoint include the detection results from a first viewpoint and the detection results from a second viewpoint. Based on the epipolar geometry constraint principle, the detection results from each viewpoint are filtered to obtain pairs of target points from the same origin, including: Based on the relative positions of the acquisition units, the fundamental matrix is obtained; The center point pixel coordinates of the specified target in the first-view detection result are used as input, and matrix multiplication is performed using the fundamental matrix to obtain the epipolar equation of the center point of the specified target in the first-view detection result. Traverse the detection results from the first-person perspective to obtain the set of center point pixel coordinates of all candidate targets in the same category group, and obtain the set of center point pixel coordinates of the candidate targets; For each candidate target's center point in the candidate target's center point pixel coordinate set, the vertical pixel distance from the corresponding epipolar line is obtained; If the specified vertical pixel distance is less than the preset tolerance threshold, then the candidate target in the second viewpoint is determined to form a candidate homogeneous point pair with the target in the first viewpoint, and a homogeneous target point pair is obtained.
[0224] In one embodiment of this application, the candidate three-dimensional spatial coordinates of the target are obtained by calculating the points of the same source target using a triangulation algorithm, including: Construct a system of linear equations containing the first-view projection matrix and the second-view projection matrix; Substitute the pixel coordinates of the target point pairs that satisfy the same source into the linear equation system, and solve the linear equation system using the least squares method or singular value decomposition method to obtain the candidate three-dimensional spatial coordinates of the target.
[0225] In one embodiment of this application, the detection result includes the target's position coordinates. Reprojection error verification is performed based on the target's candidate three-dimensional spatial coordinates to determine the three-dimensional detection result, which includes: The candidate 3D spatial coordinates of the target are projected back onto their respective image planes through the first-view projection matrix and the second-view projection matrix to obtain the reprojected pixel coordinates. The reprojection error is obtained by calculating the Euclidean distance between the reprojected pixel coordinates and the target position coordinates. If the reprojection error is less than the preset pixel threshold, the candidate 3D spatial coordinates of the corresponding target are determined to be true, and the 3D detection result is obtained. If the reprojection error is less than a preset pixel threshold, the candidate 3D spatial coordinates of the corresponding target are determined to be false, and the candidate 3D spatial coordinates of the corresponding target are removed.
[0226] In one embodiment of this application, the identification result includes a first identification result and a second identification result. The three-dimensional detection result is compared with preset standard process parameters to obtain the identification result, which includes: The spatial distance deviation between the three-dimensional detection results and the standard design coordinates is calculated to obtain the target spatial distance deviation; The first identification result is obtained when the target spatial distance deviation is within the allowable tolerance range; When the target spatial distance deviation is outside the allowable tolerance range, a second identification result is obtained.
[0227] The functions of each module in each device in the embodiments of this application can be found in the corresponding descriptions in the above methods, and will not be repeated here.
[0228] Figure 9 A structural block diagram of an electronic device according to an embodiment of this application is shown. Figure 9As shown, the electronic device includes a memory 410 and a processor 420. The memory 410 stores instructions that can be executed on the processor 420. When the processor 420 executes the instructions, it implements the method for identifying automotive welding studs in the above embodiments. The number of memories 410 and processors 420 can be one or more. This electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0229] The electronic device may also include a communication interface 430 for communicating with external devices and exchanging data. The devices are interconnected using different buses and can be mounted on a common motherboard or otherwise as needed. The processor 420 can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). The bus can be divided into address buses, data buses, control buses, etc. For ease of illustration, Figure 9 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0230] Optionally, in a specific implementation, if the memory 410, processor 420 and communication interface 430 are integrated on a single chip, the memory 410, processor 420 and communication interface 430 can communicate with each other through an internal interface.
[0231] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting Advanced Reduced Instruction Set Machines (ARM) architecture.
[0232] This application provides a computer-readable storage medium (such as the memory 410 described above) that stores computer instructions, which, when executed by a processor, implement the method provided in this application.
[0233] Optionally, memory 410 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device, etc. Furthermore, memory 410 may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory 410 may optionally include memory remotely located relative to processor 420, and these remote memories can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0234] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.
[0235] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.
[0236] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in this application, and these should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for identifying automotive welding studs, characterized in that, include: Acquire first-view image data of automotive parts from multiple perspectives; Based on the specified model, the first image data is used to detect weld point targets and stud targets, and detection results are generated from each viewpoint. Based on the principle of epipolar geometry constraint, the detection results from each perspective are filtered to obtain pairs of target points from the same source. The candidate three-dimensional spatial coordinates of the target are obtained by calculating the points of the same source using a triangulation algorithm. Based on the candidate 3D spatial coordinates of the target, the reprojection error is verified to determine the 3D detection result; The three-dimensional detection results are compared with preset standard process parameters to obtain the identification results.
2. The method according to claim 1, characterized in that, The acquisition of first image data of automotive parts from multiple perspectives includes: Multiple image acquisition units are used to capture images of automotive parts from different preset angles, obtaining multiple raw images; Geometric distortion correction is performed on the multiple original images based on pre-calibrated camera intrinsic parameters and distortion coefficients to obtain a multi-view corrected image sequence. The corrected image sequence is divided into multiple non-overlapping rectangular sub-blocks, and a histogram of each sub-block is obtained. The histogram is cropped and restricted, and the number of pixels exceeding a preset threshold is evenly distributed to other gray levels of the histogram to obtain the processed sub-blocks; The processed sub-blocks are then subjected to bilinear interpolation fusion to obtain first image data of automotive parts from multiple perspectives.
3. The method according to claim 2, characterized in that, The method for constructing the specified model includes: Obtain the initial model and training set; A channel attention module and a spatial attention module are embedded in the backbone feature extraction network of the initial model to obtain the first model. The channel attention module is used to weight each channel of the feature map, and the spatial attention module is used to weight the spatial position of the feature map. In the feature fusion network of the first model, a shallow high-resolution feature branch for detecting small targets is set to obtain the second model. The shallow high-resolution feature branch integrates deep semantic information and shallow texture information. The shallow high-resolution feature branch is used to detect solder joints or studs with pixel sizes smaller than a preset threshold. Based on the second model and the training set, the predicted bounding box and the predicted class probability are obtained; Based on the training set, determine the ground truth bounding boxes; The second model is trained using a multi-task joint loss function, which includes a localization loss component. The intersection-union ratio (IUU) is obtained by calculating the ratio of the overlap area between the predicted bounding box and the true bounding box to the union area. Calculate the Euclidean distance between the center point of the predicted bounding box and the center point of the true bounding box, and obtain the diagonal distance of the smallest closure region that can simultaneously contain both bounding boxes. Divide the Euclidean distance by the diagonal distance for normalization to obtain the center point distance penalty term. Calculate the difference between the aspect ratio of the predicted bounding box and the aspect ratio of the true bounding box to obtain the aspect ratio penalty term; The localization loss value of the sample is obtained by subtracting the center point distance penalty term and the aspect ratio penalty term from the intersection-union ratio. A dynamic modulation factor is constructed based on the predicted probability of the category, and the dynamic modulation factor is multiplied by the standard binary cross-entropy loss function term to obtain the classification loss part. The second model is optimized based on the localization loss value, the classification loss component, and the training set to obtain the specified model.
4. The method according to claim 3, characterized in that, The detection results from each perspective include the detection results from the first perspective and the detection results from the second perspective. Based on the principle of epipolar geometry constraints, the detection results from each perspective are filtered to obtain pairs of target points from the same source, including: Based on the relative positions of the acquisition units, the fundamental matrix is obtained; Using the center point pixel coordinates of the specified target in the detection result of the first viewpoint as input, matrix multiplication is performed using the basic matrix to obtain the epipolar equation of the center point of the specified target in the detection result of the first viewpoint. The center pixel coordinates of all candidate targets in the same category group in the detection results of the first viewpoint are traversed to obtain the set of center pixel coordinates of the candidate targets; For each candidate target's center point in the candidate target's center point pixel coordinate set, the vertical pixel distance from the corresponding epipolar line is obtained; If the specified vertical pixel distance is less than a preset tolerance threshold, then the candidate target in the second viewpoint and the target in the first viewpoint are determined to form a candidate homogeneous point pair, thus obtaining a homogeneous target point pair.
5. The method according to claim 4, characterized in that, The calculation of the candidate three-dimensional spatial coordinates of the target by using the triangulation algorithm on the same target point pair includes: Construct a system of linear equations containing the first-view projection matrix and the second-view projection matrix; Substitute the pixel coordinates of the corresponding target point pairs into the linear equations, and solve the linear equations using the least squares method or singular value decomposition method to obtain the candidate three-dimensional spatial coordinates of the target.
6. The method according to claim 5, characterized in that, The detection result includes the target's position coordinates. The reprojection error verification based on the target's candidate 3D spatial coordinates, and the determination of the 3D detection result, include: The candidate 3D spatial coordinates of the target are projected back onto their respective image planes through the first view projection matrix and the second view projection matrix to obtain the reprojected pixel coordinates. The reprojection error is obtained by calculating the Euclidean distance between the reprojected pixel coordinates and the target position coordinates; If the reprojection error is less than a preset pixel threshold, the candidate 3D spatial coordinates of the corresponding target are determined to be true, and a 3D detection result is obtained. If the reprojection error is greater than or equal to a preset pixel threshold, the candidate three-dimensional spatial coordinates of the corresponding target are determined to be false, and the candidate three-dimensional spatial coordinates of the corresponding target are removed.
7. The method according to claim 6, characterized in that, The identification results include a first identification result and a second identification result. The step of comparing the three-dimensional detection results with preset standard process parameters to obtain the identification results includes: The spatial distance deviation between the three-dimensional detection results and the standard design coordinates is calculated to obtain the target spatial distance deviation; If the target spatial distance deviation is within the allowable tolerance range, a first identification result is obtained; If the target spatial distance deviation is not within the allowable tolerance range, a second identification result is obtained.
8. A system for identifying automotive welding studs, characterized in that, include: The first acquisition module acquires first image data of automotive parts from multiple perspectives. The first generation module performs weld point target and stud target detection on the first image data based on a specified model, and generates detection results from various perspectives. The detection results include the target's position coordinates, target category, and confidence information. The first module obtains pairs of target points from the same source by filtering the detection results from each viewpoint based on the principle of epipolar geometry constraint. The second module obtains the candidate three-dimensional spatial coordinates of the target by using a triangulation algorithm to calculate the points of the same source target; The first determining module is used to verify the reprojection error based on the candidate three-dimensional spatial coordinates of the target and determine the three-dimensional detection result. The third module is used to compare the three-dimensional detection results with preset standard process parameters to obtain the identification results.
9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.
10. A computer-readable storage medium storing computer instructions that, when executed by a processor, implement the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Weld defect identification and positioning method and system based on deep learning network
CN113034478A
Target detection method based on improved YOLOv5 and binocular stereo vision
CN114565900A
Welding spot positioning device based on multi-view vision and detection method
CN119394179A
Method, processor and system for detecting stud welding quality
CN119579559A
Welding spot quality detection method and system
CN120182242A