Screw detection method, device, product and medium in terminal equipment production process

By dividing the detection units based on structural parameters and historical data during the production process of terminal equipment, acquiring images from multiple angles and performing enhancement processing and cross-view alignment, and utilizing a multi-task detection model, the problem of incomplete screw detection under a single viewpoint is solved, thereby improving the accuracy of screw detection.

CN122492675APending Publication Date: 2026-07-31GUANGZHOU XIYANG ELECTRONIC CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU XIYANG ELECTRONIC CO LTD
Filing Date
2026-06-08
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In the current production process of terminal equipment, single-angle machine vision inspection solutions are difficult to fully cover all screw positions, resulting in omissions or misjudgments and low accuracy.

Method used

Based on the structural parameter information and historical detection data of the terminal device, the spatial distribution area of ​​the screw is determined and the detection unit is divided. Images are acquired from multiple preset angles, and the screw detection information is fused and analyzed using a multi-task detection model through image enhancement processing and cross-view alignment.

Benefits of technology

It achieves precise positioning of screws and accurate acquisition of status information, improving the accuracy of inspection, and continuously optimizes the inspection system through the feedback mechanism of inspection reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492675A_ABST
    Figure CN122492675A_ABST
Patent Text Reader

Abstract

This invention relates to a screw detection method, equipment, product, and medium in the production process of terminal equipment, and pertains to the field of computer vision technology. The method determines the spatial distribution area of ​​the screw to be detected and divides it into detection units based on the structural parameters of the terminal equipment and historical detection data. By acquiring images of each detection unit from multiple preset angles and performing image enhancement processing based on the spatial position information of the detection units, clear image information of the screw can be obtained from different perspectives. Furthermore, by performing cross-viewpoint alignment processing on the target images from different angles to obtain alignment feature maps, and inputting these maps into a pre-trained multi-task detection model, effective fusion and comprehensive analysis of multi-viewpoint information is achieved, thereby accurately obtaining the screw's position information, state information, and three-dimensional offset. Finally, by generating a screw detection report and updating historical detection data, the accuracy of screw detection in the production process of terminal equipment is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, specifically to a screw detection method, equipment, product, and medium in the production process of terminal equipment. Background Technology

[0002] With the rapid development of consumer electronics, the structural design of terminal devices is becoming increasingly complex, and the number of screws used is also constantly increasing. As an important fastener in terminal devices, the installation quality of screws directly affects the reliability and safety of the product. Therefore, the importance of screw inspection in the production process of terminal devices is becoming increasingly prominent.

[0003] Currently, screw inspection on terminal equipment production lines primarily employs a single-angle machine vision inspection solution. This solution involves deploying fixed image acquisition devices at the inspection station to photograph the terminal equipment and analyze the screw installation status. However, due to the complexity of the terminal equipment structure, a single-angle inspection solution often struggles to comprehensively cover all screw locations. This is especially problematic in areas with dense screw distribution, leading to missed or incorrect detections and consequently, low screw inspection accuracy on the terminal equipment. Summary of the Invention

[0004] In view of this, this application provides a screw inspection method, equipment, product, and medium in the production process of terminal equipment.

[0005] In a first aspect, this application provides a screw inspection method during the production process of terminal equipment, the method comprising: Obtain structural parameter information and historical test data of the terminal device, determine the spatial distribution area of ​​each screw to be tested in the terminal device based on the structural parameter information and historical test data, and divide the spatial distribution area into multiple detection units, with each detection unit corresponding to at least one screw to be tested; When the terminal device arrives at the detection station, it captures images of each detection unit from multiple preset angles to obtain images of each detection unit at different angles. Based on the spatial position information of each detection unit, the captured image is enhanced to generate target images of each detection unit at different angles; Cross-view alignment processing is performed on the target images of each detection unit at different angles to obtain the alignment feature maps of each detection unit at different angles. The alignment feature maps of each detection unit at different angles are input into the pre-trained multi-task detection model to obtain the screw detection information corresponding to each detection unit. The screw detection information includes position information, state information and three-dimensional offset. Based on the screw detection information, a screw detection report is generated for the terminal device, and the historical detection data is updated.

[0006] By adopting the above technical solution, the spatial distribution area of ​​the screws to be inspected is determined and detection units are divided based on the structural parameter information of the terminal equipment and historical inspection data, enabling precise positioning and reasonable grouping of screw locations. By acquiring images of each detection unit from multiple preset angles and performing image enhancement processing based on the spatial position information of the detection units, clear image information of the screws from different perspectives can be obtained. Furthermore, by performing cross-viewpoint alignment processing on the target images from different angles to obtain alignment feature maps, and inputting them into a pre-trained multi-task detection model, effective fusion and comprehensive analysis of multi-viewpoint information is achieved, thereby accurately obtaining the screw's position information, state information, and three-dimensional offset. Finally, by generating screw inspection reports and updating historical inspection data, a feedback mechanism for the inspection results is established, enabling the inspection system to be continuously optimized and improved. This multi-viewpoint collaborative inspection scheme effectively solves the problem of difficulty in comprehensively covering all screw locations from a single perspective, significantly improving the accuracy of screw inspection during the production process of terminal equipment.

[0007] Optionally, the step of determining the spatial distribution area of ​​each screw to be inspected in the terminal device based on the structural parameter information and historical detection data, and dividing the spatial distribution area into multiple detection units, includes: Obtain the three-dimensional structural diagram of the terminal device and the screw specification parameters from the structural parameter information; Based on the screw anomaly records in the historical inspection data, the frequency of anomalies for each screw specification is determined. Based on the three-dimensional structural diagram and the frequency of anomalies, a heat map of screw density distribution is generated, and the area where the heat value exceeds a preset threshold is determined as the spatial distribution area of ​​each screw to be detected in the terminal device. Calculate the spatial distance and detection difficulty between adjacent screws to be detected within the spatial distribution area; Screws whose spatial distance is less than a first preset threshold and whose difference in detection difficulty is less than a second preset threshold are grouped into the same detection unit.

[0008] Optionally, the spatial location information includes the three-dimensional coordinates and surface normal vectors of the detection units. The step of performing image enhancement processing on the captured images based on the spatial location information of each detection unit to generate target images of each detection unit at different angles includes: The relative distance between each detection unit and the shooting position corresponding to each preset angle is calculated based on the three-dimensional coordinates to obtain the distance attenuation coefficient; The reflectivity is calculated based on the angle between the surface normal vector of each detection unit and the shooting direction corresponding to each preset angle. Images with reflectivity exceeding the preset threshold are marked as reflective images. Brightness saturation detection is performed on the reflected image to identify overexposed pixels; The overexposed pixels are compensated for brightness using the distance attenuation coefficient, and the compensated image data is then subjected to high-pass filtering to obtain the target images of each detection unit at different angles.

[0009] Optionally, cross-viewpoint alignment processing is performed on the target images of each detection unit at different angles to obtain alignment feature maps of each detection unit at different angles, including: The SIFT algorithm is used to extract key points and corresponding feature descriptors of the target image at different angles in each detection unit. Based on the feature descriptor, feature point matching is performed between target images from different angles to obtain matching point pairs; The RANSAC algorithm is used to filter the matching point pairs, and the homography matrix between target images at different angles is calculated based on the filtered matching point pairs. Based on the homography matrix, perspective transformation is performed on target images at different angles to align target images of the same detection unit at different angles to a unified reference viewpoint. Multi-scale feature extraction is performed on the aligned images at different angles to obtain the alignment feature maps of each detection unit at different angles.

[0010] Optionally, the step of performing multi-scale feature extraction on the aligned images at different angles to obtain alignment feature maps of each detection unit at different angles includes: The aligned images at different angles are downsampled according to different spatial resolutions; Feature extraction is performed on the downsampled image using convolution kernels of different sizes to obtain coarse-grained features and fine-grained features. The coarse-grained features are used to capture the overall outline and position information of the screw, while the fine-grained features are used to capture the texture details and edge information of the screw surface. Channel splicing is performed on the coarse-grained and fine-grained features at each angle to obtain the preliminary feature representation of each angle; Based on the angle between the shooting direction at each angle and the surface normal vector of the detection unit, the view quality weight of each angle is calculated, and the view quality weight is expanded into a weight map with the same spatial size as the preliminary feature representation, and then weighted element-wise with the preliminary feature representation of the corresponding angle. The extrinsic parameters of the camera at each angle are converted into position encoding vectors, and then concatenated with the corresponding angle-weighted feature representations along the channel dimension to obtain the aligned feature maps of each detection unit at different angles.

[0011] Optionally, the method further includes a training process for a multi-task detection model, specifically including: Collect multi-view image data from multiple terminal devices of different models, including images of each detection unit at different angles; The multi-view image data is labeled, and the labeling content includes the bounding box coordinates of the screw, the screw type, the status label and the three-dimensional position truth value. The status label includes at least one of missing, present, loose, tight, tilted and vertical. Cross-view alignment and multi-scale feature extraction are performed on the labeled multi-view image data to obtain the alignment feature map for training. A multi-task detection model is constructed, which includes a shared backbone network and multiple task branches, including a localization branch, a state classification branch, and a 3D correction branch. The alignment feature map used for training is input into the multi-task detection model. The multi-task loss function is calculated based on the difference between the model output and the corresponding labeled content. The multi-task loss function includes localization loss, classification loss and 3D offset loss. Based on the multi-task loss function, the model parameters are updated by the optimizer, and the training is repeated until the loss function converges to obtain a pre-trained multi-task detection model.

[0012] Optionally, the method further includes: Based on the screw detection information of each detection unit, calculate the consistency rate of state judgment and the variance of three-dimensional offset; Identify low-confidence screws whose state determination consistency rate is lower than the state threshold or whose three-dimensional offset variance exceeds the preset offset range; The alignment feature maps of the low-confidence screws at various angles are extracted, the sharpness score and occlusion degree of the feature maps are calculated, and the optimal detection angle is determined based on the sharpness score and occlusion degree. The location of the screws with low confidence is marked in the screw inspection report, and an image taken at the optimal inspection angle is attached as a basis for manual re-inspection. When there is an abnormal detection unit in the historical detection data where the number of times a low-confidence screw appears is greater than a preset number, the number of shooting angles of the abnormal detection unit will be increased during the next detection by the terminal device.

[0013] A second aspect of this application provides an electronic device for screw detection in the production process of a terminal device. The electronic device includes: one or more processors and a memory; the memory is coupled to the one or more processors and is used to store computer program code, the computer program code including computer instructions, wherein the one or more processors invoke the computer instructions to cause the electronic device for screw detection in the production process of the terminal device to perform the method described in the first aspect and any possible implementation thereof.

[0014] A third aspect of this application provides a computer program product containing instructions that, when run on an electronic device for screw detection during the manufacturing process of a terminal device, causes the electronic device to perform the method described in the first aspect and any possible implementation thereof.

[0015] A fourth aspect of this application provides a computer-readable storage medium including instructions that, when executed on an electronic device for screw detection during the production process of an end device, cause the electronic device to perform the method described in the first aspect and any possible implementation thereof.

[0016] In summary, one or more technical solutions provided in this application have at least the following technical effects or advantages: This application determines the spatial distribution area of ​​the screws to be inspected and divides it into inspection units based on the structural parameter information of the terminal equipment and historical inspection data, enabling precise positioning and reasonable grouping of screw locations. By acquiring images of each inspection unit from multiple preset angles and performing image enhancement processing based on the spatial position information of the inspection units, clear image information of the screws from different perspectives can be obtained. Furthermore, by performing cross-viewpoint alignment processing on the target images from different angles to obtain alignment feature maps, and inputting them into a pre-trained multi-task detection model, effective fusion and comprehensive analysis of multi-viewpoint information is achieved, thereby accurately obtaining the screw's position information, state information, and three-dimensional offset. Finally, by generating a screw inspection report and updating historical inspection data, a feedback mechanism for the inspection results is established, enabling the inspection system to be continuously optimized and improved. This multi-viewpoint collaborative inspection scheme effectively solves the problem of difficulty in comprehensively covering all screw locations from a single perspective, significantly improving the accuracy of screw inspection during the production process of terminal equipment. Attached Figure Description

[0017] Figure 1 This is a schematic flowchart of a screw inspection method in the production process of a terminal device provided in an embodiment of this application; Figure 2 This is a schematic diagram of a detection unit division provided in an embodiment of this application; Figure 3This is a schematic diagram of an exemplary hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0018] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0019] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0020] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0021] Please refer to Figure 1 This paper presents a flowchart illustrating a screw inspection method in the production process of terminal equipment. This method can be implemented using a computer program, a microcontroller, or run within a screw inspection device used in the production process of terminal equipment. The computer program can be integrated into the computer device or run as a standalone application. Specifically, the method includes steps 10 to 60, as follows: Step 10: Obtain the structural parameter information and historical test data of the terminal device. Based on the structural parameter information and historical test data, determine the spatial distribution area of ​​each screw to be tested in the terminal device, and divide the spatial distribution area into multiple test units. Each test unit corresponds to at least one screw to be tested.

[0022] It should be noted that this application is mainly applicable to the production and quality inspection process of high-precision terminal equipment, such as server production lines with high product reliability requirements, medical equipment assembly requiring shock resistance and loosening prevention, communication equipment manufacturing with strict safety level requirements, or other terminal equipment production with strict requirements on screw connection quality.

[0023] Specifically, the first step is to acquire the structural parameter information of the terminal device, including a 3D structural diagram and screw specifications. The 3D structural diagram contains the spatial coordinates and installation direction of each screw in the terminal device, while the screw specifications include the dimensions and installation requirements for different types of screws. Simultaneously, historical inspection data is acquired. This data contains accumulated screw anomaly records for this type of terminal device during the production process, recording the frequency statistics of missing, loose, or tilted screws in different locations and of different specifications.

[0024] After obtaining the above information, a screw density distribution heatmap is first generated based on the screw spatial distribution information in the 3D structural diagram of the terminal device, combined with the frequency of anomalies of screws of various specifications in historical inspection data. This heatmap uses color depth to reflect the screw distribution density and anomaly risk level in different areas, with darker colors indicating denser screw distribution or higher anomaly frequency. Areas in the heatmap where the heat value exceeds a preset threshold are identified as the spatial distribution areas of the screws to be inspected. These areas are typically key areas of focus where screws are densely installed or prone to anomalies.

[0025] After determining the spatial distribution area, it needs to be further divided into multiple detection units. The spatial distance between adjacent screws to be detected within the spatial distribution area can be calculated, and the detection difficulty of each screw can be evaluated. The detection difficulty mainly considers factors such as screw visibility, background complexity, and lighting conditions. When the spatial distance between adjacent screws is less than a first preset threshold (e.g., 10 mm) and the difference in detection difficulty is less than a second preset threshold (e.g., the difference in detection difficulty scores does not exceed 20%), these screws are grouped into the same detection unit. Each detection unit contains at least one screw to be detected; screws that are relatively independent or have significantly different detection conditions are treated as separate detection units. This detection unit division method based on structural parameters and historical data can fully utilize the structural features and historical experience of the terminal device, organizing screws with similar spatial locations and detection conditions into detection units, which is beneficial for targeted optimization of subsequent image acquisition and processing.

[0026] As an optional embodiment, the step of determining the spatial distribution area of ​​each screw to be inspected in the terminal device based on structural parameter information and historical inspection data, and dividing the spatial distribution area into multiple inspection units, may further include the following steps: Step 101: Obtain the three-dimensional structural diagram of the terminal device and the screw specifications from the structural parameter information.

[0027] Specifically, the system acquires the 3D structural diagram of the terminal device and screw specifications through a pre-defined data interface. The 3D structural diagram is stored in a standard CAD format and contains complete 3D model data of the terminal device. This data details precise spatial information such as the 3D coordinates of each screw's installation position, the normal vector of the mounting surface, and the depth and diameter of the screw hole. The screw specifications include specific information on various screw types, such as screw type (e.g., M2.5, M3), thread specification, head shape (e.g., Phillips head, hexagonal head), material properties, and standard torque value. This structural parameter information provides basic spatial positioning and specification reference data for subsequent detection unit division.

[0028] Step 102: Determine the frequency of abnormal occurrences for each screw specification based on the screw abnormality records in the historical inspection data.

[0029] Specifically, the system extracts historical inspection data from the production database, which includes screw inspection records for each terminal device. For each anomaly record, the system records detailed information such as the anomaly type (missing, loose, tilted), location, and screw specifications. Through data analysis algorithms, the system calculates the anomaly frequency for each screw specification at different locations. The specific calculation method is as follows: within a specified time window (e.g., the past 3 months), the number of times anomalies occur at each screw location is counted and divided by the total number of inspections to obtain the anomaly frequency. Simultaneously, the system also considers the temporal distribution characteristics of anomalies, assigning higher weight to more recent anomalies, making the frequency calculation more reflective of the current production status.

[0030] Step 103: Based on the three-dimensional structure diagram and the frequency of anomalies, generate a heat map of screw density distribution, and determine the area where the heat value exceeds the preset threshold as the spatial distribution area of ​​each screw to be tested in the terminal device.

[0031] Specifically, based on the acquired 3D structural map and anomaly frequency data, the system constructs a heat map of screw density distribution. First, the 3D structural map is projected onto a 2D plane to create a gridded density matrix. For each grid cell, the number of screws within its coverage area and the corresponding anomaly frequency are calculated, and these two factors are combined to calculate a heat value. The formula for calculating the heat value is: Heat Value = Screw Density Coefficient × (1 + Anomaly Frequency Weight), where the screw density coefficient reflects the density of screw distribution per unit area, and the anomaly frequency weight is determined based on the historical anomaly frequency of screws in that area. The system visualizes the heat value using different colors and sets a preset threshold (e.g., heat value > 0.8) to determine key spatial distribution areas of interest.

[0032] Step 104: Calculate the spatial distance between adjacent screws to be inspected within the spatial distribution area and the difficulty of inspection.

[0033] Specifically, within a defined spatial distribution area, the system needs to evaluate the spatial relationship and detection conditions between adjacent screws. The spatial distance is calculated using the three-dimensional Euclidean distance formula, that is, for any two screw locations P1(x1,y1,z1) and P2(x2,y2,z2), the spatial distance is calculated as follows: This distance value reflects the actual physical spacing between the screws.

[0034] The assessment of the difficulty of detection mainly considers the following key factors: 1) Visibility factor (V): Calculated based on the angle θ between the screw position and the optical axis of the detection camera. When θ≤30°, V=1; when 30°<θ≤60°, V=cos(θ); when θ>60°, V=0.2. The imaging quality of the camera on the target is closely related to the viewing angle, and the vertical viewing angle provides the best observation effect. 2) Occlusion factor (K): Calculated by analyzing the 3D structure around the screw. First, a spherical space (e.g., radius 20mm) is established with the screw center as the origin. The volume ratio r occupied by other structural components in this space is calculated, K= 1 - r / 0.8 (when r>0.8, K=0). The occlusion of surrounding structural components will affect the image acquisition quality. 3) Background Complexity Factor (C): Through local image region analysis, the gray-level gradient variance σ is calculated in the 30×30 pixel region around the screw, C = 1 - σ / σmax, where σmax is an empirical value, usually the maximum variance from historical data. Background texture complexity affects the accuracy of target detection. The final detection difficulty score S is calculated using a weighted method: S = 100 × (w1×V+w2×K+w3×C), where the weight coefficients w1=0.5, w2=0.3, and w3=0.2.

[0035] Step 105: Combine the screws to be tested that are less than the first preset threshold in terms of spatial distance and less than the second preset threshold in terms of difficulty of detection into the same detection unit.

[0036] Specifically, the system divides detection units based on calculated spatial distance and detection difficulty scores. First, two thresholds are set: a spatial distance threshold (e.g., 10mm) and a detection difficulty difference threshold (e.g., 20 points). For any two adjacent screws, if the spatial distance between them is less than the first preset threshold and the difference in detection difficulty scores is less than the second preset threshold, then these two screws are grouped into the same detection unit. The system uses a clustering algorithm for automatic division. Starting with any screw, it sequentially checks whether surrounding screws meet the merging conditions until no new screws meeting the conditions can be found, thus forming a detection unit. This division method ensures that screws within the same detection unit have similar spatial positions and detection conditions, which is beneficial for subsequent image acquisition and targeted optimization of detection algorithms.

[0037] Please see Figure 2 This is a schematic diagram of a detection unit division provided in an embodiment of this application. The left figure shows a thermal map of the spatial distribution density of screws on the surface of the terminal device, where the red area represents the area with the densest screw distribution, gradually transitioning to the periphery through yellow and green, indicating a trend of gradually decreasing screw density. The right figure shows the detection unit division result based on density distribution characteristics, with detection unit 1, detection unit 2, and detection unit 3, etc., defined by blue dashed circles.

[0038] Step 20: When the terminal device arrives at the inspection station, it acquires images of each inspection unit from multiple preset angles to obtain images of each inspection unit at different angles.

[0039] Specifically, after the terminal equipment arrives at the inspection station via the conveyor, the system determines the spatial position of each inspection unit based on pre-stored inspection unit information. The inspection station is equipped with multiple camera mounting positions, including but not limited to front-view cameras C1 and C2, oblique-view cameras C3 and C4, and top / bottom-view cameras C5 and C6. The front-view cameras (C1 / C2) are primarily used to capture screw head features, including Phillips head, hexagonal head, and brand markings; the oblique-view cameras (C3 / C4) are used to observe the fit between the screw side and the threaded hole, particularly to detect any tilting; and the top / bottom-view cameras (C5 / C6) are primarily used to check whether the screw is fully embedded in the threaded hole. In addition, the inspection station is equipped with a composite light source system consisting of a ring light source, a coaxial light source, and a low-angle light source to address the issue of metal surface reflection and enhance the contrast between the screw and the background.

[0040] The system employs an industrial-grade high-resolution camera for image acquisition, with a resolution of 4096×3072 pixels and an adjustable focal length range of 8–50 mm. When the detection program is triggered by the terminal device's arrival signal, the front-view cameras C1 and C2 first capture clear images of the screw head features. Next, the system controls the oblique-view cameras C3 and C4 to capture side images of the screw from different angles for subsequent analysis of the screw's installation tilt. Finally, the top / bottom-view cameras C5 and C6 acquire the screw's embedding depth information.

[0041] During image acquisition, the system selects the most suitable light source combination based on the specific characteristics of the detection unit. For capturing the screw head features, a coaxial light source is primarily used to provide uniform illumination; for detecting side tilt, a low-angle light source is used to highlight the screw outline; and a ring light source is used to eliminate reflections from the metal surface. The system also automatically adjusts the focal length and aperture parameters of each camera based on the spatial position information of the detection unit to ensure clear images are acquired at different shooting angles.

[0042] For each detection unit, the system acquires multiple images at each preset angle, resulting in images captured by each detection unit at different angles. All acquired images are stored according to the detection unit number and the shooting angle. For example, "Unit1_C1_20251024143000.jpg" represents the image of detection unit 1 acquired by camera C1 at 14:30:00 on October 24, 2025.

[0043] Step 30: Perform image enhancement processing on the captured images based on the spatial position information of each detection unit to generate target images of each detection unit at different angles.

[0044] Specifically, when performing image enhancement processing on images captured by each detection unit at different angles, the system first performs image preprocessing based on the spatial location information of the detection units. Since different detection units are located at different planes and depths on the terminal device, the captured images may suffer from distortion, blurring, or uneven lighting. Based on the three-dimensional coordinate information of the detection units and the camera parameters, the system uses a perspective transformation algorithm to perform geometric correction on the captured images, eliminating image distortion caused by the shooting angle.

[0045] For images captured by front-view cameras C1 and C2, the system focuses on enhancing the clarity of screw head features. Since features such as cross-shaped and hexagonal screw heads have low contrast, the system employs a local contrast enhancement algorithm. By calculating the mean and variance of grayscale values ​​in local image regions, the system dynamically adjusts the mapping relationship of pixel values ​​to highlight the structural features of the screw head. For images captured by oblique-view cameras C3 and C4, the system primarily enhances the edge features of the screw outline. A directional edge enhancement algorithm is used to selectively enhance edge information in specific directions based on the screw's spatial installation orientation, making the screw's tilt more clearly visible. For images captured by top / bottom-view cameras C5 and C6, the system focuses on improving the salience of the screw embedding depth. Through illumination compensation and shadow enhancement techniques, the system highlights the depth hierarchy between the screw and the screw hole.

[0046] The system also specifically addresses the issue of metal surface reflection. By analyzing the bright areas in the image, a detection model for reflective areas is established, and adaptive compensation is applied to the detected reflective areas. The compensation process considers the grayscale distribution characteristics of surrounding pixels and employs gradient consistency constraints to ensure a natural transition in the compensated image. Simultaneously, the system performs targeted image enhancement based on the combined characteristics of different light sources. For example, for images illuminated by a ring light source, the focus is on suppressing interference caused by ring reflections; for images illuminated by a low-angle light source, the emphasis is on enhancing the projection features of the screw outline. After these enhancement processes, the system saves the results as the target image. Through this image enhancement processing based on spatial location information, the system can overcome imaging differences caused by different detection positions and obtain standardized, high-quality detection images.

[0047] As another optional embodiment, the step of performing image enhancement processing on the captured image based on the spatial position information of each detection unit to generate target images of each detection unit at different angles may further include the following steps: Step 201: Calculate the relative distance between each detection unit and the shooting position corresponding to each preset angle based on three-dimensional coordinates to obtain the distance attenuation coefficient.

[0048] Specifically, the system calculates the relative distance based on the spatial coordinates (x1, y1, z1) of the detection unit stored in the 3D structure diagram and the camera installation position coordinates (x2, y2, z2) corresponding to each preset angle, using the 3D Euclidean distance formula: Based on the calculated distance value, the system uses an exponential decay model to calculate the distance attenuation coefficient α = exp(-d / d0), where d0 is the standard value of distance attenuation, typically taken as the nominal working distance of the detection station. This attenuation coefficient calculation method based on spatial distance takes into account the physical characteristic that light intensity decreases with increasing distance, providing an important reference for subsequent image enhancement.

[0049] Step 202: Calculate the reflection probability based on the angle between the surface normal vector of each detection unit and the shooting direction corresponding to each preset angle, and mark the captured images with a reflection probability exceeding the preset threshold as reflective images.

[0050] Specifically, the system first extracts the normal vector n(nx,ny,nz) of the detection unit surface from the 3D structure map, and simultaneously obtains the shooting direction vector v(vx,vy,vz) of each preset angle camera. The angle between the two vectors is calculated using their dot product. Based on the physical properties of reflection, reflection is most likely to occur when the angle between the incident and reflected light rays is close to the camera's viewing angle. Therefore, the system calculates the reflection probability P = exp(-|θ - θc| / σ), where θc is the critical reflection angle and σ is the probability decay parameter. When the calculated reflection probability P exceeds a preset threshold (e.g., 0.8), the system marks the captured image at that angle as a reflective image, preparing it for subsequent specialized processing.

[0051] Step 203: Perform brightness saturation detection on the reflective image to identify overexposed pixels.

[0052] Specifically, for the marked reflective images, the system performs brightness saturation detection. First, the image is converted to the HSV color space, and the V channel (brightness channel) data is extracted. The system sets a brightness saturation threshold Vth (typically 95% of the maximum brightness) and scans the image pixel by pixel. When the brightness value V(x,y) of a pixel is greater than the threshold Vth, that pixel is marked as overexposed. Simultaneously, the system analyzes the spatial distribution characteristics of overexposed pixels, calculating the area and shape characteristics of connected regions to distinguish between overexposure caused by reflection and highlight areas caused by other reasons.

[0053] Step 204: Use the distance attenuation coefficient to compensate the brightness of overexposed pixels, and perform high-pass filtering on the compensated image data to obtain the target images of each detection unit at different angles.

[0054] Specifically, the calculated distance attenuation coefficient α is used to compensate for the brightness of overexposed pixels. The compensation formula is: V'(x,y)=V(x,y)×α+β×(1-α), where V(x,y) is the original brightness value and β is the reference brightness value. This compensation method considers the influence of spatial distance on light intensity and can more accurately restore the actual brightness value. The compensated image is then subjected to high-pass filtering, using a Gaussian high-pass filter to enhance the image's detail features. The frequency response function of the filter is H(u,v)=1-exp[-(u²+v²) / 2D0²], where u and v represent the horizontal and vertical frequency components in the frequency domain, respectively, with values ​​ranging from [-M / 2, M / 2-1] and [-N / 2, N / 2-1], and M and N are the width and height of the image, respectively; D0 is the cutoff frequency, which determines the cutoff characteristics of the filter. A smaller D0 value will retain more high-frequency details, and the filtering intensity can be controlled by adjusting D0. The filtered image retains detail while avoiding information loss due to reflection, ultimately generating target images of each detection unit at different angles. These target images not only overcome interference from reflection but also improve overall image quality through distance compensation and high-pass filtering.

[0055] Step 40: Perform cross-view alignment processing on the target images of each detection unit at different angles to obtain the alignment feature maps of each detection unit at different angles.

[0056] Specifically, when performing cross-viewpoint alignment processing on target images of each detection unit at different angles, the system first establishes a coordinate mapping relationship between different viewpoints based on camera calibration parameters. For each detection unit, the system acquires its target images from six camera viewpoints, C1 to C6, and selects the image from the orthogonal viewpoint camera C1 as the reference viewpoint. Through the camera's intrinsic parameter matrix K and extrinsic parameter matrix [R|t], a projection relationship between the world coordinate system and the coordinate systems of each camera image is established, with the projection matrix P=K[R|t].

[0057] The system employs the SIFT algorithm to extract feature points from the baseline viewpoint image and generate feature descriptors. The feature points are then projected onto images from other viewpoints using a projection matrix. Nearest neighbor ratio testing and the RANSAC algorithm are used for feature matching to estimate the homography matrix H. Based on the homography matrix, the system performs perspective transformation and bilinear interpolation on the target images from different viewpoints, aligning them to the baseline viewpoint plane to obtain aligned feature maps of each detection unit at different angles. These aligned feature maps preserve the feature information of their respective viewpoints while establishing a pixel-level correspondence with the baseline viewpoint, providing accurate feature representations for subsequent state recognition.

[0058] As another optional embodiment, the step of performing cross-viewpoint alignment processing on the target images of each detection unit at different angles to obtain the alignment feature maps of each detection unit at different angles may further include the following steps: Step 301: Extract key points and corresponding feature descriptors of the target image at different angles in each detection unit using the SIFT algorithm.

[0059] Specifically, the system employs the SIFT algorithm to extract features from target images captured by each detection unit from the perspectives of six cameras (C1 to C6). First, a Gaussian difference pyramid is constructed, and the image is subjected to Gaussian blurring at different scales. The difference images are obtained by subtracting adjacent layers. Local extrema are detected in the difference images as candidate keypoints, and keypoints are optimized by removing low-contrast points and edge response points. For each keypoint, its principal orientation is calculated, and a 128-dimensional feature descriptor is generated based on the image gradient direction and magnitude statistics. The feature descriptor contains gradient histogram information in eight directions for a 4×4 sub-region of the keypoint's neighborhood, exhibiting rotation invariance and illumination invariance, effectively describing local image features.

[0060] Step 302: Based on the feature descriptor, perform feature point matching between target images from different angles to obtain matching point pairs.

[0061] Specifically, based on feature descriptors, the system performs feature matching on target images from different angles. A nearest neighbor ratio matching strategy is employed. For example, for feature point P1 in image 1, the system searches for the two feature points P2 and P2' in image 2 with the smallest Euclidean distance between their feature descriptors. If the ratio of the smallest distance to the second smallest distance is less than a threshold (usually 0.8), then P1 and P2 are considered a valid matching pair. This method effectively avoids fuzzy matching and improves matching accuracy. The system performs feature matching on each pair of camera views from C1 to C6, obtaining a complete set of matching point pairs.

[0062] Step 303: Use the RANSAC algorithm to filter the matching point pairs, and calculate the homography matrix between target images at different angles based on the filtered matching point pairs.

[0063] Specifically, the system uses the RANSAC algorithm to filter matching point pairs and eliminate incorrect matches. The process is as follows: four pairs of matching points are randomly selected, and the homography matrix H is calculated. All matching points are then transformed using H, and the reprojection error between the transformed points and their actual corresponding points is calculated. If the error is less than a threshold (usually 3 pixels), the matching point pair is marked as an inlier. This process is repeated multiple times (usually 1000 iterations), and the homography matrix with the most inliers is selected as the optimal result. The filtered matching point pairs accurately reflect the geometric transformation relationships between images from different viewpoints.

[0064] Step 304: Perform perspective transformation on the target images at different angles based on the homography matrix, and align the target images of the same detection unit at different angles to a unified reference viewpoint.

[0065] Specifically, based on the homography matrix H, the system performs perspective transformation on target images from different angles. The image at the positive angle C1 is selected as the reference angle, and images at other angles are transformed to this reference angle using the homography matrix. The transformation process employs a reverse mapping method; for each pixel (x', y') in the target image, its corresponding position (x, y) in the source image is calculated using the inverse matrix of H. Since the calculated coordinates are usually non-integer, the system uses bilinear interpolation to calculate the pixel values, ensuring the smoothness of the transformed image.

[0066] Step 305: Perform multi-scale feature extraction on the aligned images at different angles to obtain the alignment feature maps of each detection unit at different angles.

[0067] Specifically, the system performs multi-scale feature extraction on the aligned image. First, it constructs an image pyramid, then sequentially downsamples the original image by a factor of 2, generating three scale levels. At each scale level, HOG (Histogram of Oriented Gradients) and LBP (Local Binary Pattern) features are extracted. HOG features capture the image's edge and contour information, while LBP features describe local texture features. The system unfolds the features of different scales and types according to the original resolution, forming feature vectors, and finally generates an aligned feature map. This multi-scale feature extraction method can comprehensively describe the geometric and texture features of the screw, providing a reliable feature representation for subsequent state recognition.

[0068] As an optional embodiment, the step of extracting multi-scale features from the aligned images at different angles to obtain the alignment feature maps of each detection unit at different angles may further include the following steps: Step 401: Downsample the aligned images at different angles according to different spatial resolutions.

[0069] Specifically, the system performs multi-resolution downsampling on the images aligned across the six camera views (C1 to C6) for each detection unit. Using a Gaussian pyramid structure, the original image (1024×1024 pixels) is sequentially downsampled to lower resolution images of 512×512, 256×256, and 128×128 pixels. During downsampling, the image is first smoothed using a Gaussian filter with a kernel size of 5×5 and a standard deviation σ=1.6, and then pixels are sampled at a 2:1 ratio. This multi-resolution downsampling strategy can reduce computational load and provide feature representation at different scales while preserving the main structural information of the image.

[0070] Step 402: Use convolution kernels of different sizes to extract features from the downsampled image to obtain coarse-grained features and fine-grained features. Coarse-grained features are used to capture the overall outline and position information of the screw, while fine-grained features are used to capture the texture details and edge information of the screw surface.

[0071] Specifically, the system uses convolutional kernels of different sizes to extract features from the downsampled image. For coarse-grained feature extraction, larger convolutional kernels (7×7 and 5×5) with a stride of 2 are used, primarily extracting the overall outline and positional information of the screw. Coarse-grained features include: using a 7×7 convolutional kernel to extract 16 channels of shape features and using a 5×5 convolutional kernel to extract 32 channels of edge features. For fine-grained feature extraction, smaller convolutional kernels (3×3 and 1×1) with a stride of 1 are used, focusing on capturing the texture details and local edge information of the screw surface. Fine-grained features include: using a 3×3 convolutional kernel to extract 64 channels of texture features and using a 1×1 convolutional kernel to extract 128 channels of local response features. Each convolutional layer is followed by a BatchNormalization layer and a ReLU activation function to enhance the expressive power of the features.

[0072] Step 403: Perform channel splicing on the coarse-grained and fine-grained features at each angle to obtain the preliminary feature representation for each angle.

[0073] Specifically, the system concatenates the coarse-grained and fine-grained features extracted from each angle along the channel dimension. First, feature maps of different sizes are adjusted to the same spatial resolution (256×256) using bilinear interpolation. Then, the coarse-grained features from 48 channels (16+32) and the fine-grained features from 192 channels (64+128) are concatenated to obtain a preliminary feature representation of 240 channels. This multi-scale feature fusion preserves both the overall structural information of the screw and important local details.

[0074] Step 404: Calculate the view quality weight of each angle based on the angle between the shooting direction and the surface normal vector of the detection unit, and expand the view quality weight into a weight map with the same spatial size as the preliminary feature representation, and perform element-wise weighting with the preliminary feature representation of the corresponding angle.

[0075] Specifically, the system calculates the viewpoint quality weight based on the angle between each shooting angle and the surface normal vector of the detection unit. This is because the observation quality of screw features varies depending on the viewing angle: when the angle between the camera's shooting direction and the surface normal vector is small, the screw features are clearer; when the angle is close to 90 degrees, the feature information will be distorted or lost due to the tilt of the viewing angle. In practice, the surface normal vector of the detection unit and the shooting direction vector of each camera are first obtained from the 3D structural information. The angle θ is calculated by the vector dot product, and the viewpoint quality weight is calculated based on the angle: w = cos(θ)×exp(-|θ| / σ), where σ is the attenuation coefficient, usually taken as π / 4. This weight design ensures that the vertical viewing angle (angle close to 0 degrees) receives the maximum weight, while the weight of the lateral viewing angle (angle close to 90 degrees) approaches 0, which conforms to the relationship between viewpoint quality and angle in actual observation.

[0076] However, using only a single scalar weight cannot reflect the differences in feature reliability across different regions of the image. Therefore, the system employs a spatially adaptive weight expansion method, expanding the viewpoint quality weight w into a weight map with the same spatial size (256×256) as the feature representation. First, a 4×4 initial weight grid is generated. The weight value for each grid point (x,y) is calculated using distance weighting: w(x,y)=w×(1+α×exp(-d(x,y) / β)), where d(x,y) is the normalized distance from that point to the center of the feature map, α=0.2 is the center gain coefficient, and β=0.3 is the distance attenuation coefficient. This design considers that the central region of the image typically contains more reliable feature information and should therefore receive a slightly higher weight. Subsequently, a bicubic interpolation algorithm is used to expand the 4×4 weight grid into a 256×256 weight map Wi. The interpolation process employs a cubic polynomial kernel function to ensure a smooth transition in weight distribution, avoiding feature discontinuities caused by abrupt weight changes.

[0077] Finally, the system performs element-wise weighting of the weight map Wi with the preliminary feature representation Fi under the corresponding viewpoint: F'i(x,y,c)=Fi(x,y,c)×Wi(x,y), where x,y are the spatial coordinates and c is the feature channel index (c=1,2,...,240). This adaptive weighting mechanism based on viewpoint quality can effectively distinguish and utilize the reliability of features under different viewpoints. For images with good viewpoints (such as frontal views), their feature information is fully preserved; while for images with unfavorable viewpoints (such as strong side views), features that may be deformed or occluded are appropriately suppressed.

[0078] Step 405: Convert the extrinsic information of the camera at each angle into a position encoding vector, and concatenate it with the corresponding angle-weighted feature representation along the channel dimension to obtain the aligned feature map of each detection unit at different angles.

[0079] Specifically, in multi-camera systems, relying solely on image features is insufficient to fully express the geometric relationships between different viewpoints, leading to difficulties in accurately understanding the three-dimensional state of the screw. Therefore, it is necessary to encode the spatial position information of the cameras into the feature representation, enabling the network to learn the geometric constraints between viewpoints and improving the viewpoint invariance of the feature representation. The extrinsic parameters of each camera are converted into position encoding vectors, including the rotation matrix R and translation vector t. These vectors are flattened and mapped to 32-dimensional position encoding vectors through a fully connected layer. The position encoding vectors are then spatially expanded into a 256×256×32 feature map, which is concatenated with the aforementioned weighted feature representation (256×256×240) along the channel dimension, resulting in a 256×256×272 aligned feature map. This feature representation not only includes the visual features of the image but also incorporates the spatial position information of the cameras, providing a more complete feature description for subsequent state recognition.

[0080] As an optional embodiment, the screw detection method in the terminal device manufacturing process also includes a multi-task detection model training process, which may include the following steps: Step 501: Collect multi-view image data from multiple terminal devices of different models. The multi-view image data includes images of each detection unit at different angles.

[0081] Specifically, to construct a representative training dataset, the system collected multi-view image data from 1000 different models of terminal devices. Each device used six industrial cameras (C1 to C6) to simultaneously acquire images. The cameras were evenly distributed around the device at 60-degree intervals, with the shooting distance maintained within the range of 300mm to 500mm, resulting in a total of 6000 high-resolution images. During the acquisition process, various working conditions in a real production environment were simulated, including different lighting conditions and equipment placement angles, to ensure the diversity and authenticity of the data.

[0082] Step 502: Label the multi-view image data. The labeling content includes the bounding box coordinates of the screw, screw type, status label and 3D position ground truth. The status label includes at least one of missing, present, loose, tight, tilted and vertical.

[0083] Specifically, bounding boxes were created for each screw in each image, recording the center coordinates, width, and height of the bounding box in (x,y,w,h) format. Then, based on the screw's actual condition, its category (e.g., Phillips head screw, hexagonal screw, etc.) and status label were assigned. Status labels included: missing (screw not present at all), present (screw in hole), loose (thread not fully tightened), tight (thread fully tightened), tilted (screw axis not perpendicular to the mounting surface), and perpendicular (screw axis perpendicular to the mounting surface). Simultaneously, the precise 3D position coordinates (X,Y,Z) of the screw were obtained using a 3D measuring instrument as the ground truth position. The annotation process was cross-validated by multiple professionals to ensure annotation quality. The dataset contains 90% normal state samples and 10% abnormal state samples (including missing, loose, and tilted).

[0084] Step 503: Perform cross-view alignment processing and multi-scale feature extraction on the labeled multi-view image data to obtain the alignment feature map for training.

[0085] Specifically, the labeled multi-view image data is processed to generate aligned feature maps for training. First, images from different viewpoints are spatially aligned based on camera calibration parameters and 3D ground truth. Then, a multi-scale feature extraction network is applied to the aligned images, including: generating three resolution levels (1024×1024, 512×512, 256×256) using a Gaussian pyramid; extracting features at each resolution level using convolutional kernels of different sizes (7×7, 5×5, 3×3); and finally, integrating the multi-scale features into a unified aligned feature map (256×256×272) through a feature fusion network. This feature extraction method ensures that the training data contains multi-scale visual information about the screw. Step 504: Construct a multi-task detection model. The multi-task detection model includes a shared backbone network and multiple task branches, including a localization branch, a state classification branch, and a 3D correction branch.

[0086] Specifically, the system adopts an improved YOLOv8 multi-task detection model architecture, using CSPDarknet53 as a shared backbone network, which includes three parallel task branches: (1) a localization branch, used to output the bounding box (x, y, w, h) and category of the screw (e.g., front panel USB screw, rear panel main screw); (2) a state classification branch, used to output the screw state (missing / present, loose / tight, tilted / vertical); and (3) a 3D correction branch, used to output the spatial position deviation (Δx, Δy, Δz) of the screw. This multi-task architecture can share the feature extraction process and improve computational efficiency.

[0087] Step 505: Input the alignment feature map used for training into the multi-task detection model, and calculate the multi-task loss function based on the difference between the model output and the corresponding labeled content. The multi-task loss function includes localization loss, classification loss and 3D offset loss.

[0088] Specifically, the system inputs the alignment feature maps used for training into the multi-task detection model. For each training sample, the model outputs through three task branches: the localization branch outputs the predicted screw bounding box values ​​(xp, yp, wp, hp), the state classification branch outputs the six-dimensional state probability vector sp, and the three-dimensional correction branch outputs the predicted position offset values ​​(Δxp, Δyp, Δzp). These prediction results are compared with the ground truth values ​​labeled on the sample: the ground truth bounding box coordinates (xt, yt, wt, ht), the ground truth state label st, and the ground truth position offset (Δxt, Δyt, Δzt), and the loss for each task is calculated.

[0089] The specific loss calculation process is as follows: The localization loss adopts the CIoU loss function: LCIoU = 1 - IoU + d² / c² + α×v, where IoU is the intersection-union ratio of the predicted box and the ground truth box, d is the Euclidean distance between the center points of the predicted box and the ground truth box, c is the diagonal length of the smallest bounding rectangle that simultaneously contains the predicted box and the ground truth box, v is the aspect ratio similarity coefficient between the predicted box and the ground truth box, and α is the balance coefficient (set to 0.5).

[0090] The state classification loss uses weighted cross-entropy loss: Lcls = -Σ(wi×st,i×log(sp,i)), where st,i is the one-hot encoding of the true state, sp,i is the predicted state probability, and wi is the state weight (abnormal state weight 2.0, normal state weight 0.5). The 3D offset loss uses smoothed L1 loss: Loff = smoothL1(||(Δxp,Δyp,Δzp) - (Δxt,Δyt,Δzt)||). When the difference between the predicted offset and the true offset is less than a threshold of 1, squared loss is used; when it is greater than the threshold, L1 loss is used. The final multi-task loss function is: Ltotal = LCIoU + Lcls + Loff. Through this multi-task loss function design, the system can simultaneously optimize the screw detection positioning accuracy, state classification accuracy, and 3D position error, ensuring that the model achieves good performance on each task.

[0091] Step 506: Update the model parameters using the optimizer based on the multi-task loss function, repeat the training until the loss function converges, and obtain the pre-trained multi-task detection model.

[0092] Specifically, the system uses the AdamW optimizer for model training, setting the initial learning rate to 0.001 and the weight decay coefficient to 0.05. A phased training strategy is employed: first, the parameters of the backbone network are frozen for pre-training; after the model initially converges, fine-tuning is performed on the entire network. During training, model performance is evaluated by the convergence of the loss function on the validation set; training stops when the loss function stabilizes. This training strategy ensures both model convergence and improved training efficiency.

[0093] Step 50: Input the alignment feature maps of each detection unit at different angles into the pre-trained multi-task detection model to obtain the screw detection information corresponding to each detection unit. The screw detection information includes position information, state information and three-dimensional offset.

[0094] Specifically, the system inputs the aligned feature maps (each feature map size is 256×256×272) from six perspectives (C1 to C6) of each detection unit into a pre-trained YOLOv8 multi-task detection model for inference, obtaining screw detection information including position information, state information, and 3D offset. This multi-task detection model is based on the CSPDarknet53 backbone network and processes the input feature maps in parallel through three task branches to achieve comprehensive screw detection. The localization branch performs predictions at three scales (32×32, 64×64, and 128×128) using a feature pyramid structure. Each prediction head outputs bounding box coordinates (x, y, w, h) and class probability, using non-maximum suppression (NMS) to handle overlapping detection boxes with a threshold of 0.5. The system selects the detection result with the highest confidence as the final screw position information, including the screw center point coordinates, bounding box size, and class (e.g., "front panel USB screw," "rear panel main screw," etc.). The state classification branch outputs a six-dimensional state probability vector through a spatial attention-enhanced feature processing network, corresponding to six states: missing, present, loose, tight, tilted, and vertical. The system uses a soft maximum function to normalize the probability vector and selects the state with the highest probability as the screw's state determination result. For each detection unit, the state prediction results from the six perspectives are weighted and fused, with the weights proportional to the perspective quality, to obtain the final state information. The three-dimensional correction branch outputs the spatial position deviation (Δx, Δy, Δz) of each screw through a regression network.

[0095] Step 60: Generate a screw inspection report for the terminal device based on the inspection information of each screw, and update the historical inspection data.

[0096] Specifically, the screw inspection information is first integrated into structured data. For location information, the system converts bounding box coordinates and categories into standardized location descriptions, such as "second screw of front panel USB interface"; for status information, the system determines the anomaly type according to preset status judgment rules (e.g., a screw tilt angle exceeding 5 degrees is considered tilted); for 3D offset, the system converts spatial coordinate deviations into actual physical dimensions (accurate to 0.01 mm). The system generates an inspection report according to a preset format, which includes the following core content: basic information: recording the inspection timestamp, device ID (SN code), and workstation number; Anomaly Details: Records the screw location, anomaly type (missing / loose / tilted), 3D coordinate deviation (X / Y / Z) value, and confidence score (range 0-1) calculated based on the deviation value for each detection unit; Statistical Information: Calculates the occurrence rate of various anomalies, such as a missing rate of 0.5% per hour or 1.2% per day, and highlights high-frequency anomaly locations. When an anomaly is detected, the system automatically executes the following responses: sends an M57 host control signal to the production line via the PLC interface, triggers an audible and visual alarm, and stops the production line operation, awaiting manual re-inspection; writes the detection results to the SQL database, establishes a detection record indexed by the equipment SN code, and supports subsequent queries of historical anomalies for specific equipment (such as screw loosening status of a batch of terminal equipment). The system uses an incremental update mechanism to maintain historical data, automatically calculating and updating anomaly statistics every hour. For recurring anomaly locations, the system analyzes potential systemic problems using pattern recognition algorithms and highlights them prominently in the report.

[0097] As an optional embodiment, the screw inspection method in the terminal equipment manufacturing process may further include the following steps: Step 601: Calculate the consistency rate of state judgment and the variance of three-dimensional offset based on the screw detection information of each detection unit.

[0098] Specifically, the system receives screw detection information from each detection unit output by the multi-task detection model and calculates two key indicators for each screw detection unit: state judgment consistency rate and three-dimensional offset variance. Taking the second screw of the front panel USB interface as an example, the system first statistically analyzes the state judgment results from six detection perspectives (C1-C6): four perspectives determine it to be in a "loose" state, and two perspectives determine it to be in a "tight" state, resulting in a state judgment consistency rate of 66.7%. Simultaneously, the system collects the three-dimensional coordinate offsets of the screw measured at the six perspectives. For example, the measured values ​​in the Z-axis direction are 0.8mm, 0.9mm, 1.2mm, 0.85mm, 0.95mm, and 1.1mm, respectively, resulting in a Z-axis offset variance of 0.022mm². The system performs the same variance calculation process for the X-axis and Y-axis offsets. This multi-dimensional consistency analysis provides a quantitative basis for subsequent confidence assessment.

[0099] Step 602: Identify low-confidence screws whose state determination consistency rate is lower than the state threshold or whose three-dimensional offset variance exceeds the preset offset range.

[0100] Specifically, the state threshold is set to 85%, and the upper limit of the preset offset range variance is 0.015 mm². When the consistency rate of the detection unit's state determination is lower than 85% or the variance of the three-dimensional offset in any direction exceeds 0.015 mm², it is marked as a low-confidence screw. For example, the consistency rate of the state determination of the second screw of the front panel USB interface (66.7%) is lower than the threshold, and the variance of the Z-axis offset (0.022 mm²) exceeds the range, so it is identified as a low-confidence screw by the system. This dual-index evaluation mechanism can comprehensively identify the uncertainty of the detection results.

[0101] Step 603: Extract the alignment feature map of the low-confidence screw at each angle, calculate the sharpness score and occlusion degree of the feature map, and determine the optimal detection angle based on the sharpness score and occlusion degree.

[0102] Specifically, the system first extracts alignment feature maps of the low-confidence screws at various angles. For the sharpness evaluation of each alignment feature map, the system uses an improved Laplacian operator to calculate the second derivative of the image. Specifically, the alignment feature map is first denoised using Gaussian filtering, then the Laplacian operator value is calculated, and finally, the variance of the Laplacian operator is used as the sharpness score (range 0-100). A higher sharpness score indicates clearer edges and details in the image. For the evaluation of occlusion, the system uses a pre-trained deep convolutional neural network model. This model contains 5 convolutional layers and 3 fully connected layers, with the feature map as input and the occlusion score in the range of 0-1 as output. The closer the occlusion score is to 0, the lower the degree of occlusion. For example, the scores for the six perspectives are: C1 (clarity 75, occlusion 0.4), C2 (80, 0.3), C3 (85, 0.2), C4 (70, 0.5), C5 (65, 0.6), and C6 (60, 0.7). The system uses a weighted scoring mechanism: Overall score = Clarity score × 0.6 + (1 - Occlusion) × 100 × 0.4. For example, the overall score for perspective C3 is: 85 × 0.6 + (1 - 0.2) × 100 × 0.4 = 83. The calculated overall scores for the six perspectives are: C1 (69), C2 (74), C3 (83), C4 (64), C5 (61), and C6 (58). The system marks perspective C3, which has the highest overall score, as the optimal detection angle for the screw.

[0103] Step 604: Mark the location of the screws with low confidence in the screw inspection report, and attach the image taken at the optimal inspection angle as the basis for manual re-inspection.

[0104] Specifically, the system includes a dedicated "Low Confidence Warning" section in the inspection report, detailing information for each low-confidence screw. This includes: a description of the screw's location (e.g., "second screw at the front panel USB port"), the consistency rate of the status determination (66.7%), and the variance of the three-dimensional offset (X-axis 0.018mm², Y-axis 0.020mm², Z-axis 0.022mm²). Simultaneously, the system attaches the original image of the screw at its optimal inspection angle (C3 viewpoint), precisely marking the target screw's location within the image using a red rectangle (coordinates [120, 150, 180, 210]). This detailed recording and visualization provides ample reference for manual re-inspection.

[0105] Step 605: When there are abnormal detection units in the historical detection data where the number of times low-confidence screws appear is greater than the preset number, increase the number of shooting angles of the abnormal detection units in the next detection by the terminal device.

[0106] Specifically, the system tracks the confidence level of each detection unit in real time. When the database shows an abnormal detection unit where the number of low-confidence screw occurrences exceeds a preset number (3 times), the system automatically increases the shooting angle of that unit during the next terminal device detection. For example, if the second screw of the front panel USB interface is judged as having low confidence on three consecutive devices, the system adds two supplementary perspectives, C7 and C8, to the original six perspectives C1-C6. The shooting parameters of the newly added perspectives are determined through an interpolation algorithm: the horizontal angle of C7 is the median of C1 and C2, while the vertical angle remains unchanged; the horizontal angle of C8 is the median of C2 and C3, while the vertical angle remains unchanged.

[0107] This application also provides a computer storage medium that can store multiple instructions. The instructions are adapted to be loaded and executed by a processor as described above for a screw detection method in the production process of a terminal device. For details of the execution process, please refer to the specific description of the above embodiments, which will not be repeated here.

[0108] The following describes an electronic device for screw detection during the production process of a terminal device, provided by an embodiment of this application. Figure 3 This is a schematic diagram of an exemplary hardware structure of an electronic device provided in an embodiment of this application.

[0109] In some embodiments, the electronic device for screw detection during the production process of the terminal device is a computer device, or the electronic device for screw detection during the production process of the terminal device includes a computer device. The computer device includes a processor, memory, and a network interface connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores data. The network interface of the computer device is used to communicate with other external terminals or servers via a network connection. In some embodiments, the network interface can be a wired network interface; in some embodiments, the network interface can also be a wireless network interface. When the computer program is executed by the processor, it implements the methods in the embodiments of this application.

[0110] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0111] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0112] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0113] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A screw detection method in a terminal device production process, characterized by, The method includes: Obtain structural parameter information and historical test data of the terminal device, determine the spatial distribution area of ​​each screw to be tested in the terminal device based on the structural parameter information and historical test data, and divide the spatial distribution area into multiple detection units, with each detection unit corresponding to at least one screw to be tested; When the terminal device arrives at the detection station, it captures images of each detection unit from multiple preset angles to obtain images of each detection unit at different angles. Based on the spatial position information of each detection unit, the captured image is enhanced to generate target images of each detection unit at different angles; Cross-view alignment processing is performed on the target images of each detection unit at different angles to obtain the alignment feature maps of each detection unit at different angles. The alignment feature maps of each detection unit at different angles are input into the pre-trained multi-task detection model to obtain the screw detection information corresponding to each detection unit. The screw detection information includes position information, state information and three-dimensional offset. Based on the screw detection information, a screw detection report is generated for the terminal device, and the historical detection data is updated.

2. The screw detection method in a terminal device production process according to claim 1, characterized by, The spatial distribution area of ​​each screw to be inspected in the terminal device is determined based on the structural parameter information and historical detection data, and the spatial distribution area is divided into multiple detection units, including: Obtain the three-dimensional structural diagram of the terminal device and the screw specification parameters from the structural parameter information; Based on the screw anomaly records in the historical inspection data, the frequency of anomalies for each screw specification is determined. Based on the three-dimensional structural diagram and the frequency of anomalies, a heat map of screw density distribution is generated, and the area where the heat value exceeds a preset threshold is determined as the spatial distribution area of ​​each screw to be detected in the terminal device. Calculate the spatial distance and detection difficulty between adjacent screws to be detected within the spatial distribution area; Screws whose spatial distance is less than a first preset threshold and whose difference in detection difficulty is less than a second preset threshold are grouped into the same detection unit.

3. The screw detection method in a terminal device production process according to claim 1, wherein The spatial location information includes the three-dimensional coordinates and surface normal vectors of the detection units. The step of performing image enhancement processing on the captured images based on the spatial location information of each detection unit to generate target images of each detection unit at different angles includes: The relative distance between each detection unit and the shooting position corresponding to each preset angle is calculated based on the three-dimensional coordinates to obtain the distance attenuation coefficient; The reflectivity is calculated based on the angle between the surface normal vector of each detection unit and the shooting direction corresponding to each preset angle. Images with reflectivity exceeding the preset threshold are marked as reflective images. Brightness saturation detection is performed on the reflected image to identify overexposed pixels; The overexposed pixels are compensated for brightness using the distance attenuation coefficient, and the compensated image data is then subjected to high-pass filtering to obtain the target images of each detection unit at different angles.

4. The screw inspection method in the terminal equipment manufacturing process according to claim 1, characterized in that, Cross-view alignment processing is performed on the target images of each detection unit at different angles to obtain alignment feature maps of each detection unit at different angles, including: Key points and corresponding feature descriptors of the target image at different angles in each detection unit are extracted using the SIFT algorithm. Based on the feature descriptor, feature point matching is performed between target images from different angles to obtain matching point pairs; The RANSAC algorithm is used to filter the matching point pairs, and the homography matrix between target images at different angles is calculated based on the filtered matching point pairs; Based on the homography matrix, perspective transformation is performed on target images at different angles to align target images of the same detection unit at different angles to a unified reference viewpoint; Multi-scale feature extraction is performed on the aligned images at different angles to obtain the alignment feature maps of each detection unit at different angles.

5. The screw inspection method in the terminal equipment manufacturing process according to claim 4, characterized in that, The step of extracting multi-scale features from the aligned images at different angles to obtain alignment feature maps of each detection unit at different angles includes: The aligned images at different angles are downsampled according to different spatial resolutions; Feature extraction is performed on the downsampled image using convolution kernels of different sizes to obtain coarse-grained features and fine-grained features. The coarse-grained features are used to capture the overall outline and position information of the screw, while the fine-grained features are used to capture the texture details and edge information of the screw surface. Channel splicing is performed on the coarse-grained and fine-grained features at each angle to obtain the preliminary feature representation of each angle; Based on the angle between the shooting direction at each angle and the surface normal vector of the detection unit, the view quality weight of each angle is calculated, and the view quality weight is expanded into a weight map with the same spatial size as the preliminary feature representation, and then weighted element-wise with the preliminary feature representation of the corresponding angle. The extrinsic parameters of the camera at each angle are converted into position encoding vectors, and then concatenated with the corresponding angle-weighted feature representations along the channel dimension to obtain the aligned feature maps of each detection unit at different angles.

6. The screw inspection method in the terminal equipment manufacturing process according to claim 1, characterized in that, The method also includes a training process for a multi-task detection model, specifically including: Collect multi-view image data from multiple terminal devices of different models, including images of each detection unit at different angles; The multi-view image data is labeled, and the labeling content includes the bounding box coordinates of the screw, the screw type, the status label and the three-dimensional position truth value. The status label includes at least one of missing, present, loose, tight, tilted and vertical. Cross-view alignment and multi-scale feature extraction are performed on the labeled multi-view image data to obtain the alignment feature map for training. A multi-task detection model is constructed, which includes a shared backbone network and multiple task branches, including a localization branch, a state classification branch, and a 3D correction branch. The alignment feature map used for training is input into the multi-task detection model. The multi-task loss function is calculated based on the difference between the model output and the corresponding labeled content. The multi-task loss function includes localization loss, classification loss and 3D offset loss. Based on the multi-task loss function, the model parameters are updated by the optimizer, and the training is repeated until the loss function converges to obtain a pre-trained multi-task detection model.

7. The screw inspection method in the terminal equipment manufacturing process according to claim 1, characterized in that, The method further includes: Based on the screw detection information of each detection unit, calculate the consistency rate of state judgment and the variance of three-dimensional offset; Identify low-confidence screws whose state determination consistency rate is lower than the state threshold or whose three-dimensional offset variance exceeds the preset offset range; The alignment feature maps of the low-confidence screws at various angles are extracted, the sharpness score and occlusion degree of the feature maps are calculated, and the optimal detection angle is determined based on the sharpness score and occlusion degree. The location of the screws with low confidence is marked in the screw inspection report, and an image taken at the optimal inspection angle is attached as a basis for manual re-inspection. When there is an abnormal detection unit in the historical detection data where the number of times a low-confidence screw appears is greater than a preset number, the number of shooting angles of the abnormal detection unit will be increased during the next detection by the terminal device.

8. An electronic device for screw detection during the production process of terminal equipment, characterized in that, The electronic device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the electronic device to perform the method as described in any one of claims 1-7.

9. A computer program product containing instructions, characterized in that, When the computer program product is run on an electronic device for screw detection during the production process of a terminal device, the electronic device causes the electronic device to perform the method as described in any one of claims 1-7.

10. A computer-readable storage medium comprising instructions, characterized in that, When the instruction is executed on an electronic device for screw detection during the production process of a terminal device, the electronic device causes the electronic device to perform the method as described in any one of claims 1-7.