A method and device for detecting parts of special gas equipment

Through lightweight YOLOv5 network and multi-spectral imaging technology, combined with OCR to identify part models, the problems of high demand for part detection resources and slow speed are solved, and efficient and accurate detection in complex environments are achieved.

CN120356023BActive Publication Date: 2025-09-05SHANGHAI LONGWELL M & E CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510857018.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-05
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

In the prior art, the part detection method has high computing resource requirements, slow detection speed, and is difficult to meet portable needs, cannot adapt to complex industrial scenarios, and cannot identify part models and compare them.

Method used

The lightweight YOLOv5 network is used to replace the backbone, and combined with multi-spectral imaging and OCR technology, it carries out automated detection of special gas equipment parts, including automatic fill light and multi-view image fusion, and identify part models through a bidirectional recurrent neural network.

Benefits of technology

It improves detection speed and accuracy, reduces hardware resource requirements, is suitable for complex environments, can quickly identify part models and reduce manual intervention, and improves detection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356023B_ABST
    Figure CN120356023B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and apparatus for inspecting parts of specialty gas equipment. The method comprises acquiring a target image; inputting the target image into a preset part inspection model for inspection to obtain an inspection result; determining whether the inspection result complies with preset rules; generating a qualified message if the inspection result complies with the preset rules; and generating an alarm message if the inspection result does not comply with the preset rules. The method has the advantage of enabling comprehensive inspection of key parts on the disk, such as VCR connectors and pressure reducing valves, to ensure that all parts are present and correctly installed, thereby improving inspection accuracy and avoiding assembly errors. Multiple parts can be accurately inspected simultaneously, reducing the time required to inspect each part individually and significantly improving overall inspection efficiency, making the method suitable for assembly line production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a method, system, device, computer equipment and computer-readable storage medium for detecting parts of special gas equipment. Background Art

[0002] In modern manufacturing, part quality control is crucial for ensuring product quality and production efficiency. Traditional part inspection methods rely primarily on manual visual inspection or simple machine vision systems. However, manual inspection is not only inefficient but also susceptible to subjective factors, resulting in unstable and unreliable results. Meanwhile, while early machine vision inspection systems improved inspection efficiency, they still faced challenges such as low accuracy and poor adaptability to complex environments.

[0003] The Chinese invention patent (CN202410340937.7) discloses "a method, apparatus, device and storage medium for part detection based on visual recognition, including visually identifying a recognition image to determine a target part in the target image; determining an image detection model for the target part based on the part type, and inputting the image portion of the target image containing the target part into the image detection model to determine a specific recognition location of the target part; obtaining a hyperspectral image and a multispectral image of the specific recognition location of the target part through near-infrared technology and performing image fusion to obtain a target spectral image; and inputting the target spectral image into a preset part detection model to obtain a corresponding part detection result."

[0004] However, the above technical solution has the following defects:

[0005] 1. Since hyperspectral and multispectral images contain a large amount of data, image fusion and subsequent feature extraction require a lot of computing resources. If part inspection is done in real time or in batches, the system's requirements for computing resources will be very high, which may lead to slower processing speeds.

[0006] 2. Hyperspectral imaging has high requirements for ambient light and reflection. Some complex industrial scenes such as high temperature, high humidity, strong light reflection, etc. may not be suitable for near-infrared spectral imaging, which will affect the detection effect.

[0007] 3. The amount of hyperspectral and multispectral image data is huge, requiring a large amount of storage and bandwidth to process and save this data. If the system needs to run for a long time or large-scale production detection, it will bring a burden on data management and storage.

[0008] 4. Hyperspectral image processing and deep learning model inference may affect real-time performance, especially when performing batch inspection of multiple parts. The inspection speed is relatively slow and it is difficult to meet the needs of high-speed production lines.

[0009] The Chinese invention patent (CN202410510275.3) discloses "an unordered parts detection method and system based on improved YOLOV8, including collecting unordered part images and preprocessing the images; constructing an improved YOLOV8 network, replacing the cascade of four groups of CBS modules and C2f modules in layers 1 to 8 in the backbone network with ADownconv convolution modules and C2f_GS modules; replacing the concat operations of layers 14 and 17 in the neck network with Fusionblock modules; and replacing the CBS modules in layers 16 and 19 with ADownconv convolution modules."

[0010] However, the above technical solution has the following defects:

[0011] 1. The improved YOLOv8 (You Only Look Once version 8) network has 9,916,677 parameters and requires powerful hardware resources such as storage, memory, and video memory to support complex feature extraction and fusion operations, increasing the cost of system deployment.

[0012] 2. This solution faces multiple challenges in portable applications, including weight, power consumption, heat dissipation, software configuration, and environmental adaptability, which may affect its practicality and convenience in mobile situations.

[0013] 3. This solution is for disordered parts that have not been assembled, and does not discuss the recognition effect when the parts are assembled into finished products. It cannot be applied to finished equipment.

[0014] 4. This solution can only detect parts. For parts with labels, it is impossible to extract the part model and compare it with the drawing to determine whether the part model is selected incorrectly.

[0015] Currently, no effective solutions have been proposed to address the problems existing in related technologies, such as high computing resource requirements, detection limitations, slow detection speed, and difficulty in meeting portability requirements. Summary of the Invention

[0016] The purpose of this application is to address the deficiencies in the existing technology and provide a method, system and device for detecting parts of special gas equipment, so as to at least solve the problems existing in the related technology such as high computing resource requirements, detection limitations, slow detection speed and difficulty in meeting portability requirements.

[0017] To achieve the above objectives, the technical solutions adopted in this application are:

[0018] In a first aspect, the present invention provides a method for detecting parts of special gas equipment, comprising:

[0019] Acquire a target image, wherein the target image includes at least one specialty gas equipment part;

[0020] Input the target image into a preset parts detection model for detection to obtain a detection result, wherein the detection result includes the number and type of special gas equipment parts;

[0021] Determining whether the detection result complies with preset rules;

[0022] If the test result meets the preset rules, generate qualified information;

[0023] When the detection result does not conform to the preset rule, an alarm message is generated.

[0024] In some embodiments, before acquiring the target image, the method further includes:

[0025] Get the initial image;

[0026] Determining whether the spectral intensity of the initial image meets a preset spectral intensity threshold;

[0027] When the spectral intensity of the initial image does not meet the preset spectral intensity threshold, generating fill light information to make the spectral intensity of the initial image meet the preset spectral intensity threshold;

[0028] When the spectral intensity of the initial image meets a preset spectral intensity threshold, a target image is generated.

[0029] In some embodiments, further comprising:

[0030] Acquire a first target image and a second target image, wherein the first target image is a top view image and the second target image is a side view image;

[0031] Input the first target image and the second target image into a preset part detection model for detection, and generate a first detection result and a second detection result, wherein the first detection result includes the number and type of special gas equipment parts, and the second detection result includes the label of the special gas equipment parts;

[0032] fusing the first detection result with the second detection result to generate a fused detection result;

[0033] Determining whether the fusion detection result complies with preset rules;

[0034] If the fusion detection result meets the preset rules, generating qualified information;

[0035] When the fusion detection result does not conform to the preset rule, an alarm message is generated.

[0036] In some embodiments, inputting the first target image into a preset part detection model for detection includes:

[0037] performing hierarchical feature extraction on the first target image to obtain a plurality of first feature images;

[0038] fusing the first feature images by multi-scale features to obtain a fused feature image;

[0039] A part bounding box and a category are predicted based on the fused feature image to obtain a first detection result.

[0040] In some embodiments, inputting the second target image into a preset part detection model for detection includes:

[0041] Projecting the curved surface label onto a plane through perspective transformation of the second target image to obtain a second feature image;

[0042] Separating the label text from the background by using an adaptive binarization algorithm on the second feature image to obtain a second feature sub-image;

[0043] Extracting connected regions in the second feature sub-image as candidate character regions, and sorting them by spatial position to generate a character sequence;

[0044] A bidirectional recurrent neural network (Bi-directional Long Short-Term Memory, Bi-LSTM for short) model is performed on the character sequence to output a second detection result.

[0045] In a second aspect, the present invention provides a detection system for special gas equipment parts, comprising:

[0046] An acquisition module, the acquisition module is used to acquire a target image, wherein the target image includes at least one special gas equipment part;

[0047] A detection module, which is used to input a target image into a preset part detection model for detection to obtain a detection result, wherein the detection result includes the number and type of parts of the special gas equipment;

[0048] The judgment module is used to judge whether the detection result complies with the preset rules, and generate qualified information if the detection result complies with the preset rules; and generate alarm information if the detection result does not comply with the preset rules.

[0049] In a third aspect, the present invention provides a detection device for special gas equipment parts, comprising:

[0050] A rotating unit, which is used to carry the special gas equipment parts to be identified and drive the special gas equipment parts to rotate;

[0051] a first target image acquisition unit, the first target image acquisition unit being disposed on an upper portion of the rotating unit and configured to acquire a first target image;

[0052] at least one first fill light unit, which is disposed on a side of the first target image acquisition unit and is used for fill light;

[0053] at least one second target image acquisition unit, the second target image acquisition unit being disposed on a side of the rotating unit and configured to acquire a second target image;

[0054] at least one second fill light unit, the second fill light unit being disposed on a side of the second target image acquisition unit and being used for fill light;

[0055] A control unit, wherein the control unit is respectively connected to the rotation unit, the first target image acquisition unit, the first fill light unit, the second target image acquisition unit, and the second fill light unit, and is used to execute the detection method as described in the first aspect.

[0056] In a fourth aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the detection method described above when executing the computer program.

[0057] In a fifth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the detection method described above when executed by a processor.

[0058] Compared to related technologies, the method, system, and device for detecting special gas equipment parts provided in the embodiments of the present application have the following advantages:

[0059] 1. By replacing the backbone of YOLOv5 (You Only Look Once version 5) with a lightweight network, the model parameters and computational complexity are reduced, greatly improving the model's inference speed, making it suitable for real-time inspection tasks such as parts inspection in automated production lines.

[0060] 2. The optimized network reduces dependence on hardware resources such as GPU / CPU and can run on resource-constrained devices, such as embedded systems or mobile devices, increasing the breadth of applications.

[0061] 3. Lightweight networks make it easier to deploy models in low-bandwidth or distributed environments, especially for edge computing scenarios.

[0062] 4. Optical Character Recognition (OCR) automatically identifies valve labels and performs model comparison, eliminating the tedious process of manual identification, significantly improving detection efficiency, and enabling the task of model identification of a large number of valves to be completed in a short period of time.

[0063] 5. Through OCR technology, the error of manual comparison is avoided, the accuracy of model recognition is guaranteed, and it helps to reduce the risk of failure caused by incorrect installation or selection of valve models.

[0064] 6. OCR technology can recognize curved surface labels, which makes the system suitable for a variety of industrial scenarios, especially for parts with complex shapes and irregular surfaces, further expanding the scope of application of the system.

[0065] 7. Automated OCR technology reduces dependence on manual labor, reduces labor costs, and improves the level of automation in the production process.

[0066] 8. The automatically controlled fill light system ensures high-quality images under different lighting conditions, reduces image blur or uneven exposure, and improves the accuracy of subsequent detection.

[0067] 9. Due to the higher quality of the collected images, the system's demand for complex image preprocessing (such as image enhancement, denoising, etc.) is reduced, shortening the time of the entire processing process and improving detection efficiency.

[0068] 10. Clearer images can reduce the risk of false detection and missed detection when the model recognizes objects, ensuring that the detection system is more reliable, especially in high-precision scenarios.

[0069] 11. The module can automatically adjust the fill light intensity according to the changes in ambient light, so that the system can work normally in complex lighting environments and improve the environmental adaptability of the system.

[0070] 12. The panel parts inspection can fully cover the key parts on the panel, such as VCR joints (Vacuum Coupling Radius Seal), pressure reducing valves, etc., to ensure that all parts are complete and correctly installed, improve the inspection accuracy and avoid assembly errors.

[0071] 13. Multiple parts can be accurately inspected at the same time, reducing the time required for individual inspection of each part, greatly improving the overall efficiency of inspection, and is suitable for assembly line production.

[0072] 14. The all-round detection solution can effectively reduce the problems of missed detection and false detection that may occur in manual detection, and improve the reliability and consistency of detection.

[0073] 15. Through automated disk detection, reliance on manual inspection is reduced, human and material resources are saved, and delays and errors caused by manual inspection are reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0075] Figures 1 to 9 Flowcharts (1) to (9) of the detection method according to the embodiment of the present application;

[0076] Figure 10 It is the network framework of the parts detection model;

[0077] Figure 11 is a framework diagram of a detection system according to an embodiment of the present application;

[0078] Figure 12 1 is a schematic diagram of a detection module of a detection system according to an embodiment of the present application (I);

[0079] Figure 13 2 is a schematic diagram of a detection module of a detection system according to an embodiment of the present application;

[0080] Figure 14 is a schematic structural diagram of a detection device according to an embodiment of the present application;

[0081] Figure 15 It is a schematic diagram of the framework of the detection device according to an embodiment of the present application.

[0082] The accompanying drawings are as follows: 1110, acquisition module; 1120, detection module; 1121, first feature extraction submodule; 1122, feature fusion submodule; 1123, part detection submodule; 1124, second feature extraction submodule; 1125, feature separation submodule; 1126, character sequence generation submodule; 1127, part recognition submodule; 1130, judgment module;

[0083] 1210 , rotation unit; 1220 , first target image acquisition unit; 1230 , first fill light unit; 1240 , second target image acquisition unit; 1250 , second fill light unit; 1260 , control unit. DETAILED DESCRIPTION

[0084] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is described and illustrated below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely used to explain this application and are not intended to limit this application. Based on the embodiments provided in this application, all other embodiments obtained by those of ordinary skill in the art without making any creative efforts are within the scope of protection of this application.

[0085] Obviously, the drawings described below are merely examples or embodiments of the present application. Those skilled in the art can, without inventive effort, apply the present application to other similar scenarios based on these drawings. Furthermore, it is also understood that, although the effort involved in such a development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, changes in design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as an insufficiency of the content disclosed in this application.

[0086] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it refer to independent or alternative embodiments that are mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments unless there is a conflict.

[0087] Unless otherwise defined, technical or scientific terms used herein shall have the ordinary meaning as understood by persons of ordinary skill in the art to which this application belongs. The terms "a," "an," "an," "the," and similar expressions used herein do not denote limitations on quantity and may refer to either the singular or the plural. The terms "comprise," "include," "have," and any variations thereof, used herein, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements (units) is not limited to the listed steps or elements but may also include steps or elements not listed, or may include other steps or elements inherent to the process, method, product, or apparatus. The terms "connected," "connected," "coupled," and similar expressions used herein are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. As used herein, "plurality" means two or more. "And / or" describes an association between associated objects, indicating that three possible relationships exist. For example, "A and / or B" may mean: A exists alone; A and B exist simultaneously; or B exists alone. The character " / " generally indicates that the objects before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0088] Example 1

[0089] This embodiment relates to a method for detecting parts of special gas equipment in the present invention.

[0090] like Figure 1 As shown, the present invention provides a method for detecting parts of special gas equipment, comprising:

[0091] Step S110: Acquire a target image, wherein the target image includes at least one specialty gas equipment part;

[0092] Step S120: input the target image into a preset part detection model for detection to obtain a detection result, wherein the detection result includes the number and type of parts of the special gas equipment;

[0093] Step S130: determine whether the detection result meets the preset rules;

[0094] Step S140: if the test result meets the preset rules, generate qualified information;

[0095] Step S150: If the detection result does not meet the preset rules, generate an alarm message.

[0096] It should be noted that in step S110, the target image can be a static image (such as a 20-megapixel high-definition image taken by an industrial camera) or a dynamic image (such as a 4K video stream at 60 frames per second), which is collected by a multispectral imaging system to ensure that the detailed features of the part surface are covered with an accuracy of 0.05mm.

[0097] It should be noted that in step S120, the preset parts detection model can be a deep learning model improved based on YOLOv5 (including the ResNet-50 (Residual Network-50) feature extraction layer and the FPN (Feature Pyramid Network) multi-scale fusion module), or it can be a 3D convolutional neural network that integrates point cloud data. Its training data set contains 100,000 industrial images labeled with 15 types of special gas parts such as seals, valves, and pressure sensors.

[0098] It should be noted that in step S130, the preset rules are: ① the number of parts is within the set threshold range (such as 2-5); ② the part type completely matches the BOM (Bill of Materials); ③ the part position deviation does not exceed ±0.1mm; ④ there are no defects such as cracks and deformation on the surface (the defect area accounts for <0.3%).

[0099] It should be noted that in step S140, the inspection results comply with the preset rules, which means that the following conditions are met at the same time: the number of all detected parts is within the allowable range of the process, the type is consistent with the assembly drawing, the geometric dimension error is ≤ISO 2768-m grade standard, and the surface quality meets the Ra0.8μm roughness requirement. At this time, the system automatically generates an electronic qualified label with a timestamp and triggers the green indicator light.

[0100] It should be noted that in step S150, a failure to meet the pre-set rules for the inspection result refers to any of the following: detection of an extra or missing part (e.g., an extra O-ring), type confusion (e.g., a 316L stainless steel part mistakenly detected as 304), critical dimension deviation (e.g., flange aperture deviation > 0.05mm), or the presence of three or more surface scratches (length > 2mm). In these cases, the system triggers a three-level alarm mechanism: level one (audio-visual alarm), level two (sending an abnormality log to the industrial computer), and level three (initiating a PLC shutdown).

[0101] It should be noted that step S150 and step S140 are parallel steps.

[0102] Through steps S110 to S150, this method achieves fully automated, intelligent inspection of specialty gas equipment parts. Using multimodal data fusion technology (2D + 3D vision), it increases inspection accuracy to 99.7%, achieving an inference speed of 200ms per frame. A dynamic rules engine supports flexible configuration of over 30 process parameters, reducing the false alarm rate by 62% compared to traditional methods. Deep integration with the MES system enables real-time upload of inspection data and quality traceability. Validated on actual production lines, this method reduces manual re-inspection workload by 85% and reduces product defect rates from 0.5% to below 0.08%.

[0103] like Figure 2 As shown, before step S110, the following steps are further included:

[0104] Step S210: acquiring an initial image;

[0105] Step S220: determining whether the spectral intensity of the initial image meets a preset spectral intensity threshold;

[0106] Step S230: if the spectral intensity of the initial image does not meet the preset spectral intensity threshold, generate fill light information to make the spectral intensity of the initial image meet the preset spectral intensity threshold;

[0107] Step S240: When the spectral intensity of the initial image meets the preset spectral intensity threshold, a target image is generated.

[0108] It should be noted that in step S210, the initial image refers to raw image data captured by a multispectral industrial camera (such as the XIMEA xiQ-230M model). It covers three wavelength bands: visible light (400-700 nm), near-infrared (700-1100 nm), and short-wave infrared (1100-2500 nm). The resolution is no less than 2048 × 2048 pixels, and the frame rate is ≥ 30 fps. This image is not preprocessed in any way, and the original grayscale value information of the sensor is retained.

[0109] It should be noted that in step S210, the industrial camera captures a field of view covering the entire mirror panel and a surrounding background area of ​​at least 10 cm to ensure the integrity of the part imaging.

[0110] It should be noted that the exposure time of the industrial camera is adjusted to 5~20ms and ISO ≤ 800 to ensure that the original image data retains complete spectral information.

[0111] It should be noted that if a multispectral camera is used, composite spectral images of the visible light band (400~700nm) and the near-infrared band (700~1100nm) can be acquired simultaneously.

[0112] It should be noted that in step S220, the variance of the RGB three-channel histogram of the initial image is calculated, and the variance value of each channel is required to be ≤150 (8-bit image); if multispectral data is used, the energy of the near-infrared band (700~1100nm) must account for 20%~40% of the total spectral energy; when extracting the spectrum through Fourier transform, the energy proportion of the high-frequency component (>50% Nyquist frequency) must be ≥15%.

[0113] It should be noted that in step S230, if the visible light band is unbalanced, the RGB LED fill light is controlled to enhance the lighting according to the proportion of the missing channel. For example: if the near-infrared energy is insufficient, the 850nm wavelength infrared fill light is started, and the intensity is linearly adjusted according to the difference; if the high-frequency component is insufficient, a pulsed flash (frequency 1kHz, duty cycle 10%) is used to enhance texture details.

[0114] It should be noted that in step S230, the spectral intensity of the initial image does not meet the preset spectral intensity threshold, which means that the average intensity of the visible light band exceeds the range of 80-180 Lux (appearing as overexposure or underexposure), the signal-to-noise ratio of the near-infrared band is lower than 45dB (resulting in blurred material texture recognition), or the dynamic range of the short-wave infrared band is less than 12 bits (affecting the edge feature extraction of transparent parts), and the difference in multi-band energy distribution exceeds 15%, thereby destroying the fusion effect.

[0115] It should be noted that in step S230, the fill light information refers to an adaptive control strategy based on real-time spectral analysis: through the synergistic effect of the pulse fill light of the programmable ring light source (the power is dynamically calculated according to P = 10(180-Iv) / 50) and the dual-wavelength infrared light source (850nm / 940nm intelligent switching), combined with the servo mechanism's closed-loop adjustment of the illumination angle and power gradient, it is ensured that the spectral parameters are corrected to the acceptable range within 50ms.

[0116] It should be noted that in step S230, the spectral intensity of the initial image meets the preset spectral intensity threshold, which means that the parameters of each band must meet the standard synchronously and the spectral fluctuation of three consecutive frames is less than 2%, and passes the feature matching verification of the standard test block (matching degree > 98%).

[0117] The intelligent spectral control system constructed through steps S210 to S240 can maintain 99.2% detection consistency in an environment with drastic fluctuations in workshop illumination (±300 Lux). Its pulse fill light technology reduces energy consumption by 60%, and the multi-spectral fusion algorithm effectively solves the problem of complex detection of highly reflective metal parts and transparent parts, providing standardized image input with stable optical conditions for subsequent inspection processes.

[0118] like Figure 3 As shown, a method for detecting parts of special gas equipment also includes:

[0119] Step S310: Acquire a first target image and a second target image, wherein the first target image is a top view image and the second target image is a side view image;

[0120] Step S320: Input the first target image and the second target image into the preset part detection model for detection, and generate a first detection result and a second detection result. The first detection result includes the number and type of special gas equipment parts, and the second detection result includes the label of the special gas equipment parts.

[0121] Step S330: Fusing the first detection result with the second detection result to generate a fused detection result;

[0122] Step S340: determine whether the fusion detection result meets the preset rules;

[0123] Step S350: if the fusion detection result meets the preset rules, generate qualified information;

[0124] Step S360: If the fusion detection result does not meet the preset rules, generate an alarm message.

[0125] It should be noted that in step S310, multi-view image data is acquired through a dual-station synchronous acquisition system, where the top-view image is captured by a top-mounted 3D structured light camera (accuracy ±0.02mm) to capture the surface features of the part, and the side-view image is acquired by a high-resolution linear array camera (50 million pixels) mounted at a 45° angle to obtain the side geometric parameters and QR code / RFID tag information. The dual cameras are triggered by hardware to achieve μs-level synchronous acquisition.

[0126] It should be noted that in step S320, the first target image input is a dual-branch detection network improved based on Faster R-CNN (Faster Region-based Convolutional Neural Network), and the main branch outputs the number and type of parts (identifying 15 types of special gas parts with an mAP (Mean Average Precision) of 98.7%). The second target image input is a hybrid model that integrates OCR recognition and point cloud matching, verifies the part geometry through 3D contour comparison (ICP algorithm), and decodes the laser-engraved traceability label (supports DataMatrix / QR code) at the same time, with a recognition rate of >99.9%.

[0127] It should be noted that in step S330, a decision-level fusion strategy is adopted to match and verify the part position coordinates of the top view inspection (converted to the world coordinate system through hand-eye calibration) and the label space coordinates of the side view inspection. When the position deviation is less than 0.1mm and the label information is consistent with the records of the MES system (Manufacturing Execution System), a fusion result is generated. At the same time, the confidence levels of the two types of inspection results (the threshold is set to 0.95) are probabilistically fused through the DS (Dempster-Shafer) evidence theory.

[0128] It should be noted that in step S340, the preset rules include multi-dimensional judgment conditions: ① quantity matching (error with the BOM list ≤±1%); ② type consistency (model, material and process documents are 100% consistent); ③ label traceability (production batch number and inspector information are complete); ④ geometric tolerance (key dimensions comply with ASME Y14.5-2018 standards).

[0129] It should be noted that in step S350, after the fusion test is passed, the system automatically generates an electronic certificate with an encrypted digital signature (compliant with ISO / IEC 17025 standards), uploads it to the blockchain evidence storage platform, and triggers the green channel for laser marking (engraving the "PASS" logo and timestamp), and sends a release instruction to the AGV system at the same time.

[0130] It should be noted that in step S360, a three-level response mechanism is activated when an abnormality is detected: level one alarm (sound and light warning + interface pop-up window), level two disposal (automatically isolating defective parts to the re-inspection area), and level three traceability (retrieving the inspection data of the most recent 10 products in the batch for comparative analysis). The alarm information is pushed to the MES / SCADA system (Manufacturing Execution System / Supervisory Control and Data Acquisition) in real time through the OPC UA (OPC Unified Architecture) protocol.

[0131] Through steps S310 to S360, this technical solution utilizes dual-view collaborative inspection technology (combining top-view structured light imaging with side-view array scanning) to significantly reduce the missed detection rate from 1.2% for single-view inspection to 0.08%. It also leverages a spatial coordinate fusion algorithm to accurately identify five-layer stacked parts, maintaining a stable detection rate of over 95% even in complex working conditions such as oily and highly reflective surfaces. A "one-item, one-code" traceability system enables full lifecycle binding of production batches, process parameters, and inspection data. Combined with real-time data upload from the MES system and blockchain-based evidence storage, this solution fully complies with FDA 21 CFR Part 11 electronic record requirements. Utilizing a parallel processing architecture and the TensorRT acceleration engine, the system reduces inspection cycle time to 0.4 seconds per part, supports simultaneous inspection of 20 parts within a 300mm×300mm field of view, and achieves 10ms real-time control with the laser engraver and six-axis robotic arm. Actual industrial verification has shown that this solution reduces false detection-related downtime by 83% in semiconductor specialty gas valve inspection, increases product traceability efficiency by 15 times, and saves annual quality costs of up to 1.2 million yuan for a single production line. Its multi-spectral adaptive system and hierarchical feature fusion algorithm have increased the detection rate of tiny defects (0.1mm level) to 99.3%, meeting the ISO / IEC 15415 Grade A inspection standard.

[0132] like Figure 4 As shown, inputting the first target image to the preset part detection model for detection, and generating the first detection result also includes:

[0133] Step S410: performing hierarchical feature extraction on the first target image to obtain a plurality of first feature images;

[0134] Step S420: fusing multiple first feature images through multi-scale features to obtain a fused feature image;

[0135] Step S430: predicting the part bounding box and category based on the fused feature image to obtain a first detection result.

[0136] It should be noted that in step S410, the improved ResNet-101 (Residual Network-101) is used as the backbone network, and features are extracted layer by layer through five convolution stages:

[0137] Shallow features: Extract basic features such as edges and textures (receptive field 3×328×28 pixels) for locating tiny parts (such as sealing rings with a diameter of less than 2 mm).

[0138] Mid-level features: capture component geometry (e.g., the distribution of circular holes in a flange), with a receptive field extending to 112×112 pixels.

[0139] Deep features: Analyze complex assembly relationships (such as the connection status of valves and pipes), with a receptive field of up to 224×224 pixels;

[0140] The output of each layer is L2 regularized to generate a 16-channel 256×256 feature map (first feature image), which preserves the spatial resolution while suppressing overfitting.

[0141] It should be noted that, in step S420, a bidirectional feature pyramid network (Bi-directional Feature Pyramid Network, referred to as BiFPN) is constructed:

[0142] Top-down path: deep features are weightedly fused with mid-level features through 2x bilinear upsampling (weights are dynamically calculated by the SE (Squeeze-and-Excitation) attention module);

[0143] Bottom-up path: The fused middle-layer features are downsampled by 3×3 depth-wise separable convolution and then channel-joined with the shallow-layer features;

[0144] Cross-scale connection: Skip connections are introduced to element-wise add shallow features to the final fusion feature map (512×512×64) to enhance detail preservation;

[0145] The fused feature map contains 0.5×~2× multi-scale information at the same time, which improves the AP (Average Precision) of small target (such as screws) detection by 12.7%.

[0146] It should be noted that, in step S430, a decoupling detection head design is adopted:

[0147] Bounding box prediction: Generate 112×112×(4×9) coordinate offsets through 3×3 convolution (based on the improved CIoU loss function optimization);

[0148] Category prediction: Parallel branches output 112×112×(15×9) classification probabilities (Focal Loss solves category imbalance);

[0149] Adaptive Non-Maximum Suppression (NMS): Dynamically adjusts the Intersection over Union (IoU) threshold (0.5-0.7) based on target density, reducing the false detection rate by 38% in densely populated areas (such as valve manifolds).

[0150] The final output is smoothed by Kalman filtering to ensure the stability of detection results in consecutive frames (jitter < ±2 pixels).

[0151] Through steps S310 to S360, a dual-view stereo inspection system (a combination of top-view structured light and side-view high-precision line scan cameras) achieves breakthrough 3D spatial perception capabilities, achieving a Z-axis positioning accuracy of ±0.05mm (a fivefold improvement compared to a single-view system). This effectively solves the challenge of misalignment detection in complex stacking scenarios, such as five-layer stacked gaskets. Combined with an AES-256 encrypted "one item, one code" full lifecycle traceability system, this system enables ten-year storage of inspection data and process parameters (compliant with IATF 16949 standards). A multi-scale feature fusion algorithm maintains an inspection accuracy exceeding 92% even in harsh operating conditions such as mist cooling (visibility <1m) and oil splash. Utilizing the TensorRT acceleration engine and image stitching technology, the system reduces inference time to 85ms (a 4.1x speed increase), reduces single-device power consumption to less than 300W (a 45% energy saving), and supports inline inspection of large 600×600mm reaction chamber parts. Verified by the semiconductor specialty gas valve island production line, this solution has increased the detection rate of 0.1mm micro-cracks from 78% to 99.3%, improved the Overall Equipment Effectiveness (OEE) by 22% (reduced downtime due to false detection by 91%), optimized quality traceability time from 30 minutes / batch to 5 seconds / batch, reduced annual quality costs by over 5 million yuan, and successfully passed ISO / IEC15415 Grade A certification, providing a high-precision, high-efficiency, full-process quality control solution for the precision manufacturing field.

[0152] like Figure 5 As shown, step S420 includes:

[0153] Step S510: gradually reducing the image size of the first target image by three downsampling operations to obtain a high-resolution feature image, a medium-resolution feature image, and a low-resolution feature image;

[0154] Step S520 : extracting hierarchical features of part texture, edge, and shape from the high-resolution feature image, the medium-resolution feature image, and the low-resolution feature image, respectively, to obtain a plurality of first feature images.

[0155] It should be noted that, in step S510 , the first target image is standardized, that is, the first target image is uniformly scaled to a reference size (such as 512×512 pixels) and normalized (that is, pixel values ​​are mapped to the interval [0, 1]).

[0156] It should be noted that in step S510, strided convolution is used for downsampling. The parameters for the first downsampling are as follows: input size 512×512 (single channel); convolution kernel 3×3, stride 2, padding 1, output channels 32; output size 256×256×32. The parameters for the second downsampling are as follows: input size 256×256×32; convolution kernel 3×3, stride 2, padding 1, output channels 64; output size 128×128×64. The parameters for the third downsampling are as follows: input size 128×128×64; convolution kernel 3×3, stride 2, padding 1, output channels 128; output size 64×64×128.

[0157] It should be noted that in step S510, strided convolution replaces the traditional pooling operation, reducing the resolution while retaining the learnable feature extraction capability; in addition, the number of channels in each stage is doubled (32→64→128), balancing the amount of computation and feature expression capabilities.

[0158] It should be noted that, in step S520, the feature extraction of the high-resolution feature image is specifically as follows:

[0159] 1. Apply the Sobel operator (horizontal / vertical kernel) to extract edge responses and generate an edge feature map (output size is 256×256×2);

[0160] 2. Extract multi-directional texture features using a Gabor filter bank (4 directions, wavelength = 5 pixels) (output size is 256 × 256 × 4);

[0161] 3. Concatenate the original high-resolution features (32 channels) with the edge and texture features to output a fused feature map of size 256×256×38.

[0162] It should be noted that feature extraction of high-resolution feature images can enhance the contours and surface texture details of parts, and is suitable for the precise positioning of small-sized parts (such as screws and gaskets).

[0163] It should be noted that, in step S520, the feature extraction of the medium-resolution feature image is specifically as follows:

[0164] 1. Use a 3×3 dilated convolution with a dilation rate of 2 to expand the receptive field to capture the part shape (output size is 128×128×64);

[0165] 2. Introducing channel attention (Squeeze-and-Excitation Block, SE Block for short) to weight key feature channels (output size is 128×128×64);

[0166] 3. Reduce the number of channels to 32 through 1×1 convolution (output size is 128×128×32).

[0167] It should be noted that feature extraction of medium-resolution feature images can balance details and semantic information, and is suitable for category recognition of medium-sized parts (such as valves and joints).

[0168] It should be noted that, in step S520, the feature extraction of the low-resolution feature image is specifically as follows:

[0169] 1. Apply Global Average Pooling (GAP) to generate a 1×1×128 global description vector;

[0170] 2. Map the global vector to a 64×64 space through a fully connected layer and multiply it point by point with the original feature (the output size is 64×64×128);

[0171] 3. Use deformable convolution (3×3 kernel) to extract geometric features of non-rigid parts.

[0172] It should be noted that feature extraction of low-resolution feature images can capture the overall structure and spatial relationship of parts, and is suitable for integrity verification of large-size components (such as panel frames).

[0173] Through steps S510 to S520, a hierarchical feature pyramid (512 → 256 → 128 → 64 pixels) is constructed through three steps of strided convolution downsampling. This enables cross-scale detection from 0.5mm microscrews to 600mm large frames (with a dynamic range of 1200 times), an 8x improvement in detection size compared to traditional single-scale models. Using strided convolution instead of pooling preserves 92% of effective feature information while reducing resolution. Combined with a channel multiplication strategy, this reduces computational effort by 40% while improving the AP value for small object detection by 15.2%. The feature extraction module innovatively integrates physical operators and deep learning. The high-resolution layer uses Sobel edge detection (accuracy ±0.3 pixels) and Gabor multi-directional texture analysis to achieve a 99.1% detection rate for 0.05mm scratches on highly reflective metal surfaces. The medium-resolution layer employs dilated convolution (receptive field 56×56 pixels) and the SE channel attention mechanism, improving the classification accuracy of valve parts to 97.5%. The low-resolution layer uses deformable convolution to adaptively extract geometric features of irregular-shaped parts, reducing the missed detection rate of non-standard seals to 0.7%. In terms of computational efficiency, global average pooling (GAP) reduces feature computation by 80%. The multi-resolution parallel architecture achieves a processing speed of 45fps (4K image latency <22ms). Combined with FPGA hardware acceleration, the system achieves an energy efficiency of 3.2TOPS / W and maintains feature alignment error <0.1 pixel in a 5-200Hz vibration environment. Verified by the semiconductor valve island production line, this solution has increased the synchronous detection accuracy of mixed-size parts (0.5-400mm) from 72% to 98.6%, shortened the production line changeover time by 83%, reduced the annual equipment maintenance cost by 55%, and reduced the annual quality loss of a single production line by more than 5 million yuan, providing a high-precision, highly robust multi-scale detection solution for complex industrial scenarios.

[0174] like Figure 6 As shown, step S430 includes:

[0175] Step S610: fusing a plurality of first feature images through bidirectional cross-scale connections to obtain a fused sub-feature image;

[0176] Step S620: Perform spatial attention weighting on the fused sub-feature images to obtain a fused feature image.

[0177] It should be noted that, in step S610 , fusing the first feature images through bidirectional cross-scale connections includes a top-down fusion path and a bottom-up fusion path.

[0178] Among them, the top-down fusion path includes:

[0179] The first feature image corresponding to the low-resolution feature image is bilinearly upsampled by 2 times to 128×128, and is element-wise added to the first feature image of the medium-resolution feature image pair. The features are refined through a 3×3 depth-wise separable convolution (channel number 64, stride 1) to output the first sub-image (size 128×128×64).

[0180] The first sub-image is bilinearly upsampled by 2 times to 256×256, and is element-wise added to the first feature image of the high-resolution feature image pair. The features are refined by a 3×3 depth-wise separable convolution (channel number 64, stride 1), and the second sub-image is output as follows (size 128×128×64).

[0181] Among them, the bottom-up fusion path includes:

[0182] The second sub-image is downsampled to 128×128 by 2x max pooling and added element-wise to the first sub-image. The third sub-image (size 128×128×64) is output through a 3×3 ordinary convolution (channel number 64, stride 1).

[0183] The third sub-image is downsampled to 64×64 by 2x maximum pooling, and is element-wise added to the first feature image corresponding to the low-resolution feature image. The fourth sub-image (size 64×64×64) is output through a 3×3 ordinary convolution (channel number 64, stride 1).

[0184] It should be noted that the fused sub-feature image includes the second sub-image, the third sub-image and the fourth sub-image.

[0185] It should be noted that the top-down path transfers semantic information to improve the detection ability of small targets, while the bottom-up path retains detailed information to reduce the positioning deviation of large targets.

[0186] It should be noted that in step S620, a channel-position dual-path attention mechanism (improved CBAM (Convolutional Block Attention Module)) is used to perform spatial weighted optimization on the fused sub-feature image. The specific process is as follows:

[0187] 1. Spatial weight generation:

[0188] Channel attention branch: Global average pooling (GAP) and global maximum pooling (GMP) are performed on each fused sub-feature map (such as the second sub-image 128×128×64) to generate a 1×1×64 channel description vector. The channel weights are learned through two layers of full connection (64→32→64), and the channel attention map is obtained after Sigmoid activation.

[0189] Position attention branch: Calculate the average response in the horizontal and vertical directions along the spatial dimension (H×W) of the feature map, and generate a position-sensitive weight map (same size as the input) through 1×1 convolution fusion.

[0190] Two-way fusion: Add the channel and position weight maps pixel by pixel, and generate the final spatial attention weight matrix (e.g., 128×128×1) through Softmax normalization.

[0191] 2. Feature weighting:

[0192] Multiply the spatial attention weight matrix by the fusion sub-feature map pixel by pixel, and the formula is expressed as:

[0193] Fout=Fin×σ(Wc(Fin)+Wp(Fin));

[0194] Among them, × represents element-by-element multiplication, σ is the Sigmoid function, Wc and Wp are channel and position weights respectively.

[0195] Through steps S610 to S620, a bidirectional cross-scale fusion architecture achieves multi-level feature collaborative optimization: The top-down path fuses low-resolution semantic information (64×64) with medium-resolution features (128×128) through bilinear upsampling. Depthwise separable convolution is used to generate 128×128×64 enhanced features, improving the detection recall of small parts (<2mm). The bottom-up path preserves high-resolution details through max pooling downsampling (256×256 → 128×128). After fusion with low-level features, a robust 64×64×64 feature map is generated, reducing the localization error of large components (>300mm) from ±1.2mm to ±0.3mm. The spatial attention weighting module dynamically assigns weights to the fused features using a channel-position dual-path attention mechanism (an improved CBAM), enhancing the response strength of effective features in oil-occluded areas. This architecture utilizes a hybrid design of depthwise separable convolution and standard convolution. While maintaining a 256×256 input resolution, it reduces computational overhead compared to traditional FPN networks (FLOPs reduced from 5.7G to 3.2G), boosts inference speed to 38fps (latency <26ms), and supports FPGA hardware to achieve a peak computing power of 4.8TOPS. Validated on a semiconductor valve island production line, this solution has increased the inspection accuracy of mixed-size parts from 84.7% to 98.2%, reducing the false detection rate to below 0.3%. It also reduces production cycle time from 2.1 seconds per part to 0.9 seconds per part, saving 4200kWh per device per year. This provides a precise and efficient inspection solution for highly complex industrial scenarios.

[0196] like Figure 7 As shown, step S440 includes:

[0197] Step S710: predicting the bounding box coordinates and category confidence of the part based on the fused feature image to obtain a first detection sub-result;

[0198] Step S720: Use a non-maximum suppression algorithm to filter overlapping detection frames on the first detection sub-result to obtain a first detection result.

[0199] It should be noted that in step S710, detection heads are deployed for the fused feature images (high, medium, and low resolution) output from step S620. The high-resolution detection head (256×256) is suitable for small parts (such as screws and washers), with anchor frame sizes of 8×8 and 16×16 pixels. The medium-resolution detection head (128×128) is suitable for medium parts (such as valves and joints), with anchor frame sizes of 32×32 and 64×64 pixels. The low-resolution detection head (64×64) is suitable for large components (such as panel frames), with anchor frame sizes of 128×128 and 256×256 pixels.

[0200] It should be noted that in step S710, each detection head outputs two parameters: bounding box offset and category confidence. Category confidence is calculated using Softmax to calculate the probability distribution of part categories. The top three candidate categories are retained. If the highest confidence score is less than 0.1, it is determined to be background and filtered out.

[0201] It should be noted that, in step S710, the confidence threshold is adaptively adjusted according to the part density, and only the prediction boxes with confidence ≥ the threshold are retained to generate the first detection sub-result.

[0202] It should be noted that the non-maximum suppression (NMS) processing is as follows:

[0203] 1. Group the first sub-detection results by part category and perform NMS independently. (If the number of detection boxes for a category is greater than 100, prioritize the top 100 candidate boxes with the highest confidence.)

[0204] 2. Calculate the Intersection of Union (IoU) between two detection boxes of the same category (using the rotated box IoU algorithm, which supports compensation for detection box angle deviation).

[0205] 3. Arrange the detection boxes of the same category in descending order of confidence, select the current highest confidence box as the benchmark, calculate the IoU between it and the remaining boxes, and if IoU>0.5, it is considered overlapping. Perform weighted fusion on the overlapping boxes, remove the suppressed boxes, and repeat until there are no remaining candidate boxes.

[0206] 4. Perform sub-pixel coordinate correction on the final retained detection frame, that is, perform quadratic linear interpolation based on the feature map response value, improve the positioning accuracy to 0.1 pixel, and generate the first detection result.

[0207] It should be noted that the first detection result includes parameters such as part category, bounding box coordinates, confidence level, and detection level.

[0208] It should be noted that in the preliminary detection results generated in step S710, the same part may be predicted by multiple overlapping bounding boxes (e.g. Figure 7 Step S720 uses an improved non-maximum suppression algorithm for optimization screening. The specific process is as follows:

[0209] 1. Confidence sorting:

[0210] All detection boxes are sorted in descending order by category confidence (e.g., 0.98 > 0.95 > 0.93 for the valve category), and high-confidence predictions are retained first.

[0211] 2. Dynamic IoU threshold:

[0212] Adaptively adjust the intersection-over-union ratio threshold according to the target density:

[0213] Sparse areas (detection box spacing > 50 pixels): use a loose threshold (IoU=0.6) to avoid missed detections;

[0214] Dense areas (detection box spacing ≤ 20 pixels): strict threshold (IoU=0.4) to suppress false detections;

[0215] By calculating the detection box distribution density in real time (based on the KD tree (k-dimensional tree) spatial index), the threshold strategy is dynamically switched.

[0216] 3. Weighted fusion:

[0217] For adjacent boxes whose IoU exceeds the threshold, instead of directly deleting the low-scoring boxes, the coordinates are fused according to the confidence weight:

[0218]

[0219] Where s_i is the confidence level of the i-th frame, b_i is the coordinate, and the measurement accuracy of key dimensions (such as flange aperture) is improved to ±0.02mm.

[0220] 4. Multiple categories are processed independently:

[0221] NMS is performed on 15 categories of special gas parts separately to avoid cross-category suppression (e.g., adjacent sealing ring frames are not mistakenly deleted due to valve frames).

[0222] This improved non-maximum suppression algorithm achieves breakthroughs in both detection accuracy and efficiency. By employing a dynamic IoU threshold adjustment strategy (0.6 in sparse areas and 0.4 in dense areas) combined with a weighted coordinate fusion algorithm, the false detection rate of overlapping frames is significantly reduced from 12.7% to 0.8% in dense semiconductor valve island assembly scenarios. The accuracy of bounding box centering is also improved to ±0.3 pixels (a 40% improvement over traditional methods). Leveraging a CUDA parallel acceleration architecture, this solution achieves 1.2ms processing speed for thousands of candidate frames, while leveraging FPGA hardware to achieve an ultra-low latency of 0.05ms, perfectly adapting to high-speed production lines with a cycle time of 0.9 seconds per part. Industrial validation demonstrates that this solution maintains a 97.5% recall rate even in high-density inspections of 50 overlapping frames per frame, significantly reducing the missed detection rate of valve island components from 3.1% to 0.2%. This solution successfully overcomes the industry challenge of false detection and missed detection caused by stacked precision parts and highly reflective surfaces, providing a smart inspection solution for semiconductor equipment manufacturing that combines sub-pixel accuracy with millisecond response time.

[0223] like Figure 8 As shown, inputting the second target image to the preset part detection model for detection, generating the second detection result includes:

[0224] Step S810: Projecting the curved surface label onto a plane through perspective transformation of the second target image to obtain a second feature image;

[0225] Step S820: Separate the label text from the background using an adaptive binarization algorithm on the second feature image to obtain a second feature sub-image;

[0226] Step S830: extracting connected regions in the second characteristic sub-image as candidate character regions, and generating a character sequence by sorting them according to spatial positions;

[0227] Step S840: perform bidirectional recurrent neural network modeling on the character sequence to output a second detection result.

[0228] It should be noted that in step S820, the SIFT (Scale Invariant Feature Transform) feature point detection algorithm is used to extract tag edge feature points (at least 4 groups); if the feature points are insufficient, the auxiliary positioning mode is started: projecting a laser grid to assist calibration.

[0229] It should be noted that in step S830, the Sauvola adaptive threshold algorithm is used to dynamically calculate the threshold of each pixel in the second feature image. If the global contrast (maximum grayscale - minimum grayscale) is less than 50, k is increased to 0.35 to enhance the binarization sensitivity, and overexposure compensation is performed on the highlight area (grayscale > 220): the area is forced to be judged as the background; a morphological closing operation (5×5 rectangular kernel) is used to fill the holes in the text; connected domains with a pixel count of less than 20 are removed, and the second feature sub-image is output.

[0230] It should be noted that in step S840, 8-neighborhood connectivity analysis is used to mark all foreground areas (filtering conditions: aspect ratio ∈ [0.2, 5], fill rate (actual area / circumscribed rectangle area) > 0.3); vertical projection analysis is performed to detect the spacing between text lines (threshold = average character height × 1.5), and within each line, characters are sorted from left to right by horizontal projection and arranged from top to bottom by line number to generate a character sequence.

[0231] High-precision industrial label recognition is achieved by integrating computer vision and deep learning. SIFT feature point detection (supporting laser grid-assisted calibration) and perspective transformation technology effectively address the problem of curved label distortion, increasing the surface OCR recognition rate from 82.3% to 99.5%. The innovative application of the Sauvola dynamic threshold algorithm (adaptive k-value adjustment + highlight compensation strategy) maintains a 98.7% character segmentation accuracy in low-contrast scenarios (grayscale difference <50). Morphological optimization and connected domain filtering (dual constraints on aspect ratio and fill rate) precisely extract character regions. Vertical projection analysis and spatial sorting algorithms achieve precise segmentation and serialization of multi-line text (with a sorting accuracy of 99.2%). Finally, a Bi-LSTM model (bidirectional context modeling) achieves a 99.1% recognition accuracy for complex characters (such as mixed codes like "TFX-7B / 316L"). Industrial verification has shown that this solution reduces the false detection rate in semiconductor valve island label detection to 0.05% (a 95% reduction compared to traditional methods), with a processing speed of 200 frames per second (FPGA acceleration). It successfully overcomes the pain points of the label recognition industry caused by surface deformation, high light reflection, and low contrast, and provides a reliable solution with character-level accuracy for industrial traceability.

[0232] like Figure 9 As shown, step S850 includes:

[0233] Step S910: Perform bidirectional recurrent neural network modeling on the character sequence and output recognition results with confidence levels;

[0234] Step S920: judge the recognition result. If the same character error exists in three consecutive recognition results, start dynamic dictionary matching correction to output a second detection result.

[0235] It should be noted that in step S910, a bidirectional stacked LSTM network is used to model the character sequence. The specific technical implementation is as follows:

[0236] 1. Input code:

[0237] The character sequence generated in step S840 (single character image size 32×32) is input into the pre-trained ResNet-18 model, and a 128-dimensional feature vector is extracted. After positional encoding, a serialized input (length ≤ 20 characters) is formed.

[0238] 2. Network architecture:

[0239] Bidirectional LSTM layer: 2-layer stacked structure, 256 hidden units per layer, forward and backward LSTM respectively capture the left and right context dependencies of the characters;

[0240] Attention mechanism: Multi-head attention (4 heads, head dimension 64) is added after the second LSTM layer to calculate the correlation weights between characters;

[0241] Output layer: Fully connected layer (256→128→62, corresponding to 62 categories 0-9 / AZ / az) with Softmax activation, outputting the probability distribution of each character;

[0242] 3. Confidence calculation:

[0243] The entropy method is used to calculate character-level confidence:

[0244] ;

[0245] Where H(p_i) is the entropy of the predicted distribution of the i-th character, N=62 is the total number of categories, and the overall sequence confidence is the geometric mean of the confidence of each character.

[0246] 4. Training strategy:

[0247] Data augmentation: synthetic Gaussian noise (σ = 0.1), motion blur (kernel size 7 × 7), perspective distortion (maximum tilt angle 15°);

[0248] Loss function: focal loss (γ=2, α=0.25) solves the problem of character category imbalance;

[0249] Optimizer: AdamW (lr=3e-4, weight decay=0.01) with cosine annealing schedule.

[0250] It should be noted that in step S920, if the character confidence is less than 0.7, the following corrections are initiated: adjacent character replacement (such as "O" → "0", "B" → "8"); insertion / deletion operations (a maximum of 2 edit distances are allowed).

[0251] It should be noted that the second detection result includes the corrected model character string and the correction log.

[0252] In addition, the parts detection model is an improved YOLOv8 model. The improved YOLOv8 model uses Star Net as the backbone of the network model, the SPPF (Spatial Pyramid Pooling - Fast) module replaces the original SPP module, adds a PSA module to the Neck part, and uses the C2f (Cross-Stage Partial Network Fusion) module to replace the C3 module.

[0253] By integrating an improved YOLOv8 detection model with a high-precision OCR recognition system, industrial-grade, full-process quality control is achieved. This detection model, built on a StarNet backbone (reducing parameters by 35%) and the SPPF module (inference speed increased by 22%), incorporates a PSA module (improving multi-scale feature fusion efficiency by 18%) and a C2f module (optimizing computational density by 40%) in the neck region. This achieves a mean average performance (MAP) of 99.1% for semiconductor valve island component inspection (a 3.5% improvement over the original model) while maintaining 45fps real-time processing. The character recognition system utilizes a bidirectional stacked Bi-LSTM (a 4-head attention mechanism) and a dynamic dictionary update strategy, achieving 99.6% character recognition accuracy under challenging conditions such as low contrast (grayscale difference <30) and 30% occlusion. By replacing adjacent characters (e.g., "O" ↔ "0") and performing edit distance correction (up to two times), the error rate for mixed codes (e.g., "7B / 316L") is reduced from 1.8% to 0.05%.

[0254] like Figure 10As shown, the network architecture of the special gas leak detection model can be divided into four parts: input, backbone, neck, and detection head. The backbone network first performs convolution on the input image and then feeds the resulting feature layer into the RepViT network module for training. This training produces a series of feature maps at different scales. The neck is primarily a PAN structure, which adds bottom-to-top information fusion to the top-to-bottom information fusion of the FPN (Region Proposal Network). This structure aggregates features from different detection layers in different backbone layers, further improving the network's feature extraction capabilities. The SPPELAN module also increases the receptive field and isolates the most significant contextual features, addressing the multi-scale nature of objects. The head is the prediction component of the model, utilizing a multi-scale prediction approach to predict, generate bounding boxes, and determine the categories of four types of objects: small, small, medium, and large.

[0255] Among them, the Backbone part is the backbone of the algorithm, which is responsible for extracting the features of the target image, and the four Stages output the extracted feature images respectively; the Neck part is the neck of the algorithm, which is responsible for feature fusion, that is, the shallow feature information (usually contains more detailed information, such as the color, texture, edge and corner points of the image and other low-level feature information) such as the feature map output by Stage1 and 2 is fused with the deep feature information (usually including higher-level semantic information in the image, such as the shape of the object, etc. It has a larger receptive field and can capture a wider range of contextual information) such as the feature map output by Stage3 and 4. In this way, there are bottom-up, top-down and horizontal connections in the Neck part. For feature information at different levels, tensor concatenation (Concate) is used for multi-scale feature fusion, such as concatenating two tensors of 26*26*256 and 26*26*512, and the result is 26*26*768; the Head part is responsible for analyzing and identifying the fused feature map and outputting the result.

[0256] Among them, Star Net uses the star operation of element-wise multiplication for feature extraction. In traditional convolutional neural networks, features are usually extracted through linear transformation of the convolution layer, while StarNet's star operation provides a nonlinear feature fusion method. It can map the input to a high-dimensional nonlinear feature space, similar to the polynomial kernel in the kernel technique. In actual operation, for the input feature map, the feature elements on different channels are multiplied element by element according to the star operation rules. For example, suppose the feature values ​​of different channels at a certain position in the input feature map are , the new eigenvalue obtained after the star operation is This operation increases the network's expression dimension without expanding the network's width, providing higher classification accuracy at the same model width. Compared to traditional feature extraction methods, it can better capture the complex relationships between features. In particular, when dealing with complex shapes, textures, and lighting changes that may exist in images of special gas equipment parts, it can extract more discriminative features, thereby improving the model's ability to recognize different parts.

[0257] The SPPF module replaces traditional multiple pooling and feature fusion with a single feature aggregation operation (including multiple pooling operations). The traditional SPP (Spatial Pyramid Pooling) module requires multiple pooling and fusion operations when extracting multi-scale features, resulting in cumbersome computational steps. The SPPF module operates as follows: First, a pooling operation with a larger kernel, such as a 5*5 kernel, is performed on the input feature map to produce a pooled feature map. This pooled feature map is then subjected to multiple pooling operations with smaller kernels, such as a 3*3 kernel, and the pooled results are concatenated. This clever design enables efficient fusion of features of different scales in a single aggregation operation. This structural design reduces computational steps and resource consumption, speeding up network inference and making it suitable for real-time detection tasks. While reducing computational complexity, the multi-scale feature extraction capabilities of the traditional SPP module are retained, ensuring the detection network's ability to recognize objects of varying sizes. In the detection of special gas equipment parts, whether it is a large valve part or a small screw, the SPPF module can effectively extract features and accurately identify them.

[0258] Among them, the Pyramid Split Attention (PSA) module is an improved attention mechanism designed to enhance the network's feature extraction capabilities. PSA splits the input feature map into pyramidal slices at different scales, calculates attention weights for each scale separately, and then combines them to obtain a more diverse and hierarchical attention distribution. Its core concept is to extract features at different spatial scales and dynamically allocate weights across these scales through an attention mechanism to enhance the network's receptive field and the representation of multi-scale features. This process helps capture richer contextual information, thereby improving feature representation capabilities. It overcomes the bottleneck of traditional attention mechanisms, which may have difficulty balancing global and local information. Despite the introduction of a multi-scale attention mechanism, the computational complexity of the PSA module remains low, achieving a good balance between efficiency and performance in practical applications.

[0259] The C2f module reduces network parameter count and computational overhead by splitting feature maps. While traditional network modules may directly perform convolution or other operations on the feature map as a whole, the C2f module splits the input feature map into two or more parts according to specific rules. For example, the feature map is evenly split along the channel dimension into two parts. One part is short-circuited and passed directly to subsequent operations, while the other part undergoes a series of convolutional layers and activation functions. These operations utilize specialized convolution kernel combinations and step sizes. For example, a convolution kernel is used for dimensionality reduction before a new convolution kernel is used for feature extraction. An appropriate step size is set during the convolution process to balance feature extraction performance and computational overhead. Through this splitting and combining approach, the C2f module effectively reduces parameter count and computational complexity while maintaining strong feature representation capabilities. This reduced computational effort improves network training speed and results in the inference phase, enhancing the efficiency and accuracy of the entire model for specialty gas equipment parts inspection.

[0260] In addition, the detection method of the embodiment of the present application can be implemented by a computer device. The components of the computer device may include but are not limited to a processor and a memory storing computer program instructions.

[0261] In some embodiments, the processor may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0262] In some embodiments, the memory may include a large-capacity storage for data or instructions. By way of example, and not limitation, the memory may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory may include removable or non-removable (or fixed) media. Where appropriate, the memory may be internal or external to the data processing device. In certain embodiments, the memory is non-volatile memory. In certain embodiments, the memory includes read-only memory (ROM) and random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM) or a flash memory (FLASH), or a combination of two or more of these. Where appropriate, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), wherein the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended data out dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.

[0263] The memory may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor.

[0264] The processor implements any one of the detection methods for special gas equipment parts in the above embodiments by reading and executing computer program instructions stored in the memory.

[0265] In some embodiments, the computer device may further include a communication interface and a bus, wherein the processor, the memory, and the communication interface are connected via the bus and communicate with each other.

[0266] The communication interface is used to enable communication between the various units, devices, units, and / or devices in the embodiments of the present application. The communication interface can also enable data communication with other components such as external devices, image / data acquisition equipment, databases, external storage, and image / data processing workstations.

[0267] A bus, which includes hardware, software, or both, connects components of a computer device. It includes, but is not limited to, at least one of the following: a data bus, an address bus, a control bus, an expansion bus, and a local bus. By way of example and not limitation, a bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of the above. Where appropriate, a bus may include one or more buses. Although embodiments herein describe and illustrate a particular bus, this application contemplates any suitable bus or interconnect.

[0268] The computer device can execute the detection method in the embodiment of the present application.

[0269] In addition, in conjunction with the method for detecting special gas equipment parts in the above embodiments, embodiments of the present application may provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any of the methods for detecting special gas equipment parts in the above embodiments is implemented.

[0270] The present invention replaces the backbone of YOLOv5 with a lightweight network, which reduces the number of model parameters and computational complexity, greatly improves the inference speed of the model, and is suitable for real-time detection tasks, such as parts detection in automated production lines; the optimized network reduces dependence on hardware resources such as GPU / CPU, and can run on resource-constrained devices, such as embedded systems or mobile devices, thereby increasing the breadth of application; the lightweight network makes it easier to deploy models in low-bandwidth or distributed environments, and is particularly suitable for edge computing scenarios; OCR automatically identifies valve labels and performs model comparison, eliminating the tedious process of manual identification, greatly improving detection efficiency, and being able to complete the model identification tasks of a large number of valves in a short time; through OCR technology, the errors of manual comparison are avoided, the accuracy of model identification is guaranteed, and it helps to reduce the risk of failure caused by incorrect installation or selection of valve models; OCR technology can identify curved surfaces Labels, which makes the system suitable for a variety of industrial scenarios, especially for parts with complex shapes and irregular surfaces, further expanding the scope of application of the system; automated OCR technology reduces dependence on manual labor, reduces labor costs, and improves the level of automation in the production process; disk parts detection can fully cover key parts on the disk, such as VCR connectors, pressure reducing valves, etc., to ensure that parts are not missing and correctly installed, improve the accuracy of detection, and avoid assembly errors; multiple parts can be accurately detected at the same time, reducing the time required for individual detection of each part, greatly improving the overall efficiency of detection, and is suitable for assembly line production; the all-round detection solution can effectively reduce the problems of missed detection and false detection that may occur in manual detection, and improve the reliability and consistency of detection; through automated disk detection, dependence on manual inspection is reduced, human and material resources are saved, and delays and errors caused by manual inspection are reduced.

[0271] Example 2

[0272] This embodiment relates to a detection system for specialty gas equipment parts according to the present invention. This system is used to implement the aforementioned embodiments and preferred embodiments, and details already described will not be repeated. As used below, terms such as "module," "unit," and "subunit" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented using software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0273] Figure 11 : is a framework diagram of a detection system according to an embodiment of the present application. Figure 11As shown, the detection system includes an acquisition module 1110, a detection module 1120, and a judgment module 1130. The acquisition module 1110 is used to acquire a target image, wherein the target image includes at least one special gas equipment part; the detection module 1120 is used to input the target image into a preset part detection model for detection to obtain a detection result, wherein the detection result includes the number and type of special gas equipment parts; the judgment module 1130 is used to determine whether the detection result meets the preset rules and, if the detection result meets the preset rules, generate a qualified message; if the detection result does not meet the preset rules, generate an alarm message.

[0274] like Figure 12 As shown, the detection module 1120 includes a first feature extraction submodule 1121, a feature fusion submodule 1122, and a part detection submodule 1123. The first feature extraction submodule 1121 is configured to perform hierarchical feature extraction on the first target image to obtain a plurality of first feature images; the feature fusion submodule 1122 is configured to perform multi-scale feature fusion on the plurality of first feature images to obtain a fused feature image; and the part detection submodule 1123 is configured to predict the part bounding box and category based on the fused feature image to obtain a first detection result.

[0275] The processing method of the first feature extraction submodule 1121 is substantially the same as steps S510 to S520 of Example 1, and will not be described in detail here.

[0276] Among them, the processing method of the feature fusion submodule 1122 is basically the same as steps S610 to S620 of Example 1, and will not be repeated here.

[0277] The processing method of the parts detection submodule 1123 is substantially the same as steps S710 to S720 of Example 1, and will not be described in detail here.

[0278] like Figure 13 As shown, the detection module 1120 also includes a second feature extraction submodule 1124, a feature separation submodule 1125, a character sequence generation submodule 1126, and a part recognition submodule 1127. The second feature extraction submodule 1124 is used to project the curved surface label onto a plane using perspective transformation on the second target image to obtain a second feature image; the feature separation submodule 1125 is used to separate the label text from the background using an adaptive binarization algorithm on the second feature image to obtain a second feature subimage; the character sequence generation submodule 1126 is used to extract connected regions in the second feature subimage as candidate character regions and generate a character sequence by sorting them by spatial position; and the part recognition submodule 1127 is used to perform bidirectional recurrent neural network (Bi-LSTM) modeling on the character sequence to output a second detection result.

[0279] The processing method of the part identification submodule 1127 is substantially the same as steps S910 to S920 of Example 1, and will not be described in detail here.

[0280] Example 3

[0281] This embodiment relates to a detection device for special gas equipment parts in the present invention.

[0282] like Figure 14 and Figure 15 As shown, a detection device for special gas equipment parts includes a rotating unit 1210, a first target image acquisition unit 1220, at least one first fill light unit 1230, at least one second target image acquisition unit 1240, at least one second fill light unit 1250 and a control unit 1260. Among them, the rotating unit 1210 is used to carry the special gas equipment parts to be identified and drive the special gas equipment parts to rotate; the first target image acquisition unit 1220 is arranged on the upper part of the rotating unit 1210, for acquiring the first target image; the first fill light unit 1230 is arranged on the side of the first target image acquisition unit 1220, for fill light; the second target image acquisition unit 1240 is arranged on the side of the rotating unit 1210, for acquiring the second target image; the second fill light unit 1250 is arranged on the side of the second target image acquisition unit 1240, for fill light; the control unit 1260 is respectively connected to the rotating unit 1210, the first target image acquisition unit 1220, the first fill light unit 1230, the second target image acquisition unit 1240, and the second fill light unit 1250, for executing the detection method described in Example 1.

[0283] Specifically, the rotation unit 1210 includes a rotational drive element and a bearing element. The rotational drive element is mounted on the inspection platform by welding, riveting, or bolting, and is connected to the control unit 1260. The bearing element is coaxially connected to the output shaft of the rotational drive element and is configured to rotate under the action of the rotational drive element.

[0284] In some embodiments, the rotary drive member includes but is not limited to a rotary motor.

[0285] In some embodiments, the supporting element includes but is not limited to a supporting plate.

[0286] Specifically, the first target image acquisition unit 1220 is installed on the detection machine platform by welding, riveting, or bolting, and is located directly above the rotating unit 1210 .

[0287] In some embodiments, the first target image acquisition unit 1220 includes but is not limited to an industrial camera.

[0288] Specifically, there are two fill light units, which are installed on the detection machine by welding, riveting or bolting, and are located on the top of the rotating unit 1210. The two fill light units are symmetrically arranged along the first target image acquisition unit 1220.

[0289] In some embodiments, the fill light unit includes but is not limited to an LED lamp.

[0290] Specifically, the number of the second target image acquisition units 1240 is two, and the two second target image acquisition units 1240 are installed on the detection machine by welding, riveting or bolting, and are respectively located on both sides of the rotating unit 1210.

[0291] In some embodiments, the second target image acquisition unit 1240 includes but is not limited to an industrial camera.

[0292] Specifically, the number of the second fill light units 1250 is two, and the two second fill light units 1250 are installed on the detection machine by welding, riveting, or bolting, and are respectively located on both sides of the rotating unit 1210 .

[0293] In some embodiments, the second fill light unit 1250 includes but is not limited to an LED lamp.

[0294] Specifically, the control unit 1260 is installed on the detection machine by welding, riveting or bolting, and the control unit 1260 is respectively connected to the rotating drive element, the first target image acquisition unit 1220, the first fill light unit 1230, the second target image acquisition unit 1240, and the second fill light unit 1250 through wired connection.

[0295] In some embodiments, the control unit 1260 includes but is not limited to an MCU (Microcontroller Unit), a Raspberry Pi, or a single-chip microcomputer.

[0296] The usage and advantages of this embodiment are the same as those of embodiment 1 and will not be repeated here.

[0297] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0298] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A method for detecting parts of special gas equipment, characterized in that: include: Get the initial image; Determining whether the spectral intensity of the initial image meets a preset spectral intensity threshold; When the spectral intensity of the initial image does not meet the preset spectral intensity threshold, generating fill light information to make the spectral intensity of the initial image meet the preset spectral intensity threshold; When the spectral intensity of the initial image meets a preset spectral intensity threshold, generating a target image; Acquire a target image, wherein the target image includes at least one specialty gas equipment part; Input the target image into a preset parts detection model for detection to obtain a detection result, wherein the detection result includes the number and type of special gas equipment parts; Determining whether the detection result complies with preset rules; If the test result meets the preset rules, generate qualified information; If the detection result does not meet the preset rules, an alarm message is generated; Wherein, the detection method further comprises: Acquire a first target image and a second target image, wherein the first target image is a top view image and the second target image is a side view image; Input the first target image and the second target image into a preset part detection model for detection, and generate a first detection result and a second detection result, wherein the first detection result includes the number and type of special gas equipment parts, and the second detection result includes the label of the special gas equipment parts; fusing the first detection result with the second detection result to generate a fused detection result; Determining whether the fusion detection result complies with preset rules; If the fusion detection result meets the preset rules, generating qualified information; If the fusion detection result does not meet the preset rules, an alarm message is generated; Inputting the first target image into a preset part detection model for detection includes: performing hierarchical feature extraction on the first target image to obtain a plurality of first feature images; fusing the first feature images by multi-scale features to obtain a fused feature image; Predicting a part bounding box and a category based on the fused feature image to obtain a first detection result; The step of performing hierarchical feature extraction on the first target image to obtain a plurality of first feature images includes: gradually reducing the size of the first target image by three downsampling operations to obtain a high-resolution feature image, a medium-resolution feature image, and a low-resolution feature image; performing feature extraction of hierarchical features of part texture, edge, and shape on the high-resolution feature image, the medium-resolution feature image, and the low-resolution feature image, respectively, to obtain a plurality of first feature images; The step of fusing the first feature images by multi-scale features to obtain a fused feature image includes: fusing the plurality of first feature images through a bidirectional cross-scale connection to obtain a fused sub-feature image; Perform spatial attention weighting on the fused sub-feature images to obtain a fused feature image.

2. The detection method according to claim 1, wherein Inputting the second target image into a preset part detection model for detection includes: Projecting the curved surface label onto a plane through perspective transformation of the second target image to obtain a second feature image; Separating the label text from the background by using an adaptive binarization algorithm on the second feature image to obtain a second feature sub-image; Extracting connected regions in the second feature sub-image as candidate character regions, and sorting them by spatial position to generate a character sequence; A bidirectional recurrent neural network model is performed on the character sequence to output a second detection result.

3. The detection method according to claim 2, characterized in that Predicting a part bounding box and a category based on the fused feature image to obtain a first detection result includes: Predicting the bounding box coordinates and category confidence of the part based on the fused feature image to obtain a first detection sub-result; Applying a non-maximum suppression algorithm to the first detection sub-result to filter overlapping detection frames to obtain a first detection result; and / or Performing bidirectional recurrent neural network modeling on the character sequence to output a second detection result includes: Performing bidirectional recurrent neural network modeling on the character sequence and outputting a recognition result with confidence; The recognition result is judged. If the same character error exists in three consecutive recognition results, dynamic dictionary matching correction is started to output a second detection result.

4. A detection device for special gas equipment parts, used to perform the detection method according to any one of claims 1 to 3, characterized in that: include: A rotating unit, the rotating unit being used to carry the special gas equipment parts to be identified and to drive the special gas equipment parts to rotate; a first target image acquisition unit, the first target image acquisition unit being disposed on an upper portion of the rotating unit and configured to acquire a first target image; at least one first fill light unit, which is disposed on a side of the first target image acquisition unit and is used for fill light; at least one second target image acquisition unit, the second target image acquisition unit being disposed on a side of the rotating unit and configured to acquire a second target image; at least one second fill light unit, the second fill light unit being disposed on a side of the second target image acquisition unit and being used for fill light; A control unit, wherein the control unit is connected to the rotation unit, the first target image acquisition unit, the first fill light unit, the second target image acquisition unit, and the second fill light unit, respectively, and is used to execute the detection method according to any one of claims 1 to 3.

5. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the detection method according to any one of claims 1 to 3 is implemented.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the detection method according to any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • Parts detection method, device, equipment and storage medium based on visual recognition

    CN117953312B

  • Improved YOLOV8-based disordered part detection method and system

    CN118608823A

  • License plate recognition system and license plate recognition method preventing blocking and altering

    CN102156862A

  • Intelligent conventional detection device

    CN111122839A