Machine vision-based special equipment part fine granularity identification method and system
Through multi-spectral image acquisition and dynamic occlusion compensation technology, combined with RGB and infrared cameras, high-precision and real-time recognition of special equipment components are achieved, and the problems of poor environmental adaptability and insufficient real-time performance are solved. It supports the rapid expansion of new component types and reduces hardware costs.
Patent Information
- Application Number
- CN202510597916.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-08
AI Technical Summary
In the inspection and detection of special equipment, the existing technology has problems such as poor environmental adaptability, weak generalization ability of traditional algorithms, large model parameters, insufficient real-time real-time occlusion scene recognition accuracy, and high fine-grained labeling cost, making it difficult to achieve high-precision and real-time component recognition.
Multi-spectral image acquisition, component area positioning and occlusion compensation, multi-modal feature fusion and enhancement, fine-grained dynamic classification model reasoning, confidence calibration and incremental learning methods are adopted, combined with RGB and infrared cameras, and lightweight ViT architecture and incremental learning mechanism are used to achieve dynamic occlusion compensation and feature fusion, supporting new component types online.
It realizes high-precision component recognition in complex environments, with a 40% recognition rate and a 3-4-fold increase in real-time performance, supports rapid expansion of new component types, reduces hardware costs by 30%, and meets the diversified needs of special equipment.
Smart Images

Figure CN120451742A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of special equipment identification, and in particular to a method and system for fine-grained identification of special equipment components based on machine vision. Background Art
[0002] In the inspection and testing of special equipment (such as elevators, cranes, and pressure vessels), accurate component identification is a key step in achieving intelligent and automated inspection. Traditional methods mainly rely on manual visual inspection or machine vision technology based on a single sensor, which has the following significant problems: 1. Poor environmental adaptability: Low light and oil interference: Low light and oil coverage are common in inspection and testing scenarios, making it difficult for traditional RGB cameras to clearly capture component details. While infrared imaging can penetrate some obstructions, its single spectral data lacks texture information, resulting in a high false detection rate. Overexposure in strong light: In outdoor or highly reflective environments, traditional HDR algorithms have limited dynamic range, resulting in overexposure or loss of detail in key areas.
[0003] 2. Traditional algorithms have poor generalization capabilities: Rule-based feature extraction methods (such as edge detection and template matching) are sensitive to changes in component shape and size and are difficult to adapt to different equipment models (such as the differences between OTIS and Mitsubishi elevator traction motors).
[0004] 3. Existing models (such as Faster R-CNN and Mask R-CNN) have large parameters and inference latency exceeding 100ms on edge devices (such as AR glasses), failing to meet real-time requirements. They also perform poorly in occluded and partially visible scenes: When a component is obscured by more than 20%, model recognition accuracy plummets.
[0005] 4. Difficulty in fine-grained labeling: Minor differences in similar components (such as the meshing structure of door locks from different brands) require a large amount of labeled data, resulting in high manual labeling costs. New component types require retraining the entire model, which can take weeks and makes it difficult to adapt to the needs of rapid equipment upgrades. Summary of the Invention
[0006] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method and system for fine-grained recognition of special equipment components based on machine vision.
[0007] The object of the present invention is achieved through the following technical solutions: In a first aspect, the present invention provides: a method for fine-grained recognition of special equipment components based on machine vision, comprising the following steps: During the multispectral image acquisition phase, the acquisition equipment is configured, different acquisition strategies are adopted for different lighting environments, and motion jitter is compensated in real time to reduce image offset errors, thereby obtaining multispectral images and component spatial position data. In the component area positioning and occlusion compensation stage, component area positioning, dynamic occlusion compensation, and key point annotation are performed to obtain the component feature map after background removal, occlusion mask, and key point coordinates; In the multimodal feature fusion and enhancement stage, multispectral feature fusion, feature enhancement, and occlusion robustness training are performed to obtain the fused feature vector and component structure heat map; During the fine-grained dynamic classification model inference phase, the model is deployed in a lightweight manner and inference optimization is performed in real time. Feature vectors are input into the dynamic classification network, and a dynamic attention mechanism is introduced to automatically switch attention heads based on component types to obtain the component category probability distribution. During the confidence calibration and incremental learning update phase, the confidence levels are graded and different processing strategies are adopted for different levels. The incremental learning mechanism and anomaly detection module are introduced for calibration, and finally the calibrated component name, confidence level and incremental learning update status are obtained.
[0008] Preferably, the motion jitter is compensated in real time by an IMU sensor; the acquisition device includes an RGB camera and an infrared camera. The HDR mode of the RGB camera is used to capture component details in a strong light environment, and the infrared camera is used to penetrate oil and dust to enhance the component outline in a weak light environment.
[0009] Preferably, the component area positioning and occlusion compensation stage further includes the following steps: The multispectral image is input into the pre-trained YOLOv8n model to output the rough positioning bounding box of the component. For occluded scenes, the main area of the component is extracted through saliency detection, and the background interference is suppressed to obtain the component feature map. When a component is occluded, a pre-trained 3D model of the component is used to generate a virtual outline of the occluded area. The LSTM network is used to analyze the previous and next frame data and predict the structural integrity of the occluded component to obtain an occlusion mask. The key functional points of the components are marked to assist in fine-grained classification to obtain the key point coordinates.
[0010] Preferably, during the training phase, occlusion areas are randomly added to the complete component image to simulate a real scene; a contrast loss function is used to enhance the consistency of features of similar components; and the multimodal feature fusion and enhancement phase further includes the following steps: Input the RGB image and infrared pattern into the two-branch convolutional network and output a 4-channel feature map; dynamically assign spectral weights through SE Block; EfficientNet-B4 is used to extract the overall structure of the component to obtain global features; high-resolution sub-images are cropped based on the coordinates of key points, and detailed textures are extracted through ResNet-18 to obtain local features; global features and local features are spliced together to obtain feature vectors.
[0011] Preferably, the fine-grained dynamic classification model reasoning stage further includes the following steps: Using the knowledge distillation method, ViT-Base is used as the teacher model and MobileNetV3 is used as the student model. The model size is compressed to 10MB, and TensorRT is quantized to INT8, achieving lightweight deployment and improving inference speed. Use Qualcomm XR2 chip Hexagon DSP to accelerate matrix operations and achieve inference optimization.
[0012] Preferably, the anomaly detection module calculates feature distribution differences based on Mahalanobis distance, identifies unknown component types, triggers AR interface warnings, records and transmits them to the cloud for processing; the confidence calibration and incremental learning update phase also includes the following steps: For the first confidence level, the recognition result is directly output and the label is displayed on the AR interface; for the second confidence level, multi-model voting is triggered to improve the consistency of the recognition results; for the third confidence level, manual review is prompted, and the marked data is returned to the training set; For newly added component types, the classification head is updated through comparative learning, multi-device collaborative training, and encrypted uploaded features.
[0013] Preferably, the backbone network of the dynamic classification network is Vision Transformer.
[0014] A second aspect of the present invention provides: a machine vision-based fine-grained recognition system for special equipment components, used to implement any of the above-mentioned machine vision-based fine-grained recognition methods for special equipment components, comprising: The multispectral image acquisition module is used to configure the acquisition equipment, adopt different acquisition strategies for different lighting environments, and compensate for motion jitter in real time to reduce image offset errors, thereby obtaining multispectral images and component spatial position data; The component area positioning and occlusion compensation module is used to perform component area positioning, dynamic occlusion compensation and key point annotation to obtain the component feature map after removing the background, occlusion mask and key point coordinates; Multimodal feature fusion and enhancement module, used for multispectral feature fusion, feature enhancement and occlusion robustness training, to obtain the fused feature vector and component structure heat map; The fine-grained dynamic classification model inference module is used for lightweight model deployment and real-time inference optimization. It inputs feature vectors into the dynamic classification network and introduces a dynamic attention mechanism to automatically switch attention heads based on component types to obtain component category probability distribution. The confidence calibration and incremental learning update module is used to grade confidence and adopt different processing strategies for different levels. The incremental learning mechanism and anomaly detection module are introduced for calibration, and finally the calibrated component name, confidence level and incremental learning update status are obtained.
[0015] The third aspect of the present invention provides: a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are loaded and executed by a processor, any of the above-mentioned fine-grained recognition methods for special equipment components based on machine vision is implemented.
[0016] A fourth aspect of the present invention provides: a computer program product comprising instructions, which, when running on a terminal, enables the terminal to execute any of the above-mentioned methods for fine-grained identification of special equipment components based on machine vision.
[0017] The beneficial effects of the present invention are: 1) Through multi-spectral dynamic fusion, occlusion compensation algorithm and lightweight ViT architecture, high-precision component recognition in complex inspection and testing scenarios is achieved.
[0018] 2) High Precision and Strong Robustness, Multispectral Fusion: Combining RGB and infrared imaging, it penetrates complex environments such as oil pollution and low light, achieving recognition rates exceeding 85% in low light (<50 lux), and a 40% improvement in oil pollution scenarios. Dynamic Occlusion Compensation: Through GAN-generated virtual contours and LSTM time series prediction, it can maintain a recognition accuracy of ≥95% when the component is 80% visible, solving the occlusion problem.
[0019] 3) Real-time and lightweight, edge-end optimization: Using MobileNetV3 and TensorRT INT8 quantization, the model size is compressed to 10MB, and the single-frame inference time is ≤25ms, which is compatible with edge devices such as AR glasses; Hexagon DSP acceleration: Preprocessing and inference speed are increased by 3-4 times, meeting the real-time requirements of inspection and testing scenarios.
[0020] 4) Flexible scalability and privacy security. Incremental learning: supports online addition of new component types (expanding 50+ subcategories per month) without retraining the entire model, with a classification error rate of ≤2%. Federated learning: Through differential privacy (ε=0.3) and feature encryption, multi-device collaborative training is achieved, reducing the risk of data privacy leakage by 70%.
[0021] 5) Fine-grained classification capability: Based on the adaptive adjustment of ViT's attention head, it can distinguish subtle differences between similar components (such as the meshing structure of elevator door locks of different brands), improving fine-grained classification accuracy by 15%.
[0022] 6) Low cost and easy deployment, hardware collaborative design: No need to rely on 3D sensors or laser assistance, only RGB + infrared cameras are required, reducing hardware costs by 30%; cross-platform compatibility: Supports mainstream AR glasses and industrial equipment, with a low deployment threshold.
[0023] 7) Multi-scenario coverage, adapting to full-scale recognition from 10cm-level tiny components (such as screws) to 10m-level large structures (such as storage tanks), with a defect detection accuracy of ≥90%, meeting the diverse needs of special equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 This is a flow chart of the fine-grained recognition method for special equipment components based on machine vision. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work shall fall within the scope of protection of the present invention.
[0026] See Figure 1 The first aspect of the present invention provides: a method for fine-grained recognition of special equipment components based on machine vision, comprising the following steps: During the multispectral image acquisition phase, the acquisition equipment is configured, different acquisition strategies are adopted for different lighting environments, and motion jitter is compensated in real time to reduce image offset errors, thereby obtaining multispectral images and component spatial position data. In the component area positioning and occlusion compensation stage, component area positioning, dynamic occlusion compensation, and key point annotation are performed to obtain the component feature map after background removal, occlusion mask, and key point coordinates; In the multimodal feature fusion and enhancement stage, multispectral feature fusion, feature enhancement, and occlusion robustness training are performed to obtain the fused feature vector and component structure heat map; During the fine-grained dynamic classification model inference phase, the model is deployed in a lightweight manner and inference optimization is performed in real time. Feature vectors are input into the dynamic classification network, and a dynamic attention mechanism is introduced to automatically switch attention heads based on component types to obtain the component category probability distribution. During the confidence calibration and incremental learning update phase, the confidence levels are graded and different processing strategies are adopted for different levels. The incremental learning mechanism and anomaly detection module are introduced for calibration, and finally the calibrated component name, confidence level and incremental learning update status are obtained.
[0027] In this embodiment, RGB and infrared cameras are combined, and a cross-modal attention mechanism dynamically assigns spectral weights (for example, increasing the infrared weight to 70% in oil-stained scenes) to enhance the recognizability of component textures and contours. A Generative Adversarial Network (GAN) is used to generate virtual outlines of occluded areas (for example, generating the complete structure of an elevator door when it is 20% obscured). LSTM time-series prediction analyzes previous and next frame data to compensate for dynamic occlusion interference, ensuring a recognition rate of ≥95% under 80% visibility conditions.
[0028] Based on the Vision Transformer (ViT) architecture, attention heads are automatically adjusted based on component type (e.g., focusing on microtexture for "wire rope" and overall appearance for "control cabinet"), improving fine-grained classification accuracy (error rate ≤ 2%). ViT is compressed into MobileNetV3 using knowledge distillation technology. Combined with TensorRT INT8 quantization and Hexagon DSP acceleration, single-frame inference time is ≤ 25ms, and the model size is compressed to 10MB, making it suitable for real-time execution on edge devices.
[0029] Online incremental expansion: New component types (such as new traction machines) are fine-tuned through contrastive learning (SupCon), supporting monthly expansion of 50+ subcategories without retraining the entire model. Privacy protection mechanism: Federated learning uses differential privacy (ε=0.3) to encrypt feature data and aggregate classification head parameters in the cloud, balancing model iteration efficiency and data security.
[0030] In some embodiments, motion jitter is compensated in real time by an IMU sensor; the acquisition device includes an RGB camera and an infrared camera. The HDR mode of the RGB camera is used to capture component details in a strong light environment, and the infrared camera is used to penetrate oil and dust to enhance the component outline in a weak light environment.
[0031] In this embodiment, the hardware configuration includes: an RGB camera with a 120° wide-angle lens, 4K resolution (3840×2160), and HDR support, ensuring the capture of component details (such as wire rope texture) even in bright light. An infrared camera with an 850nm wavelength penetrates oil and dust, enhancing component outlines in low-light environments (such as the bumper in an elevator pit). A dynamic data collection strategy ensures that the user receives real-time shooting guidance (e.g., "Aim to the left of the traction machine") through the AR glasses interface, ensuring ≥80% visibility of the target component. An IMU sensor (accelerometer + gyroscope) compensates for motion jitter, maintaining an image offset error of ≤3 pixels. Output: RGB + infrared dual-channel image streams (resolution 3840×2160) and component spatial position data (JSON format).
[0032] In some embodiments, the component area positioning and occlusion compensation stage further includes the following steps: The multispectral image is input into the pre-trained YOLOv8n model to output the rough positioning bounding box of the component. For occluded scenes, the main area of the component is extracted through saliency detection, and the background interference is suppressed to obtain the component feature map. When a component is occluded, a pre-trained 3D model of the component is used to generate a virtual outline of the occluded area. The LSTM network is used to analyze the previous and next frame data and predict the structural integrity of the occluded component to obtain an occlusion mask. The key functional points of the components are marked to assist in fine-grained classification to obtain the key point coordinates.
[0033] In this embodiment, component region localization uses a pre-trained YOLOv8n model as input for multispectral imagery and outputs a coarse bounding box for the component (e.g., "traction machine: x1, y1, x2, y2"). For occlusion scenarios, saliency detection (U²-Net) is used to extract the component's main area and suppress background interference. Dynamic occlusion compensation uses a GAN virtual generation algorithm. If a component is occluded (e.g., a 20% occlusion of a car door), a pre-trained 3D model of the component is used to generate a virtual outline of the occluded area. Time series prediction uses an LSTM network to analyze previous and subsequent frame data and predict the structural integrity of the occluded component. Keypoint annotation uses an annotation algorithm to label key functional points of components (e.g., the traction sheave of a traction machine or the rope end fixture of a wire rope) to assist in fine-grained classification. Outputs include a background-removed component feature map (1024×1024 resolution), an occlusion mask (binary mask), and keypoint coordinates.
[0034] In some embodiments, during the training phase, occluded areas are randomly added to the complete component image to simulate a real scene; a contrast loss function is used to enhance the consistency of features of similar components; and the multimodal feature fusion and enhancement phase further includes the following steps: Input the RGB image and infrared pattern into the two-branch convolutional network and output a 4-channel feature map; dynamically assign spectral weights through SE Block; EfficientNet-B4 is used to extract the overall structure of the component to obtain global features; high-resolution sub-images are cropped based on the coordinates of key points, and detailed textures are extracted through ResNet-18 to obtain local features; global features and local features are spliced together to obtain feature vectors.
[0035] In this example, multispectral feature fusion is performed as follows: The early fusion network feeds RGB and infrared images into a two-branch convolutional network, outputting a four-channel feature map (R / G / B / IR). Cross-modal attention dynamically assigns spectral weights using SE blocks, for example, increasing the infrared channel weight to 70% in oil-contaminated scenarios. Local-global feature enhancement: Global features are extracted using EfficientNet-B4 (e.g., the control cabinet's shape). Local features are extracted using keypoints to extract high-resolution sub-images (e.g., floor door locks) and detailed textures using ResNet-18. Feature concatenation is performed by concatenating global features (1024 dimensions) and local features (512 dimensions) to reduce their dimensionality to 256. Occlusion robustness training is performed during training by randomly adding occlusions (20% of the area) to the complete component image to simulate real-world scenarios. Contrastive loss is used to enhance the consistency of features across similar components. Output: A fused feature vector (256 dimensions) and a component structure heatmap.
[0036] In some embodiments, the fine-grained dynamic classification model inference stage further includes the following steps: Using the knowledge distillation method, ViT-Base is used as the teacher model and MobileNetV3 is used as the student model. The model size is compressed to 10MB, and TensorRT is quantized to INT8, achieving lightweight deployment and improving inference speed. Use Qualcomm XR2 chip Hexagon DSP to accelerate matrix operations and achieve inference optimization.
[0037] In this embodiment, the dynamic classification network uses the Vision Transformer (ViT-Base) backbone network as input, which outputs a probability distribution of component categories. A dynamic attention mechanism automatically switches attention heads based on component type, focusing on texture features for "wire rope" and contours for "control cabinet." Lightweight deployment utilizes knowledge distillation technology, using ViT-Base as the teacher model and MobileNetV3 as the student model, reducing the model size to 10MB. TensorRT is quantized to INT8 for improved inference speed. Real-time inference optimization utilizes the Qualcomm XR2 chip's Hexagon DSP to accelerate matrix operations, achieving single-frame inference time of ≤25ms. Offline operation is supported, eliminating cloud-based dependencies. Output: Component name (e.g., "traction machine," "car door") and confidence score (0-1).
[0038] In some embodiments, the anomaly detection module calculates feature distribution differences based on Mahalanobis distance, identifies unknown component types, triggers an AR interface warning, and records and transmits the information to the cloud for processing. The confidence calibration and incremental learning update phase also includes the following steps: For the first confidence level, the recognition result is directly output and the label is displayed on the AR interface; for the second confidence level, multi-model voting is triggered to improve the consistency of the recognition results; for the third confidence level, manual review is prompted, and the marked data is returned to the training set; For newly added component types, the classification head is updated through comparative learning, multi-device collaborative training, and encrypted uploaded features.
[0039] In this embodiment, the confidence grading strategy is as follows: High confidence (≥90%): direct output of the result, with a green label displayed in the AR interface; Medium confidence (70%-90%): triggering multi-model voting (ViT + EfficientNet) to improve result consistency; Low confidence (<70%): prompting manual review, and the annotated data is returned to the training set. Incremental learning mechanism: Online fine-tuning: New component types (such as new traction machines) are updated through contrastive learning (SupCon); Federated learning support: Multi-device collaborative training, encrypted uploaded features (not raw images) to protect data privacy. Anomaly detection module: Calculating feature distribution differences based on Mahalanobis distance to identify unknown component types; triggering a yellow warning in the AR interface, and recording to the cloud for processing. Output: calibrated component name, confidence level, and incremental learning update status.
[0040] In some embodiments, the backbone network of the dynamic classification network is Vision Transformer.
[0041] A second aspect of the present invention provides: a machine vision-based fine-grained recognition system for special equipment components, used to implement any of the above-mentioned machine vision-based fine-grained recognition methods for special equipment components, comprising: The multispectral image acquisition module is used to configure the acquisition equipment, adopt different acquisition strategies for different lighting environments, and compensate for motion jitter in real time to reduce image offset errors, thereby obtaining multispectral images and component spatial position data; The component area positioning and occlusion compensation module is used to perform component area positioning, dynamic occlusion compensation and key point annotation to obtain the component feature map after removing the background, occlusion mask and key point coordinates; Multimodal feature fusion and enhancement module, used for multispectral feature fusion, feature enhancement and occlusion robustness training, to obtain the fused feature vector and component structure heat map; The fine-grained dynamic classification model inference module is used for lightweight model deployment and real-time inference optimization. It inputs feature vectors into the dynamic classification network and introduces a dynamic attention mechanism to automatically switch attention heads based on component types to obtain component category probability distribution. The confidence calibration and incremental learning update module is used to grade confidence and adopt different processing strategies for different levels. The incremental learning mechanism and anomaly detection module are introduced for calibration, and finally the calibrated component name, confidence level and incremental learning update status are obtained.
[0042] The third aspect of the present invention provides: a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are loaded and executed by a processor, any of the above-mentioned fine-grained recognition methods for special equipment components based on machine vision is implemented.
[0043] A fourth aspect of the present invention provides: a computer program product comprising instructions, which, when running on a terminal, enables the terminal to execute any of the above-mentioned methods for fine-grained identification of special equipment components based on machine vision.
[0044] The foregoing description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein and should not be construed as excluding other embodiments. Rather, the present invention can be used in various other combinations, modifications, and environments and can be modified within the scope of the concept described herein through the above teachings or techniques or knowledge in the relevant field. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention are intended to be protected by the appended claims.
Claims
1. A fine-grained recognition method for special equipment components based on machine vision, characterized by: The following steps are involved: During the multispectral image acquisition phase, the acquisition equipment is configured, different acquisition strategies are adopted for different lighting environments, and motion jitter is compensated in real time to reduce image offset errors, thereby obtaining multispectral images and component spatial position data. In the component area positioning and occlusion compensation stage, component area positioning, dynamic occlusion compensation, and key point annotation are performed to obtain the component feature map after background removal, occlusion mask, and key point coordinates; In the multimodal feature fusion and enhancement stage, multispectral feature fusion, feature enhancement, and occlusion robustness training are performed to obtain the fused feature vector and component structure heat map; During the fine-grained dynamic classification model inference phase, the model is deployed in a lightweight manner and inference optimization is performed in real time. Feature vectors are input into the dynamic classification network, and a dynamic attention mechanism is introduced to automatically switch attention heads based on component types to obtain the component category probability distribution. During the confidence calibration and incremental learning update phase, the confidence levels are graded and different processing strategies are adopted for different levels. The incremental learning mechanism and anomaly detection module are introduced for calibration, and finally the calibrated component name, confidence level and incremental learning update status are obtained.
2. The method for fine-grained recognition of special equipment components based on machine vision according to claim 1 is characterized in that: The IMU sensor is used to compensate for motion jitter in real time. The acquisition equipment includes an RGB camera and an infrared camera. The HDR mode of the RGB camera is used to capture component details in a strong light environment, and the infrared camera is used to penetrate oil and dust to enhance the component outline in a weak light environment.
3. The method for fine-grained recognition of special equipment components based on machine vision according to claim 1 is characterized in that: The component area positioning and occlusion compensation stage further includes the following steps: The multispectral image is input into the pre-trained YOLOv8n model to output the rough positioning bounding box of the component. For occluded scenes, the main area of the component is extracted through saliency detection, and the background interference is suppressed to obtain the component feature map. When a component is occluded, a pre-trained 3D model of the component is used to generate a virtual outline of the occluded area. The LSTM network is used to analyze the previous and next frame data and predict the structural integrity of the occluded component to obtain an occlusion mask. The key functional points of the components are marked to assist in fine-grained classification to obtain the key point coordinates.
4. The method for fine-grained identification of special equipment components based on machine vision according to claim 1 is characterized in that: During the training phase, occlusion regions are randomly added to the complete component image to simulate a real scene; a contrast loss function is used to enhance the consistency of features of similar components; the multimodal feature fusion and enhancement phase further includes the following steps: Input the RGB image and infrared pattern into the two-branch convolutional network and output a 4-channel feature map; dynamically assign spectral weights through SE Block; EfficientNet-B4 is used to extract the overall structure of the component to obtain global features; high-resolution sub-images are cropped based on the coordinates of key points, and detailed textures are extracted through ResNet-18 to obtain local features; global features and local features are spliced together to obtain feature vectors.
5. The method for fine-grained recognition of special equipment components based on machine vision according to claim 1 is characterized in that: The fine-grained dynamic classification model inference phase also includes the following steps: Using the knowledge distillation method, ViT-Base is used as the teacher model and MobileNetV3 is used as the student model. The model size is compressed to 10MB, and TensorRT is quantized to INT8, achieving lightweight deployment and improving inference speed. Use Qualcomm XR2 chip Hexagon DSP to accelerate matrix operations and achieve inference optimization.
6. The method for fine-grained recognition of special equipment components based on machine vision according to claim 1, characterized in that: The anomaly detection module calculates feature distribution differences based on Mahalanobis distance, identifies unknown component types, triggers AR interface warnings, records them, and transmits them to the cloud for processing; The confidence calibration and incremental learning update phase further includes the following steps: For the first confidence level, the recognition result is directly output and the label is displayed on the AR interface; for the second confidence level, multi-model voting is triggered to improve the consistency of the recognition results; for the third confidence level, manual review is prompted, and the marked data is returned to the training set; For newly added component types, the classification head is updated through comparative learning, multi-device collaborative training, and encrypted uploaded features.
7. The method for fine-grained recognition of special equipment components based on machine vision according to any one of claims 1 to 6, characterized in that: The backbone network of the dynamic classification network is Vision Transformer.
8. A fine-grained recognition system for special equipment components based on machine vision, characterized by: A method for implementing fine-grained recognition of special equipment components based on machine vision according to any one of claims 1 to 7, comprising: The multispectral image acquisition module is used to configure the acquisition equipment, adopt different acquisition strategies for different lighting environments, and compensate for motion jitter in real time to reduce image offset errors, thereby obtaining multispectral images and component spatial position data; The component area positioning and occlusion compensation module is used to perform component area positioning, dynamic occlusion compensation and key point annotation to obtain the component feature map after removing the background, occlusion mask and key point coordinates; Multimodal feature fusion and enhancement module, used for multispectral feature fusion, feature enhancement and occlusion robustness training, to obtain the fused feature vector and component structure heat map; The fine-grained dynamic classification model inference module is used for lightweight model deployment and real-time inference optimization. It inputs feature vectors into the dynamic classification network and introduces a dynamic attention mechanism to automatically switch attention heads based on component types to obtain component category probability distribution. The confidence calibration and incremental learning update module is used to grade confidence and adopt different processing strategies for different levels. The incremental learning mechanism and anomaly detection module are introduced for calibration, and finally the calibrated component name, confidence level and incremental learning update status are obtained.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are loaded and executed by the processor, the fine-grained recognition method for special equipment components based on machine vision as described in any one of claims 1 to 7 is implemented.
10. A computer program product comprising instructions, characterized in that: When the computer program product is run on a terminal, the terminal is enabled to execute the fine-grained recognition method for special equipment components based on machine vision according to any one of claims 1 to 7.
Citation Information
Cited By
Deep learning-based tower crane detection method and system
CN120931658A
Machine tool fixture visual perception positioning method and system based on marker
CN121236169A