AR point inspection intelligent camera control and AI identification method and related equipment

By leveraging the environmental perception and adaptive image acquisition technology of AR inspection smart cameras, combined with lightweight AI models and augmented reality interfaces, the problems of low recognition accuracy and poor environmental adaptability in industrial inspections have been solved, enabling efficient and accurate equipment status monitoring and operational compliance verification.

CN122367981APending Publication Date: 2026-07-10TANGSHAN LINGYUN TWIN TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610495558.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-15
Publication Date
2026-07-10

Smart Images

  • Figure CN122367981A_ABST
    Figure CN122367981A_ABST
Patent Text Reader

Abstract

This invention discloses an AR inspection intelligent camera control and AI recognition method and related equipment, belonging to the interdisciplinary field of industrial intelligent inspection and augmented reality technology. The method includes: real-time monitoring of the light intensity in the current environment of the inspected equipment and the reflectivity of the target material of the inspected equipment; adaptively activating the corresponding acquisition module to acquire target image data of the inspected equipment; real-time recognition of the target image data through a lightweight AI model deployed on an edge computing device to obtain equipment status information and operation behavior information of the inspected equipment; compliance verification of the operation behavior information; if a violation is found, a warning message is output through the augmented reality interface, and the corresponding defect handling process is triggered. This invention can complete the inspection task in real time and efficiently, comprehensively improving the shortcomings of existing technologies such as low recognition accuracy, weak operation guidance, and poor environmental adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial intelligent inspection and augmented reality technology, and in particular to an AR inspection intelligent camera control and AI recognition method and related equipment. Background Technology

[0002] Industrial inspection is a crucial link in ensuring the safe operation of production equipment. Traditional inspection methods mainly rely on manual visual inspection or taking pictures with ordinary cameras for analysis by back-end technicians. This approach is not only inefficient but also easily affected by subjective factors, making it difficult to guarantee the accuracy and consistency of inspection results. With the improvement of industrial automation, higher requirements are placed on the real-time performance, accuracy, and intelligence of inspection.

[0003] To address these issues, some advanced solutions attempt to combine artificial intelligence image recognition technology with augmented reality (AR) displays, assisting operators by overlaying virtual information onto captured images. These solutions typically utilize visible light cameras to acquire device images, perform recognition using cloud-based or local AI models, and then overlay the recognition results as static information onto the AR interface, attempting to provide some guidance.

[0004] However, these existing solutions generally suffer from three major pain points in practical applications. First, low recognition accuracy. Relying solely on a single visible light camera, the image quality deteriorates significantly in complex industrial environments, such as under direct sunlight, in dim lighting, or when the target surface contains reflective or transparent materials. This leads to a substantial decrease in the accuracy of AI recognition. Second, weak operation guidance. Existing augmented reality systems mostly overlay pre-set static information and cannot dynamically adjust based on the operator's real-time actions. They lack real-time compliance verification of the order of operation steps, the correctness of tool use, and key parameter thresholds, making it difficult to effectively prevent misoperation. Third, poor environmental adaptability. These systems lack the ability to adapt to changes in ambient lighting and target material characteristics, and cannot automatically adjust data acquisition strategies according to on-site conditions, resulting in insufficient robustness of the system in different scenarios.

[0005] In summary, existing technical solutions are often simply viewed as a conventional integration of "artificial intelligence models and augmented reality displays," failing to fundamentally address the core pain points of industrial inspection sites. Therefore, a high-precision intelligent inspection method capable of adapting to environmental changes and providing real-time operational compliance verification is needed. Summary of the Invention

[0006] This invention aims to address the shortcomings of existing technologies by providing an AR point-of-sight intelligent camera control and AI recognition method and related equipment, as detailed below: 1) In a first aspect, the present invention provides an AR point-of-sight intelligent camera control and AI recognition method, the specific technical solution of which is as follows: The AR inspection smart camera uses an environmental perception unit to monitor in real time the light intensity in the current environment where the inspected equipment is located and the reflective properties of the target material of the inspected equipment. Based on preset switching trigger conditions and light intensity and reflection characteristics, the corresponding acquisition module is adaptively activated from the visible light camera, infrared camera and depth sensor integrated in the AR inspection smart camera to acquire target image data of the inspected device. By using a lightweight AI model deployed on edge computing devices to identify target image data in real time, the device status information and operation behavior information of the inspected equipment can be obtained. Based on the built-in operation compliance rule library, the operation behavior information is verified for compliance. If the verification result is a violation, a warning message will be output through the augmented reality interface, and the corresponding defect handling process will be triggered according to the severity of the violation.

[0007] The beneficial effects of the AR point-of-sight intelligent camera control and AI recognition method provided by this invention are as follows: By using an environmental perception unit to monitor light intensity and the reflectivity of target materials in real time, and adaptively activating visible light cameras, infrared cameras, or depth sensors based on preset switching trigger conditions for data fusion and acquisition, the system effectively overcomes the interference of complex environments on imaging quality, ensuring high-quality acquisition of target image data. This improves the recognition accuracy and robustness of the lightweight AI model for device status and operational behavior information. Based on a built-in operational compliance rule base, real-time compliance verification of operational behavior information is performed, enabling dynamic monitoring and proactive error prevention of the operation process, avoiding safety risks caused by incorrect steps or misuse of tools. When the verification result is a violation, a warning message is output through the augmented reality interface, and the corresponding defect handling process is triggered according to the severity, directly converting the identification and verification results into management actions, improving problem response and processing efficiency. Simultaneously, the lightweight AI model is deployed on edge computing devices, reducing reliance on cloud resources and lowering network latency. It can complete inspection tasks in real time and efficiently, comprehensively improving the shortcomings of existing technologies such as low recognition accuracy, weak operation guidance, and poor environmental adaptability.

[0008] Based on the above solution, the AR point inspection smart camera control and AI recognition method of the present invention can be further improved as follows.

[0009] Furthermore, based on preset switching trigger conditions and considering light intensity and reflection characteristics, the corresponding acquisition modules are adaptively activated from the visible light camera, infrared camera, and depth sensor integrated into the AR point-of-sight smart camera to acquire target image data. Specifically, this includes: comparing the light intensity with a preset low light threshold; if the light intensity is lower than the low light threshold, the infrared camera is activated as the acquisition module; acquiring the reflection characteristics of the target material; if the target material is determined to be a metallic reflective material based on the reflection characteristics, the visible light camera and depth sensor are activated for data fusion acquisition; if the light intensity is higher than the low light threshold and the target material is a non-metallic reflective material, the visible light camera is activated as the acquisition module.

[0010] The beneficial effects of adopting the above-mentioned further solution are as follows: By monitoring the light intensity and the reflectivity of the target material in real time, the most suitable acquisition module is automatically selected based on preset switching trigger conditions. When the light intensity is below the low light threshold, the infrared camera is activated to ensure clear imaging even in low-light environments; when metallic reflective materials are detected, the visible light camera and depth sensor are activated to perform data fusion acquisition, using depth information to overcome image overexposure caused by specular reflection and preserve complete target structural information; under normal lighting and non-metallic reflective materials, the visible light camera is activated to acquire high-fidelity color images. This adaptive activation strategy ensures that the AR inspection smart camera always acquires target image data in the optimal way, guaranteeing image quality from the source and laying the foundation for subsequent accurate identification. At the same time, it avoids the limitations of a single sensor in complex environments and improves the system's adaptability to different industrial scenarios.

[0011] Furthermore, the process of acquiring lightweight AI models includes: The original AI model is pruned using a structured approach based on channel importance assessment to remove redundant neurons and obtain the pruned model. The pruned model is subjected to 8-bit integer quantization, which compresses the model weights and activation values ​​from 32-bit floating-point data to 8-bit integer data, resulting in a lightweight AI model.

[0012] The beneficial effects of adopting the above-mentioned further solution are as follows: By performing structured pruning based on channel importance assessment on the original AI model, redundant neurons are accurately removed, effectively reducing the number of model parameters and computational complexity, resulting in a pruned model. Further 8-bit integer quantization is applied to the pruned model, compressing the model weights and activation values ​​from 32-bit floating-point data to 8-bit integer data, significantly reducing the model's storage size and memory usage. This dual optimization of pruning and quantization enables the lightweight AI model to maintain a recognition accuracy rate above 92% while achieving high-speed real-time inference on resource-constrained edge computing devices, meeting the processing requirements of more than 20 frames per second. This provides AR inspection smart cameras with low-latency, high-energy-efficiency intelligent recognition capabilities, ensuring the real-time performance and reliability of inspection tasks.

[0013] Furthermore, before outputting warning information through the augmented reality interface, a spatial calibration step is included for the content displayed on the augmented reality interface. Specifically, this includes: initial positioning by identifying fixed markers on the device being inspected to obtain coarse pose information; acquiring 3D point cloud data of the device being inspected through a depth sensor, and iteratively registering the 3D point cloud data with the CAD model of the device being inspected to obtain fine pose information; and dynamically correcting the virtual information displayed in the augmented reality interface based on the fine pose information to ensure that the virtual information is aligned with the device being inspected at the sub-centimeter level.

[0014] The beneficial effects of adopting the above-mentioned further scheme are as follows: By spatially calibrating the content displayed on the augmented reality interface, initial positioning is first achieved by identifying fixed markers on the device being inspected, obtaining coarse pose information to provide an accurate initial estimate for subsequent registration. Then, a depth sensor is used to acquire 3D point cloud data of the device being inspected, and this 3D point cloud data is iteratively registered with the nearest point of contact of the device's CAD model to obtain fine pose information, achieving a precise conversion from the 3D point cloud data coordinate system to the CAD model coordinate system. Based on the fine pose information, the virtual information displayed in the augmented reality interface is dynamically corrected to ensure that the virtual information is always aligned with the device being inspected at the sub-centimeter level. This two-stage calibration method effectively compensates for deviations caused by user movement or minor device deformation, enabling the indicator arrows, text prompts, and other content superimposed on the augmented reality interface to be accurately attached to the corresponding physical components, improving the accuracy and usability of information guidance, and ensuring the reliability of the inspection operation.

[0015] 2) Secondly, the present invention also provides an AR point-of-sight intelligent camera control and AI recognition system, the specific technical solution of which is as follows: It includes a monitoring module, a trigger acquisition module, an identification module, a compliance verification module, and an alert module; The monitoring module is used to: monitor in real time the light intensity in the current environment where the inspected equipment is located and the reflective properties of the target material of the inspected equipment through the environmental perception unit of the AR inspection smart camera; The trigger acquisition module is used to: adaptively activate the corresponding acquisition module from the visible light camera, infrared camera and depth sensor integrated in the AR inspection smart camera according to the preset switching trigger conditions and based on the light intensity and reflection characteristics, so as to acquire the target image data of the inspected device. The recognition module is used to: perform real-time recognition of target image data using a lightweight AI model deployed on edge computing devices, and obtain equipment status information and operation behavior information of the inspected equipment; The compliance verification module is used to: verify the compliance of operational behavior information based on the built-in operational compliance rule library; The alert module is used to: output alert information through the augmented reality interface if the verification result is a violation, and trigger the corresponding defect handling process according to the severity of the violation.

[0016] 3) In a third aspect, the present invention also provides an electronic device, the electronic device including a processor coupled to a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor, so as to enable the electronic device to implement any of the above-mentioned AR point inspection smart camera control and AI recognition methods.

[0017] 4) In a fourth aspect, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described AR point inspection smart camera control and AI recognition method.

[0018] It should be noted that the beneficial effects of the technical solutions of the second to fourth aspects of the present invention and their corresponding possible implementations can be found in the above description of the technical effects of the first aspect and its corresponding possible implementations, and will not be repeated here. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments of the present invention will be briefly introduced below: Figure 1 This is a flowchart illustrating an AR inspection smart camera control and AI recognition method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of an AR inspection smart camera control and AI recognition system according to an embodiment of the present invention. Detailed Implementation

[0020] The principles and features of the present invention are described below. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0021] The technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the present invention will now be described with reference to the accompanying drawings.

[0022] like Figure 1 As shown in the figure, an AR inspection smart camera control and AI recognition method according to an embodiment of the present invention includes the following steps: S1. The environmental perception unit of the AR inspection smart camera monitors the light intensity and reflectivity of the target material of the inspected device in real time. The specific implementation process is as follows: S10. The photosensitive sensor integrated in the environmental sensing unit collects real-time data on the ambient light intensity of the environment in which the inspected equipment is located. The photosensitive sensor is a semiconductor device based on the photoelectric effect. Its core working principle is that when incident light shines on the sensor's photosensitive surface, photon energy excites electron transitions within the material, generating photogenerated carriers and thus forming a photocurrent proportional to the incident light intensity. This photocurrent signal is extremely weak. It is first converted into a voltage signal by a high-precision transimpedance amplifier and simultaneously amplified. The amplified voltage signal then enters an analog-to-digital converter (ADC). The ADC samples and quantizes the analog voltage at a fixed sampling rate, converting it into a discrete digital quantity. This digital quantity represents the intensity of the current ambient light, measured in lux (Lux). The sampling rate can be 20 times per second. To ensure measurement accuracy, the spectral response characteristics of the photosensitive sensor are calibrated before leaving the factory to match the standard visual function V(λ) curve, thus ensuring that the measured light intensity truly reflects the brightness perceived by the human eye. The digital light intensity data output from the analog-to-digital converter is transmitted to the central processing unit (CPU) of the AR point-of-sight smart camera via a serial peripheral interface bus. The CPU appends a precise timestamp to each received data point and stores it in a circular buffer for subsequent processing. Simultaneously, to eliminate drastic fluctuations in measurement values ​​caused by light source flicker or momentary occlusion, the CPU performs a moving average filter on five consecutive sampled values. The filter calculation formula is as follows: ,in, This represents the digital value of the original illumination intensity from the i-th sample. This indicates the current effective ambient light intensity after filtering. This step ensures that the light intensity data provided to the system for decision-making is stable and reliable.

[0023] S11. The visible light camera of the AR inspection smart camera continuously acquires images containing the target area of ​​the inspected equipment, and obtains the reflectivity of the target material of the inspected equipment in real time based on image analysis. The visible light camera captures images at a rate of thirty frames per second, and each frame is sent to the image signal processor for preprocessing. The image signal processor first performs black level correction to eliminate the influence of sensor dark current, then performs lens shadow correction to compensate for edge brightness attenuation, then applies an automatic white balance algorithm to restore the true color of the image, and finally converts the Bayer array data into a complete RGB color image through a demosaicing algorithm. The preprocessed color image is sent to a dedicated digital signal processor or neural network processing unit that runs a reflectivity analysis algorithm. The reflectivity analysis first converts the RGB image into a grayscale image, using the standard luminance equation: Where R, G, and B represent the pixel values ​​of the red, green, and blue channels in a color image, respectively. This is the converted grayscale value.

[0024] The algorithm then delineates the target area for the equipment being inspected. This area may be obtained through pre-defined fixed locations or automatic identification by a target detection network. Within the target area, the algorithm calculates a grayscale histogram and extracts statistical features from it. The formula for calculating the grayscale mean μ is: ,in, Let be the grayscale value of the i-th pixel within the target region, and N be the total number of pixels within the target region. The formula for calculating the grayscale standard deviation σ is: .

[0025] In addition to grayscale statistical features, the algorithm also specifically targets highlight areas on reflective metallic materials. Highlight area detection is achieved by setting a high grayscale threshold. This is achieved by marking pixels with grayscale values ​​greater than a threshold as candidate highlight points, then performing region growing based on the spatial continuity of these points to ultimately determine the highlight region. The proportion of the highlight region area to the total target region area is then considered. The calculation formula is: ,in, This represents the number of pixels in the highlight area. This represents the total number of pixels in the target area. The system's built-in classifier determines reflectivity based on preset rules, taking into account the mean grayscale value, standard deviation, and highlight ratio. For example, if... Greater than the preset metal reflection detection threshold And μ is greater than the highlight threshold If the target material's reflectivity is determined to be metallic and reflective, then if μ is below the dimness threshold... If the reflection characteristic is strong, it is classified as a strong absorbing material; otherwise, it is classified as a diffuse reflecting material. After each frame of image analysis is completed, the reflection characteristic classification result is also timestamped and stored in the buffer.

[0026] S12. Continuously monitor the update status of the two data buffers. When a new illumination intensity sample value or a new reflectance characteristic analysis result is received, it aligns the two according to the timestamp. Since reflectance characteristic analysis relies on visible light images, its update rate is usually equal to the camera frame rate, while the illumination intensity sampling rate may be higher. The central processing unit uses the most recently valid illumination intensity value to pair with the current reflectance characteristic result. The paired data pairs are encapsulated into a unified environment-aware data packet, which contains the ambient illumination intensity value. And reflectivity category identifier. This data packet serves as the basis for subsequent adaptive start-up of the acquisition module. In addition, to ensure continuous monitoring, the environmental sensing unit continuously repeats steps S10 and S11 during system operation, forming a continuously updated data stream, thereby achieving real-time monitoring of light intensity and reflectivity.

[0027] Among them, the AR inspection smart camera refers to a specialized industrial device that integrates augmented reality display functions, multimodal image acquisition sensors, and embedded edge computing processing capabilities. It can acquire visible light images, infrared thermal images, and depth data of the equipment being inspected in real time, analyze them through a built-in lightweight artificial intelligence recognition model, and accurately overlay the recognition results and guidance information onto the real field of view in the form of virtual images to assist operators in equipment inspection and maintenance.

[0028] The environmental perception unit is a functional module within the AR inspection smart camera, consisting of a photosensor, related signal conditioning circuitry, and image processing algorithms. Its task is to measure the physical parameters of the environment in which the inspected device is located in real time, primarily including light intensity and the reflectivity of the target material obtained through image analysis. These parameters are then provided to the system's main controller for adaptive adjustment of the camera's operating mode.

[0029] Illumination intensity refers to the strength of visible light in the current environment where the equipment being inspected is located. It is measured as the luminous flux received per unit area, and the SI unit is lux (Lux). This parameter directly determines whether a visible light camera can obtain high-quality images and is a key indicator for determining whether an infrared camera or auxiliary lighting source needs to be used.

[0030] The equipment being inspected refers to various mechanical, electrical, or electronic devices in industrial production sites that require regular status checks, parameter measurements, or operational compliance monitoring. These devices are the monitoring targets of the AR inspection smart camera; the condition of their surfaces, readings, and the interactions between them and the operators are all aspects of the system's focus.

[0031] The target material refers to the type of material and its physical properties used in the specific area of ​​the equipment surface to be inspected during the inspection process. Different materials, such as metals, plastics, and glass, have different optical properties, which can affect the imaging effect and the accuracy of artificial intelligence recognition. Therefore, it is necessary to distinguish them through reflection characteristic analysis.

[0032] Among them, reflection characteristics refer to the reflection behavior of the target material surface to incident light, including specular reflection, diffuse reflection, absorption and other modes. For reflective metallic materials, their surfaces will produce strong directional reflection, which can easily lead to local overexposure in visible light images. Therefore, the system needs to automatically activate multimodal data fusion strategies by monitoring reflection characteristics in real time, such as combining depth sensor data to overcome reflection interference.

[0033] S2. Based on preset switching trigger conditions and considering light intensity and reflection characteristics, the corresponding acquisition module is adaptively activated from the visible light camera, infrared camera, and depth sensor integrated into the AR inspection smart camera to acquire target image data of the inspected device. Specifically, the light intensity is compared with a preset low light threshold. If the light intensity is lower than the low light threshold, the infrared camera is activated as the acquisition module. The reflection characteristics of the target material are obtained. When the target material is determined to be a metallic reflective material based on the reflection characteristics, the visible light camera and depth sensor are activated for data fusion acquisition. If the light intensity is higher than the low light threshold and the target material is a non-metallic reflective material, the visible light camera is activated as the acquisition module. The specific implementation process is as follows: The central processing unit of the S20 AR inspection smart camera continuously maintains a shared data buffer, which stores environmental perception data packets that have undergone time synchronization and filtering. Each data packet contains two key fields: the current effective ambient light intensity. The target material reflection property category derived from the analysis of the current frame image. The reflectivity category is an enumerated variable, which can take values ​​such as metallic reflective material, non-metallic reflective material, or strongly absorbing material. The central processing unit reads the latest pair of data from this buffer at a fixed control cycle, for example, every fifty milliseconds, as input for the adaptive startup decision of the current cycle. This step ensures that all subsequent decisions are based on real-time and accurate environmental information.

[0034] S21. Read the light intensity With the preset low light threshold Compare the preset low light threshold. It is a value stored in non-volatile memory, and its specific value can be set during system configuration according to different application scenarios. For example, in a typical industrial indoor environment, it can be set to [value missing]. The central processing unit performs a comparison operation: if If the ambient light intensity is insufficient, the visible light imaging quality cannot be guaranteed, thus meeting the low light condition in the preset switching trigger conditions. At this point, the system immediately executes the infrared camera startup procedure. The infrared camera startup procedure includes: sending an enable signal to the power management chip of the infrared camera module via a general-purpose input / output interface to turn on the infrared camera power; writing initialization configuration parameters, such as integration time, gain factor, and frame rate, to the infrared camera register via the integrated circuit bus; and after waiting for the infrared camera to output a stable image, switching the data stream gating switch to the infrared camera data channel so that subsequent target image data acquisition and artificial intelligence recognition both use the infrared image as the input source. If the light intensity is sufficient, the system continues to execute the judgment in step S22.

[0035] S22. Under conditions of sufficient light intensity, further classify the reflectance characteristics. The process involves the central processing unit (CPU) reading the reflection characteristic classification result and checking if it equals a preset metallic reflective material identifier. This identifier is internally represented by an integer value, for example, 1 for metallic reflective material. If... If the target material of the inspected equipment is determined to be a reflective metallic material, meeting the metallic reflectivity condition in the preset switching trigger conditions, the system then initiates data fusion acquisition using a visible light camera and a depth sensor. Specifically, the central processing unit first activates the visible light camera, configuring it to operate with exposure parameters suitable for the current lighting conditions; simultaneously, it activates the depth sensor, configuring its operating mode to be synchronized with the visible light camera's trigger mode. Typically, the acquisition times of the two sensors are aligned via hardware trigger lines. The system then enters the fusion acquisition state, with the visible light camera outputting a high-resolution color image, and the depth sensor outputting a depth map aligned with the pixel coordinates of the color image. This depth map can be registered with the color image through projection transformation, providing three-dimensional spatial information for each pixel for subsequent artificial intelligence recognition, thus overcoming the loss of information in locally overexposed areas caused by specular reflection from the metal surface.

[0036] S23. If the light intensity is higher than the low light threshold and the reflectivity is determined to be that of a non-metallic reflective material, that is... and If the material is not a metallic reflective material, the system activates the visible light camera as the acquisition module. The central processing unit executes the standard startup procedure for the visible light camera, including enabling power, loading default configuration parameters such as auto exposure and auto white balance, and switching the data stream to the visible light camera channel. In this mode, the depth sensor remains in standby or low-power state, not outputting data to save energy and computing resources. The acquired visible light images are directly fed into a lightweight artificial intelligence model for real-time recognition.

[0037] S24. After adaptive startup, the system continuously monitors environmental changes and dynamically adjusts the acquisition modules. The above judgment and switching process is repeated in each control cycle. If the ambient light intensity changes from bright to dark, crossing the low light threshold, the system will switch from visible light mode or fusion mode to infrared mode; conversely, if the light intensity recovers and the reflectivity is non-metallic, it will switch back to visible light mode. When the system is in infrared mode, if the depth sensor data needs to be used for a specific recognition task, the depth sensor and infrared camera can be temporarily fused, but the core decision logic always follows the preset switching trigger conditions. All module switching processes are managed by a software state machine to ensure smooth, flicker-free switching and guarantee the continuity and stability of the target image data.

[0038] Among them, the visible light camera is an imaging module in the AR inspection smart camera used to capture the reflected light in the visible spectrum of the inspected equipment. Its output is a three-channel color image of red, green and blue, which is consistent with human visual perception and is suitable for routine inspection tasks in scenes with sufficient lighting and no reflective interference from the target material.

[0039] Among them, the infrared camera is an imaging module in the AR inspection smart camera used to detect the thermal radiation or reflected infrared light of the equipment being inspected. It does not rely on ambient visible light illumination and can produce clear images in complete darkness or extremely low light conditions, making it particularly suitable for detecting abnormal heating of equipment or observing targets through partially obstructed objects.

[0040] Among them, the depth sensor is an active sensing module in the AR inspection smart camera used to measure the distance between the camera and various points on the surface of the inspected equipment. It acquires three-dimensional spatial point clouds through structured light or time-of-flight technology, which can provide depth information for the image and help overcome imaging interference from special materials such as metallic reflections.

[0041] In this context, the acquisition module refers to a combination of one or more sensor units in the AR point-of-sight smart camera that are actually activated and used to acquire image data of the current target. The acquisition module can be a single visible light camera, a single infrared camera, or a collaborative combination of a visible light camera and a depth sensor.

[0042] The target image data refers to the sequence of raw digital images output by the acquisition module, reflecting the appearance, temperature, or three-dimensional shape of the equipment being inspected. This data serves as the foundational input for subsequent lightweight AI models to identify equipment status and analyze operational behavior.

[0043] The preset switching trigger conditions refer to a set of logical rules pre-stored in the AR inspection smart camera's memory. This set of rules defines which acquisition module should be activated under what combination of environmental parameters, specifically including low light threshold, metallic reflective material judgment threshold, and the logical relationships between each condition.

[0044] Non-metallic reflective materials refer to target surface materials other than metallic reflective materials, such as engineering plastics, rubber, and coatings. These materials typically exhibit diffuse reflection, providing good visible light imaging quality and eliminating the need for depth sensors.

[0045] S3. Real-time identification of target image data is performed using a lightweight AI model deployed on edge computing devices to obtain equipment status information and operation behavior information of the inspected equipment. The process of obtaining a lightweight AI model includes: performing structured pruning on the original AI model based on channel importance assessment to remove redundant neurons and obtain a pruned model; and performing 8-bit integer quantization on the pruned model to compress the model weights and activation values ​​from 32-bit floating-point data to 8-bit integer data to obtain a lightweight AI model. Specifically: 1) First method: If the original AI model is a pre-trained model, then: S30. Obtain the original AI model. The original AI model is a deep convolutional neural network pre-trained on a large general dataset or a specific industrial inspection dataset, possessing high-precision image recognition capabilities. The network structure consists of multiple stacked convolutional modules, each containing a convolutional layer, a batch normalization layer, and a ReLU activation function layer. At the end of the network, there are typically global average pooling layers and fully connected layers for outputting classification or regression results. The convolutional layer is the core of the model, and its parameters are four-dimensional weight tensors. ,in, Indicates the number of output channels. Indicates the number of input channels. This indicates the size of the convolutional kernel. The batch normalization layer normalizes the output of the convolutional layer and introduces a trainable scaling factor. Translation factor These parameters have been determined and stored after pre-training.

[0046] S31. Perform structured pruning on the original AI model based on channel importance evaluation. This step uses the scaling factor in the batch normalization layer. As a metric for measuring the importance of each channel, each channel i in the corresponding batch normalization layer of each convolutional layer has a scaling factor. The absolute value of the scaling factor The larger the value, the greater the contribution of that channel to the final output, and therefore the more important it is. The system iterates through all convolutional layers, collecting the values ​​of all channels in each layer. The values ​​are sorted in ascending order. The global pruning ratio is determined based on a preset percentage. ,For example This means removing 50% of the channels. For each layer, the channels that are sorted and placed first will be... The channels with proportional characteristics are marked as redundant channels. Subsequently, the convolutional kernels corresponding to these redundant channels are removed from the weight tensor. Remove slices corresponding to the output dimension and adjust the number of input channels in the next layer accordingly. After pruning, the pruned model is obtained. The network layer structure of the pruned model is the same as the original model, but the number of output channels in each convolutional layer is reduced, resulting in a significant decrease in model size and computational cost.

[0047] S32. Perform 8-bit integer quantization on the pruned model. Quantization aims to convert all weight parameters and intermediate activation values ​​in the model from 32-bit floating-point data to 8-bit integer data, further reducing the model size and accelerating inference on edge devices. The specific implementation includes the following sub-steps: First, for the weight tensor of each layer... Calculate its maximum value and minimum value Then determine the quantization parameters for that layer, i.e., the scaling factor. and zero point The calculation formula is: , Where 255 is the maximum range that an 8-bit integer can represent. This indicates rounding to the nearest integer.

[0048] Then, each floating-point number in the weight tensor Quantized to 8-bit integer : and ensure The values ​​are pruned to the range of 0 to 255. For activation values, since they change dynamically during inference, their distribution range needs to be statistically analyzed. To this end, a small calibration dataset is prepared, consisting of representative images collected from actual inspection scenarios. The calibration dataset is input into the pruned model for forward inference, while simultaneously recording the activation value distribution of each layer's output. This yields the maximum and minimum values ​​of each activation tensor, allowing for the calculation of the corresponding quantization parameters. After calculating the quantization parameters for all layers, the model's weights and activation values ​​can be represented and manipulated using 8-bit integers, resulting in the final lightweight AI model. This model can be directly deployed on edge computing devices of AR inspection smart cameras to achieve high-speed, real-time inference.

[0049] 2) The second method: If the original AI model is not pre-trained, then: S33. The network layer structure of the original AI model is the same as the first method, consisting of multiple convolutional modules. Each module contains a convolutional layer, a batch normalization layer, and a ReLU activation layer. The model weights are initialized using a random method, such as He initialization, to give the weights a suitable initial distribution. At this point, the model has not undergone any training, and all parameters are random values.

[0050] S34. Warm-up training of the original AI model. Since the model weights are randomized, effective channel importance assessment cannot be performed directly. Therefore, it is necessary to first train the model using the target task's dataset for a small number of iterations, allowing the weights to initially possess feature extraction capabilities. Warm-up training typically lasts only a few epochs, such as five epochs, with a small learning rate, and the optimization objective is to minimize the loss function of the point inspection task. After warm-up training, the scaling factor in the model's batch normalization layer... This begins to reflect the relative importance of different channels, providing a basis for subsequent pruning.

[0051] S35. Perform structured pruning based on channel importance evaluation on the pre-trained model. This step is the same as step S31 in the first method, utilizing the scaling factor of the batch normalization layer. As an important indicator, the pruning ratio is set according to the preset proportion. Redundant channels are removed to obtain the pruned model. The number of channels in the network layer structure of the pruned model is reduced.

[0052] S36. Since the model has not yet converged, simple post-training quantization would lead to a significant loss of accuracy. Therefore, a quantization-aware training method is adopted to simulate the quantization effect during training, allowing the model to adapt to low-precision representation. The specific steps are as follows: First, pseudo-quantized nodes are inserted into the pruned model. During forward propagation, these nodes quantize the weights and activation values ​​to 8 bits and then dequantize them to floating-point, thus simulating quantization error. During backpropagation, the gradient is still transmitted in floating-point form, and weight updates are also performed in floating-point. For the weights, their quantization parameters... and In each training iteration, calculations are performed in real-time based on the current weights; for activation values, a moving average method is used to statistically determine their range. The model is then fully trained using the dataset from the point inspection task, optimizing the loss function until convergence. After training, the weights in the model have adapted to quantization noise, at which point they can be converted to 8-bit integers using the last calculated quantization parameters, and activation values ​​are dynamically quantized during inference using statistically obtained quantization parameters. The result is a lightweight AI model that combines a compact channel structure with low-bit quantization characteristics, enabling efficient operation and high accuracy on edge devices.

[0053] The original AI model refers to the basic deep neural network model used before lightweight processing. This model has a complete network layer structure and high computational complexity. It is usually pre-trained on large-scale datasets or initialized according to standard design. It has strong feature representation capabilities and is the starting point for subsequent pruning and quantization operations.

[0054] Among them, structured pruning based on channel importance assessment is a model compression technique. By analyzing the contribution of each channel in each convolutional layer of the network to the final output, channels with smaller contributions and their associated convolutional kernels are removed as a whole, thereby reducing the network width, achieving the goal of reducing model size and computational cost, while maintaining model accuracy as much as possible.

[0055] Redundant neurons refer to channels in the convolutional layer that contribute little to the final output. These channels are identified as removable after importance evaluation, and removing them has little impact on model performance, thus effectively simplifying the network structure.

[0056] The activation value refers to the intermediate feature map output after each layer of computation when the input data propagates forward through the network. These values ​​are continuously generated during the inference process, and their range directly affects the setting of quantization parameters, which has an important impact on quantization accuracy.

[0057] The quantization parameters include scaling factor and zero point, which are used to map floating-point numbers to the integer range. The scaling factor determines the step size, and the zero point determines the offset between integer and floating-point numbers. These are the key coefficients for realizing the conversion between floating-point and integer numbers.

[0058] The calibration dataset is a small, representative set of data used for post-training quantization. By running this dataset on the model, the distribution of activation values ​​in each layer is statistically analyzed, thereby determining the quantization parameters of the activation values. This avoids using the entire training set and improves efficiency.

[0059] Preheating training refers to training a small number of iterations before the model is fully trained, so that the model weights are initially effective and provide meaningful parameter basis for subsequent pruning based on importance assessment. It is a necessary step when dealing with untrained models.

[0060] Among them, quantization-aware training is a technique that simulates quantization operations during training. By inserting pseudo-quantization nodes in the forward propagation, the model can take into account the impact of quantization errors when updating parameters, thereby obtaining a model that still maintains high accuracy under low-precision inference.

[0061] The specific implementation process of S3 is as follows: 1) A lightweight AI model is loaded into the inference engine of the edge computing device. During system startup or model update, the edge computing device reads the pre-stored lightweight AI model file from non-volatile memory. This file contains a description of the model structure after structured pruning and 8-bit integer quantization, the quantization parameters for each layer, and the weight data represented by 8-bit integers. The inference engine on the edge computing device parses the model file, constructs a computational graph in memory, and allocates buffers for input, output, and intermediate results for each layer. The inference engine is also configured with underlying hardware acceleration units, such as neural network processors or digital signal processors, enabling it to efficiently perform 8-bit integer convolution, pooling, and fully connected operations. After loading is complete, the inference engine enters a ready state, waiting to receive target image data for inference.

[0062] 2) Edge computing devices acquire the latest frame of target image data from the acquisition module via a high-speed serial interface. This data may be a color image output from a visible light camera, a grayscale image output from an infrared camera, or a depth map output from a depth sensor. First, according to the requirements of the lightweight AI model input layer, the original image is scaled to a specified size, such as 224×224 pixels. The scaling uses a bilinear interpolation algorithm to ensure that the image content is not distorted. Then, the scaled image is converted from the original pixel format to the 8-bit integer data format required by the model input. For color images, the red, green, and blue channels are usually processed separately. If mean subtraction or scaling was used during model training, the pixel values ​​will be adjusted according to the stored quantization parameters. However, most quantization models have already integrated these operations into the first layer, so the pixel values ​​only need to be directly copied to the input buffer, maintaining a range of 0 to 255. The preprocessed data is placed in the input buffer, triggering the inference engine to start calculation.

[0063] 3) The inference engine computes layer by layer, starting from the input layer, following the topological order of the computation graph. For each layer, the inference engine calls a kernel function optimized for 8-bit integers. Taking a convolutional layer as an example, the kernel function converts the input activation value tensor... with weight tensor Perform convolution operations, where and All values ​​are 8-bit integers. Multiplication and accumulation use 32-bit integers to accumulate intermediate results to avoid overflow. After convolution, the accumulated result is dequantized and biased according to the pre-stored quantization parameters of the layer, and then requantized into an 8-bit integer as the input to the next layer. If the layer contains batch normalization, the batch normalization parameters are already integrated into the weights and biases of the convolutional layer and do not need to be calculated separately. Activation functions such as ReLU are also implemented through comparison operations. Throughout the inference process, the data flow is entirely in integer form between each level of cache and computation unit, making full use of the parallel computing capabilities of edge computing devices, so that the forward propagation time of a single frame image is controlled within 20 milliseconds, meeting the processing requirements of more than 20 frames per second. When the computation reaches the output layer, the output layer generates the original score vector or feature map.

[0064] 4) The raw output layer needs to be decoded according to the specific task. For device status information recognition tasks, the output layer is usually a fully connected layer, outputting a value of length [missing information]. score vector ,in This represents the total number of predefined device status categories, such as different reading ranges for pressure gauges or different color categories for indicator lights. The post-processing unit applies a softmax function to the score vector, transforming it into a probability distribution. ,in, Indicates the first The original scores for each category, Indicates the prediction is the first The system selects the category with the highest probability. This serves as the device status information for the current frame, and the corresponding confidence level is recorded. For tasks involving the recognition of operator behavior information, the model might be designed to simultaneously detect the operator's actions and the tools used. The output layer might contain multiple branches; for example, one branch might output the location and category of the target detection box, while another branch might output the action classification result. The post-processing unit applies non-maximum suppression to the detection results, removes overlapping detection boxes, and decodes the operator's hand position, tool type, and action sequence. Finally, the device status information and operator behavior information are encapsulated into structured data, including timestamps, recognition result categories, and confidence scores.

[0065] 5) Edge computing devices generate corresponding virtual instructions based on device status information, such as overlaying reading numbers above the dashboard or prompting for the next operation based on operation behavior information. The operation behavior information is then compared with a built-in rule base to determine if any violations exist. The entire process operates continuously in a pipeline manner, starting the next frame immediately after processing each one, ensuring uninterrupted real-time identification of the inspected devices.

[0066] Edge computing devices refer to hardware platforms deployed near data sources, possessing limited but efficient computing power, capable of independently completing data collection, processing, and decision-making without relying on remote cloud environments. In AR point-of-sight smart cameras, edge computing devices are typically integrated within the camera itself, including multi-core CPUs, graphics processors, and neural network accelerators, designed specifically for running lightweight AI models in real time.

[0067] Among them, equipment status information refers to the quantitative or classification results of the current operating status of the inspected equipment obtained by analyzing the target image data through a lightweight AI model, such as instrument readings, switch status, indicator light color, equipment temperature value, etc., which are used to determine whether the equipment is working properly.

[0068] Among them, operational behavior information refers to the action description obtained by recognizing the interaction process between the operator and the equipment being inspected through a lightweight AI model. This includes the completion status of the operation steps, whether the tools are used correctly, and whether the sequence of actions conforms to the standards, and is used to verify the compliance of operations in real time.

[0069] S4. Based on the built-in operation compliance rule base, the operation behavior information is verified for compliance. The specific implementation process is as follows: S40. Construct and initialize the operation compliance rule base. The operation compliance rule base is a structured dataset stored in the non-volatile memory of the edge computing device, and its content is derived from the standard operation procedure documents of the inspected equipment. The rule base is organized using a hierarchical data structure, with each inspected equipment corresponding to an independent rule table. In the rule table, a complete operation compliance rule consists of multiple fields, including rule number, applicable device identifier, operation step sequence, expected tool type, safety threshold parameter, and timing constraints. The operation step sequence records the operation code for each step in the form of an ordered list, such as "turn on the power", "read the pressure gauge", and "press the start button". Each operation code is defined separately in the rule base and associated with a set of allowed pre- and post-steps. The expected tool type field specifies the type of tool that must be used to perform the step, such as "multimeter" or "wrench". The safety threshold parameter is a set of key-value pairs used to store the upper or lower limits of the physical quantities that need to be monitored, such as "maximum pressure: 1.0 MPa" or "minimum temperature: 20 degrees Celsius". The timing constraints record the maximum allowed time interval between adjacent steps or the completion time limit of the entire process. The rule base is loaded into memory and indexed when the system starts up to support fast lookups. Simultaneously, the system maintains a runtime state machine for each inspected device. The state machine records the sequence of completed steps and the next step currently pending; its initial state is empty.

[0070] S41. Receive and parse the real-time identified operation behavior information. After the lightweight AI model performs inference on each frame of target image data, it outputs a structured operation behavior information data packet. This data packet contains a timestamp. Operational behavior categories Operation object identifier Confidence level and additional numerical parameters Among them, the operational behavior categories It is an enumeration value, such as "hand contact", "tool use", "knob rotation"; the operation object identifier. Specify the equipment component to which the action is performed, such as "pressure gauge dial" or "power switch"; additional numerical parameters. Used when the action involves reading numerical values, such as the current instrument reading or rotation angle. First, it is determined based on the confidence level of the data packet. The system filters input behavior information, retaining only those exceeding a preset threshold (e.g., 0.8) to ensure input reliability. Subsequently, it parses all valid behavior information from the current frame, creating a candidate list to be processed.

[0071] S42. Read the runtime state machine corresponding to the currently inspected equipment and obtain the sequence of steps that have been completed. and the set of next steps expected at present. The expected set of next steps. It consists of all legal subsequent steps derived from the operational compliance rule base based on the sequence of completed steps. For each operational behavior information in the candidate list, its operational behavior category is determined. With the expected next step set The steps in the code are compared. If a match is found, the operation is initially determined to be compliant in sequence, and further verification is performed. If no match is found, the operation is directly determined to be in violation of sequence, the violation type is recorded as "incorrect step sequence", and a violation result is generated.

[0072] S43. For successfully matched operation behavior information, the system extracts detailed constraints associated with that step from the rule base. First, it performs tool type validation; the rule base specifies the tool category that must be used to execute this step. The operation object identifier in the operation behavior information. This often implicitly includes tool information; for example, if it detects "a handheld multimeter near the terminals," the tool type is inferred to be "multimeter." The system then matches the actual tool type used with... The comparison is performed; if they do not match, the violation type is determined to be "tool error." Next, a security threshold check is performed; if this step involves numerical parameters... For example, when reading a stress value, the system compares it with a preset safety threshold in the rule base to check if the inequality constraint is satisfied. Assume the rule requires a stress value... Not exceeding ,like If the violation is found to be "exceeding the threshold", then the violation type is determined to be "exceeding the threshold". Finally, a time-series constraint check is performed, recording the timestamp of the current step. Calculate the time taken to complete the previous step. Time difference If the rule base specifies a maximum allowed interval Then check If the operation fails to meet the requirements, it is classified as a "timeout". Once all checks pass, the operation is marked as compliant.

[0073] S44. For operations deemed compliant, the system identifies the operation steps. Add to the currently completed sequence of steps In the process, the expected set of next steps is updated based on the rule base. After the state machine is updated, a data packet containing operation behavior information and verification results is generated, with the result indicating "compliant" or the specific "violation type". For operations judged as violations, the state machine does not update, but a result data packet with the violation type is still generated. The entire verification process runs in an event-driven manner, triggering a verification once for each frame of operation behavior information received, ensuring real-time monitoring of the operation process of the inspected equipment.

[0074] The operational compliance rule base refers to a structured knowledge base stored in edge computing devices, which contains the standard operating procedure definition for specific inspected equipment, including the order of operation steps, the tools allowed to be used in each step, the safety thresholds of key physical quantities, and time constraints, providing a basis for comparison for automatic verification.

[0075] In this context, a state machine is an abstract mathematical model used to track the current execution progress of the operational process of the inspected equipment. It contains a sequence of completed steps and all possible next steps, and is dynamically updated as compliance operations proceed, providing context for real-time verification.

[0076] Among them, safety threshold verification refers to comparing the numerical parameters carried in the operation behavior information, such as pressure, temperature or voltage, with the upper or lower limits preset in the rule base to ensure that the equipment operating parameters are within the safe range and prevent operation beyond the limits.

[0077] Among them, timing constraint verification refers to checking whether the time interval between adjacent operation steps exceeds the maximum value allowed in the rule base, or whether the entire process is completed within the specified time, in order to detect operation delays or interruptions.

[0078] S5. If the verification result is a violation, a warning message will be output through the augmented reality interface, and the corresponding defect handling process will be triggered according to the severity of the violation. The specific implementation process is as follows: S50. After verifying the current operation information, a structured verification result data packet will be generated. This data packet contains multiple fields: timestamp. Record the exact moment the violation occurred; identify the equipment being inspected. Specify which device the violation occurred on; operational behavior information. The description details the operations deemed to be in violation, including the operation category, the target object, and possible numerical parameters; the type of violation. It is an enumerated value, such as "incorrect step order", "tool error", "threshold exceeded", or "operation timeout"; in addition, the packet also contains a violation identifier. Set to true. Parse the violation type and operational behavior information to prepare materials for generating warning messages. Simultaneously, parse the inspected equipment identifier, timestamp, violation type, etc., for subsequent severity assessment and initiation of the handling process.

[0079] S51. Includes a built-in severity rating table, which is based on the type of violation. The severity level of the current violation is calculated based on the context in which the violation occurred and preset weighting factors. The severity rating table defines three levels: minor violation, moderate violation, and severe violation. For example, "operation timeout" might be defined as a minor violation; "tool error" and "incorrect step sequence" might be defined as moderate violations; and "exceeding the threshold," if it significantly exceeds the safety range, might be defined as a severe violation. To achieve more precise evaluation, the system introduces a severity scoring function: ,in, These are the base scores for different types of violations. For example, a minor violation corresponds to 10 points, a general violation corresponds to 30 points, and a serious violation corresponds to 80 points. The penalty score is based on the number of historical violations. If the same device frequently commits the same type of violation recently, additional points will be added. This score is estimated based on the potential consequences of the current operation. For example, if the threshold is exceeded and the device is close to its limit, this score will be higher. These are preset weighting coefficients used to balance various factors. The total score is then calculated. Then, the system compares it with two preset thresholds. and Comparison: If If it is, it is judged as a minor violation; if If it is, it is judged as a general violation; if If so, it will be judged as a serious violation. The final severity level... It is attached to the verification result data packet for use in subsequent steps.

[0080] S52. Upon receiving a violation information indicating its severity level, the augmented reality rendering engine is invoked, selecting different warning information presentation methods based on the severity. For minor violations, the rendering engine displays a pale yellow semi-transparent warning box near the corresponding device component in the user's field of vision, containing text such as "Operation timed out, please speed up." The warning box fades out automatically after three seconds. For general violations, the rendering engine pops up a prominent orange warning panel in the center of the field of vision, displaying the violation type and correction suggestions. Simultaneously, a flashing red outline appears around the device component, accompanied by a brief vibration or beeping sound. The warning information continues to be displayed until the user confirms or the operation is corrected. For serious violations, the rendering engine immediately pauses the current operation guidance, the entire field of vision background turns semi-transparent red, and a large red prohibition symbol and the text warning "Serious violation, stop immediately!" are displayed in the center. At the same time, a high-intensity voice warning is issued through the speech synthesis module, such as "Pressure exceeded, danger!". All warning information is ensured to be precisely superimposed near the actual location of the violation, allowing the operator to visually identify the problem.

[0081] S53. According to severity level Different actions are taken depending on the violation. For minor violations, the system only records the violation event to the local log file and updates the operating statistics of the inspected equipment, without triggering external handling procedures. For general violations, the system automatically generates a defect work order, which includes the identifier of the inspected equipment. Time of violation Types of violations The system collects descriptions of violations and on-site photos or video clips. By calling the application programming interface (API) of the enterprise resource planning (ERP) or maintenance management system, the work order is pushed to the designated maintenance task queue and assigned a default responsible person. Simultaneously, the system sends a notification to the responsible person's mobile terminal, informing them of the new work order pending processing. For serious violations, in addition to generating and pushing the defective work order, the system immediately sends an emergency alert SMS or application push to the mobile terminals of the duty supervisor and relevant safety personnel. The notification includes the equipment location, violation details, and a link to real-time on-site footage. Furthermore, the system triggers interlocking protection actions, such as forcibly suspending the operation of the inspected equipment via a programmable logic controller (PLC) signal until safety is confirmed. All actions are recorded, creating a traceable audit trail.

[0082] S54. Continuously monitor the status of triggered work orders. When the responsible maintenance personnel accept the order, arrive on-site to handle the issue, complete the repair, and close the work order, they will receive a status update callback. This status information is rendered into corresponding virtual prompts and displayed in the operator's view, such as "Work order assigned, maintenance personnel will arrive within 5 minutes." If the violation is resolved, the system can also prompt the operator through an augmented reality interface to continue with subsequent inspection steps. The entire process ensures that the entire process from violation occurrence to problem resolution is traceable and manageable, achieving seamless integration of operational compliance verification and maintenance management.

[0083] Augmented Reality (AR) interface refers to an interactive interface that uses an AR-enabled smart camera display device to overlay computer-generated virtual information onto the real physical world scene seen by the user in real time and accurately. Users can simultaneously observe the real equipment and related guidance, prompts, or annotations through this interface. Warning messages are alerts and prompts issued to operators via the AR interface in a visual, auditory, or tactile manner when the system detects a violation, aiming to remind operators to pay attention and correct erroneous behavior. A violation refers to an operator's actions, tool usage, operational values, or time intervals during the inspection or operation of the equipment being inspected, which do not conform to the predefined standards in the operational compliance rule base and are judged by the system as a breach of regulations. Severity refers to a quantitative assessment of the potential consequences, risk level, and urgency of the violation, typically categorized as minor, moderate, and severe, used to determine subsequent warning methods and handling intensity. The defect handling process refers to a series of standardized steps performed from the discovery of a violation until the problem is finally resolved, including management activities such as recording, assessment, creating a work order, notifying the responsible party, tracking the processing progress, and closing the work order. Enumerated values ​​categorize violations, describing their specific nature, such as incorrect step sequence, tool error, exceeding thresholds, or operation timeouts, providing a basis for severity assessment. The severity grading table is a lookup table defining the mapping relationship between different violation types and base scores, as well as the historical penalty factors and influence factor weights required to calculate the final severity level. A defect work order is an electronic document automatically generated by the system containing detailed information about a violation event, used to notify maintenance personnel and record the entire maintenance process; it is a standard work unit in enterprise resource planning or maintenance management systems. An application programming interface (API) is a standardized interface for data exchange and function calls between the system and external enterprise software. It allows defect work orders to be pushed to the maintenance management system or for querying work order processing status.

[0084] Before outputting warning information through the augmented reality interface, a spatial calibration step is also included for the displayed content of the augmented reality interface, specifically including: S050. Initial positioning is achieved by identifying fixed markers on the equipment under inspection to obtain coarse pose information; three-dimensional point cloud data of the equipment under inspection is acquired through a depth sensor, and the three-dimensional point cloud data is iteratively registered with the nearest point of the CAD model of the equipment under inspection to obtain fine pose information; the virtual information displayed in the augmented reality interface is dynamically corrected based on the fine pose information to ensure that the virtual information and the equipment under inspection are aligned at the sub-centimeter level.

[0085] S0500. Initial positioning is achieved by identifying fixed markers on the inspected equipment to obtain coarse pose information. In this step, fixed markers refer to marks with specific patterns, such as QR codes or feature graphics, pre-set on the surface of the inspected equipment. The AR inspection smart camera acquires an image containing the fixed marker through its visible light camera. The image recognition algorithm integrated into the system analyzes the image, detects, and decodes the fixed marker. Since the size, shape, and precise position of the fixed marker in the world coordinate system are known in advance, the pose of the camera relative to the fixed marker, i.e., the rotation matrix and translation vector, can be calculated by solving the correspondence between the two-dimensional projection of the fixed marker in the image plane and its three-dimensional world coordinates. This calculated pose information is the coarse pose information, which provides an initial estimate close to the true value for subsequent fine calibration, ensuring the convergence speed and accuracy of the spatial calibration process.

[0086] S0501. Acquire 3D point cloud data of the equipment under inspection using a depth sensor. Iteratively register the 3D point cloud data with the CAD model of the equipment under inspection using the nearest neighbor method to obtain fine pose information. After obtaining coarse pose information, the system activates the depth sensor. The depth sensor emits light waves towards the equipment under inspection and receives the reflected signals. By measuring information such as the time of flight of the light waves or structured light distortion, the 3D spatial coordinates of each sampling point on the surface of the equipment are calculated. This massive set of coordinate points constitutes the 3D point cloud data of the equipment under inspection. Simultaneously, the system pre-stores a precise computer-aided design model (CAD model) of the equipment under inspection. Iterative nearest neighbor registration is an algorithm used to align two sets of point clouds. In this step, the 3D point cloud data serves as the source point cloud, and the target point cloud generated by dense sampling from the CAD model of the equipment under inspection serves as the reference point cloud. In each iteration, the algorithm finds the nearest neighbor point in the target point cloud for each point in the source point cloud, forming corresponding point pairs. Based on these corresponding point pairs, a rigid body transformation is calculated by minimizing an error function. This rigid body transformation includes a rotation matrix R and a translation vector t. The error function is typically defined as the sum of squared Euclidean distances between all corresponding point pairs after the transformation, i.e.: ,in, This represents a point in a 3D point cloud dataset. Indicates and The nearest neighbor point on the CAD model of the corresponding equipment being inspected. This represents the total number of corresponding point pairs. This step applies the calculated rigid body transformation to the source point cloud, and then repeats the process of finding the nearest neighbor and calculating the rigid body transformation. The algorithm stops when the error function converges below a preset threshold or when the number of iterations reaches its upper limit. The accumulated rigid body transformation result at this point, that is, the precise transformation relationship from the 3D point cloud data coordinate system to the CAD model coordinate system of the inspected equipment, is the fine pose information. This fine pose information achieves sub-centimeter-level precise alignment between the camera coordinate system and the CAD model coordinate system of the inspected equipment.

[0087] S0502. Based on the fine pose information, the virtual information displayed in the augmented reality interface is dynamically corrected to ensure sub-centimeter alignment between the virtual information and the inspected device. After obtaining the fine pose information, the system accurately knows the relative position and orientation of the AR inspection smart camera and the inspected device in three-dimensional space. Virtual information, such as indicator arrows for instrument readings, text prompts for operation steps, and perspective overlays of the device's internal structure, is pre-defined in the coordinate system of the inspected device's CAD model during its content creation stage. When rendering the augmented reality interface, a projection matrix is ​​constructed from the three-dimensional world coordinate system where the virtual information is located to the camera's imaging plane using the calculated fine pose information. Through this projection matrix, all virtual information is accurately drawn on the display screen at the position corresponding to the physical device. When the user moves while wearing the AR device or the device itself undergoes slight deformation, the depth sensor continuously acquires new three-dimensional point cloud data and performs a new round of iterative nearest-point registration with the CAD model of the inspected device, updating the fine pose information in real time. This process ensures that the virtual information always dynamically adjusts its display position to follow the changes in the physical device's position, thereby achieving continuous and stable sub-centimeter-level alignment between the virtual information and the actual device being inspected in the augmented reality interface.

[0088] In this context, fixed markers refer to patterned marks with specific geometric shapes and color contrasts pre-set on the surface of the device being inspected. These marks serve as a spatial positioning reference; their size, shape, and precise position in the world coordinate system are known. They are used by the vision system of the AR inspection smart camera for rapid identification and capture, allowing for the calculation of the camera's initial spatial position and orientation relative to the marker. Initial positioning refers to the process of preliminarily determining the relative spatial relationship between the AR inspection smart camera and the device being inspected by identifying fixed markers. This process relies on the detection and pose calculation of known markers in a two-dimensional image. The result is a coarse pose information that approximates the true value but has limited accuracy, providing initial values ​​for subsequent, more precise registration steps. The coarse pose information refers to data calculated through the initial positioning steps that describes the spatial position and orientation of the AR inspection smart camera relative to the device being inspected. This information is typically represented in the form of a rotation matrix and translation vector. Its accuracy is sufficient to meet the convergence requirements of subsequent fine registration algorithms, but it has not yet reached the sub-centimeter accuracy required for precise information overlay in augmented reality interfaces. A depth sensor is a device capable of actively acquiring three-dimensional spatial information of objects in a scene. It calculates the distance from the sensor to various points on the device surface by emitting light of a specific wavelength into the device being inspected and analyzing the reflected signals, thus generating a dataset consisting of a large number of points with three-dimensional coordinates. The three-dimensional point cloud data is a massive set of points described by a depth sensor, depicting the spatial position and shape of the device surface. Each point contains its coordinate information in three-dimensional space, and the set of these points constitutes a discrete geometric representation of the device surface in the camera coordinate system, serving as the raw data for subsequent fine registration. Iterative nearest neighbor registration is a mathematical optimization algorithm used to optimally align two sets of point cloud data. This algorithm iteratively optimizes the alignment between the two sets of point clouds by repeatedly iterating through two main steps: finding the nearest neighbors of points in the source point cloud in the target point cloud, and calculating the rigid body transformation that minimizes the sum of the distances between these corresponding point pairs. Ultimately, it obtains an accurate coordinate transformation parameter. Fine-grained pose information refers to the data describing the precise spatial relationship between the AR inspection smart camera and the inspected device, calculated by precisely aligning the 3D point cloud data acquired by the depth sensor with the CAD model of the inspected device through an iterative nearest-neighbor registration algorithm. This information includes high-precision rotation matrices and translation vectors, forming the basis for accurate overlay of virtual information. Virtual information refers to non-physical digital content generated and overlaid on the augmented reality interface to assist in inspection operations. This content can include text, icons, 3D models, dynamic arrows, etc., and its content, position, and display method are dynamically generated based on the precise 3D spatial position and state of the inspected device. Dynamic correction is the process of continuously adjusting the display position and orientation of virtual information in the augmented reality interface based on the real-time calculated fine-grained pose information.This process ensures that virtual information remains precisely attached to its corresponding real-world physical device components, maintaining a visually seamless integration effect, even when the user moves or the environment undergoes subtle changes. Sub-centimeter alignment refers to a spatial positional error of less than one centimeter between the virtual information's display position in the augmented reality interface and its corresponding real-world inspected device component. This high-precision alignment is a prerequisite for effective information guidance and accurate status labeling, guaranteeing the usability and reliability of the augmented reality system.

[0089] Compared with the prior art, the present invention has the following beneficial effects: 1) Through the adaptive startup and data fusion acquisition of the visible light camera, infrared camera and depth sensor integrated in the AR inspection smart camera (i.e. multimodal camera), the interference of complex lighting and special materials is effectively overcome, so that the AR inspection smart camera can stably output high-quality target image data in various industrial scenarios, and ensure that the lightweight AI model can perform real-time recognition of target image data with a high accuracy of more than 92%.

[0090] 2) Based on the built-in operation compliance rule base, the operation compliance verification function transforms the inspection of the equipment to be inspected from "post-inspection" to "in-process intervention", which can proactively prevent misoperation. It has been verified that the misoperation rate can be reduced by 75%.

[0091] 3) The lightweight AI model enables high-performance recognition capabilities to be pushed to edge computing devices, reducing reliance on cloud computing and network latency; the accurate lightweight AI model significantly reduces reliance on professionals for real-time recognition of target image data and compliance verification of operations, reducing manual review costs by 60%.

[0092] 4) It provides a complete defect handling process that triggers corresponding actions based on the severity of violations, directly transforming technical findings into management actions and improving the overall efficiency and reliability of the maintenance of the inspected equipment.

[0093] In the above embodiments, although the steps are numbered S1, S2, etc., they are only specific embodiments given by the present invention. Those skilled in the art can adjust the execution order of S1, S2, etc. according to the actual situation. The scheme after adjusting the order is also within the protection scope of the present invention. It can be understood that in some embodiments, some or all of the above embodiments may be included.

[0094] like Figure 2 As shown in the figure, an AR inspection smart camera control and AI recognition system according to an embodiment of the present invention includes a monitoring module, a trigger acquisition module, a recognition module, a compliance verification module and an alert module; The monitoring module is used to: monitor in real time the light intensity in the current environment where the inspected equipment is located and the reflective properties of the target material of the inspected equipment through the environmental perception unit of the AR inspection smart camera; The trigger acquisition module is used to: adaptively activate the corresponding acquisition module from the visible light camera, infrared camera and depth sensor integrated in the AR inspection smart camera according to the preset switching trigger conditions and based on the light intensity and reflection characteristics, so as to acquire the target image data of the inspected device. The recognition module is used to: perform real-time recognition of target image data using a lightweight AI model deployed on edge computing devices, and obtain equipment status information and operation behavior information of the inspected equipment; The compliance verification module is used to: verify the compliance of operational behavior information based on the built-in operational compliance rule library; The alert module is used to: output alert information through the augmented reality interface if the verification result is a violation, and trigger the corresponding defect handling process according to the severity of the violation.

[0095] Optionally, in the above technical solution, the trigger acquisition module is specifically used to: compare the light intensity with a preset low light threshold; if the light intensity is lower than the low light threshold, activate the infrared camera as the acquisition module; acquire the reflection characteristics of the target material; when the target material is determined to be a metallic reflective material based on the reflection characteristics, activate the visible light camera and the depth sensor to perform data fusion acquisition; if the light intensity is higher than the low light threshold and the target material is a non-metallic reflective material, activate the visible light camera as the acquisition module.

[0096] Optionally, the above technical solution also includes a model acquisition module, which is used for: The original AI model is pruned using a structured approach based on channel importance assessment to remove redundant neurons and obtain the pruned model. The pruned model is subjected to 8-bit integer quantization, which compresses the model weights and activation values ​​from 32-bit floating-point data to 8-bit integer data, resulting in a lightweight AI model.

[0097] Optionally, the above technical solution also includes a spatial calibration module, which is used to: perform initial positioning by identifying fixed markers on the device under inspection to obtain coarse pose information; acquire three-dimensional point cloud data of the device under inspection through a depth sensor, and perform iterative nearest-point registration between the three-dimensional point cloud data and the CAD model of the device under inspection to obtain fine pose information; and dynamically correct the virtual information displayed in the augmented reality interface based on the fine pose information to ensure that the virtual information and the device under inspection are aligned at the sub-centimeter level.

[0098] It should be noted that the beneficial effects of the AR point-of-sight intelligent camera control and AI recognition system provided in the above embodiments are the same as those of the AR point-of-sight intelligent camera control and AI recognition method described above, and will not be repeated here. Furthermore, the system provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the system can be divided into different functional modules according to the actual situation to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments, and will not be repeated here.

[0099] An electronic device according to an embodiment of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any of the above-mentioned AR point inspection smart camera control and AI recognition methods.

[0100] An embodiment of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the above-described AR point-of-sight intelligent camera control and AI recognition methods.

[0101] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.

[0102] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for controlling and recognizing an AR-based intelligent inspection camera using AI, characterized in that, include: The AR inspection smart camera's environmental perception unit monitors in real time the light intensity in the current environment where the inspected device is located and the reflective properties of the target material of the inspected device. Based on the preset switching trigger conditions and the light intensity and reflection characteristics, the corresponding acquisition module is adaptively activated from the visible light camera, infrared camera and depth sensor integrated in the AR inspection smart camera to acquire the target image data of the inspected device. The target image data is identified in real time by a lightweight AI model deployed on an edge computing device to obtain the device status information and operation behavior information of the inspected device. Based on the built-in operation compliance rule library, the operation behavior information is verified for compliance. If the verification result is a violation, a warning message will be output through the augmented reality interface, and the corresponding defect handling process will be triggered according to the severity of the violation.

2. The AR point-of-sight intelligent camera control and AI recognition method according to claim 1, characterized in that, Based on preset switching trigger conditions and the light intensity and reflection characteristics, the corresponding acquisition module is adaptively activated from the visible light camera, infrared camera, and depth sensor integrated in the AR point-of-sight smart camera to acquire target image data. Specifically, this includes: comparing the light intensity with a preset low light threshold; if the light intensity is lower than the low light threshold, activating the infrared camera as the acquisition module; acquiring the reflection characteristics of the target material; if the target material is determined to be a metallic reflective material based on the reflection characteristics, activating the visible light camera and the depth sensor for data fusion acquisition; if the light intensity is higher than the low light threshold and the target material is a non-metallic reflective material, activating the visible light camera as the acquisition module.

3. The AR point-of-sight intelligent camera control and AI recognition method according to claim 1, characterized in that, The process of acquiring a lightweight AI model includes: The original AI model is pruned using a structured approach based on channel importance assessment to remove redundant neurons and obtain the pruned model. The pruned model is subjected to 8-bit integer quantization to compress the model weights and activation values ​​from 32-bit floating-point data to 8-bit integer data, resulting in a lightweight AI model.

4. The AR point-of-sight intelligent camera control and AI recognition method according to claim 1, characterized in that, Before outputting warning information through the augmented reality interface, the process includes a step of spatial calibration of the displayed content of the augmented reality interface. Specifically, this includes: initial positioning by identifying fixed markers on the device being inspected to obtain coarse pose information; acquiring three-dimensional point cloud data of the device being inspected through the depth sensor, and performing iterative nearest-point registration between the three-dimensional point cloud data and the CAD model of the device being inspected to obtain fine pose information; and dynamically correcting the virtual information displayed in the augmented reality interface based on the fine pose information to ensure that the virtual information is aligned with the device being inspected at the sub-centimeter level.

5. An AR inspection intelligent camera control and AI recognition system, characterized in that, It includes a monitoring module, a trigger acquisition module, an identification module, a compliance verification module, and an alert module; The monitoring module is used to: monitor in real time the light intensity in the current environment where the inspected device is located and the reflective properties of the target material of the inspected device through the environmental perception unit of the AR inspection smart camera; The trigger acquisition module is used to: adaptively activate the corresponding acquisition module from the visible light camera, infrared camera and depth sensor integrated in the AR inspection smart camera according to the preset switching trigger conditions, based on the light intensity and the reflection characteristics, so as to acquire the target image data of the inspected device; The identification module is used to: perform real-time identification of the target image data using a lightweight AI model deployed on an edge computing device, and obtain the device status information and operation behavior information of the inspected device; The compliance verification module is used to: perform compliance verification on the operation behavior information based on the built-in operation compliance rule library; The warning module is used to: if the verification result is a violation, output warning information through the augmented reality interface, and trigger the corresponding defect handling process according to the severity of the violation.

6. The AR inspection intelligent camera control and AI recognition system according to claim 5, characterized in that, The trigger acquisition module is specifically used to: compare the light intensity with a preset low light threshold; if the light intensity is lower than the low light threshold, activate the infrared camera as the acquisition module; acquire the reflectivity of the target material; when the target material is determined to be a metallic reflective material based on the reflectivity, activate the visible light camera and the depth sensor to perform data fusion acquisition; if the light intensity is higher than the low light threshold and the target material is a non-metallic reflective material, activate the visible light camera as the acquisition module.

7. The AR inspection intelligent camera control and AI recognition system according to claim 5, characterized in that, It also includes a model acquisition module, which is used for: The original AI model is pruned using a structured approach based on channel importance assessment to remove redundant neurons and obtain the pruned model. The pruned model is subjected to 8-bit integer quantization to compress the model weights and activation values ​​from 32-bit floating-point data to 8-bit integer data, resulting in a lightweight AI model.

8. The AR inspection intelligent camera control and AI recognition system according to claim 5, characterized in that, It also includes a spatial calibration module, which is used to: perform initial positioning by identifying fixed markers on the device under inspection to obtain coarse pose information; acquire three-dimensional point cloud data of the device under inspection through the depth sensor, and perform iterative nearest-point registration between the three-dimensional point cloud data and the CAD model of the device under inspection to obtain fine pose information; and dynamically correct the virtual information displayed in the augmented reality interface according to the fine pose information to keep the virtual information and the device under inspection aligned at the sub-centimeter level.

9. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the AR point-inspection smart camera control and AI recognition method according to any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the AR point-inspection smart camera control and AI recognition method according to any one of claims 1 to 4.