A method and system for dynamically labeling dangerous states of AR glasses based on multi-source perception

Through the multi-source perception AR glasses dangerous state dynamic labeling method, combined with multi-sensor data fusion and collaborative synchronization technology, the problems of inaccurate hazard identification and inconsistent equipment coordination in industrial safety operations are solved, achieving accurate hazard identification and efficient team collaboration, and improving user experience and equipment endurance.

CN120543957BActive Publication Date: 2025-10-03FUJIAN RUIXIN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511049518.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-10-03
Estimated Expiration
2045-07-29

AI Technical Summary

Technical Problem

Existing technologies in industrial safety operations have problems such as one-sided hazard identification, inaccurate AR labeling, and poor coordination among multiple devices. This leads to missed detections and false detections when relying on a single sensor, static threshold assessments that are difficult to reflect dynamic risks, AR labeling that is prone to drift, information from multiple devices being out of sync, chaotic team collaboration, unreasonable allocation of computing resources, and poor scenario adaptability.

Method used

A multi-source perception dynamic labeling method for dangerous states of AR glasses is adopted. The environmental depth image, three-dimensional point cloud data and dangerous gas concentration values ​​are synchronously obtained through the depth camera, gas sensor, inertial measurement unit and low-light night vision module. The visual inertial odometry technology and the entropy-weighted iterative nearest point algorithm are combined to achieve multi-device collaborative synchronization and dynamically adjust the labeling style to adapt to user and environmental changes.

Benefits of technology

It achieves comprehensive accuracy in hazard identification, reduces the risk of missed detection or false detection, ensures accurate alignment of AR annotations with real scenes, improves efficient teamwork collaboration and user experience, and extends device battery life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543957B_ABST
    Figure CN120543957B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of augmented reality and industrial safety monitoring technology, and in particular to a method and system for dynamically labeling dangerous states of AR glasses based on multi-source perception. The method uses a depth camera, gas sensor, and other multi-source devices to synchronously collect depth images, gas concentration, and other data. An improved dangerous target detection model extracts rotationally invariant point cloud features, and multimodal data is integrated to generate dynamic danger levels. VIO-based spatial precision registration of annotations and real-world scenes is achieved, and adaptive annotation styles are generated based on danger levels. An entropy-weighted ICP algorithm is used to synchronize the annotation poses of multiple devices, and the transparency and size of annotations are dynamically optimized based on user fatigue values ​​and ambient lighting. The present invention improves the comprehensiveness of hazard identification and the accuracy of annotations, supports team collaboration, and is suitable for scenarios such as power inspections and hazardous chemical transportation, ensuring operational safety and efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of augmented reality and industrial safety monitoring, and specifically to a method and system for dynamically labeling dangerous states of AR glasses based on multi-source perception. Background Art

[0002] In industrial safety operations (such as power inspections and hazardous chemical transportation), existing technologies suffer from one-sided hazard identification, inaccurate AR annotation, and poor multi-device coordination. Specifically, reliance on a single sensor can lead to missed or false detections, static threshold assessments fail to reflect dynamic risks, AR annotations are prone to drift, and fixed styles are inadequate to adapt to the environment and user status; information from multiple devices is out of sync, disrupting team coordination; and improper allocation of computing resources leads to delays or short battery life, coupled with poor scenario adaptability.

[0003] Therefore, to address the above problems, a method and system for dynamic labeling of dangerous states of AR glasses based on multi-source perception are proposed. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for dynamically labeling dangerous states of AR glasses based on multi-source perception to solve the problems raised in the above background technology.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A method for dynamically labeling dangerous states of AR glasses based on multi-source perception includes the following steps:

[0007] S1. Synchronous acquisition of multi-source data: Through the depth camera, gas sensor, inertial measurement unit and low-light night vision module, the environmental depth image, 3D point cloud data, hazardous gas concentration value and six-degree-of-freedom posture information are synchronously acquired;

[0008] S2. Dangerous state fusion identification:

[0009] Input the depth image and 3D point cloud into the pre-trained dangerous target detection model, and output the target bounding box and category label;

[0010] Fusion of gas concentration values, equipment operation sound spectrum gradients, and point cloud temperature data to generate a dynamic hazard level score;

[0011] S3. Dynamic annotation space registration:

[0012] Calculate the three-dimensional coordinates of the virtual annotation in the world coordinate system based on visual inertial odometry technology;

[0013] Adaptively generate annotation styles based on hazard level scores, including color transparency, flashing frequency, and 3D arrow direction;

[0014] S4. Multi-device collaborative synchronization: Real-time synchronization of the position and status of annotations between multiple AR glasses is achieved through the entropy-weighted iterative closest point algorithm;

[0015] S5. Real-time optimization of annotations: Dynamically adjust the transparency and display size of annotations based on the user's eye fatigue value and ambient light intensity.

[0016] As a preferred solution, in step S2, the dangerous target detection model includes the following processing steps:

[0017] Rotation-invariant point cloud feature extraction:

[0018] For each point in the point cloud, with the origin of its local reference system as the center, calculate the weighted value of the spherical harmonic basis function of the point coordinates and the center point, and multiply it by the Gaussian radial basis function attenuation coefficient;

[0019] Sum the calculation results of all points according to the order of spherical harmonic basis functions to generate rotationally invariant eigenvectors;

[0020] Multimodal fusion decision:

[0021] The normalized gas concentration value is divided by the preset concentration threshold, and then the hyperbolic tangent function is taken and multiplied by the gas weight coefficient;

[0022] Calculate the first-order gradient norm of the device sound spectrum and multiply it by the sound weight coefficient;

[0023] Extract the difference between the maximum point cloud temperature and the preset temperature threshold, activate it through the linear rectification function, and then multiply it by the temperature weight coefficient;

[0024] The above three results are added together to generate a comprehensive risk score.

[0025] As a preferred solution, in step S3, space registration and dynamic annotation include:

[0026] Space coordinate mapping:

[0027] Using the transformation matrix from the camera to the world coordinate system, perform a reverse projection operation on the two-dimensional coordinates of the target in the depth image, and calculate the world three-dimensional coordinates of the annotation point in combination with the depth value;

[0028] Occlusion-aware transparency generation:

[0029] Calculate the absolute value of the dot product of the user's sight direction vector and the target direction vector, divide it by the product of the two vectors' moduli, and then multiply it by the sight weight coefficient to generate the first transparency component;

[0030] The inverse of the hazard score is multiplied by the hazard weight coefficient to generate a second transparency component;

[0031] The first transparency component is added to the second transparency component to generate the final annotation transparency.

[0032] As a preferred solution, in step S4, multiple devices collaboratively and synchronously adopt an entropy weighted registration method:

[0033] For the matching point clouds between the server and the client, calculate the information entropy of the neighborhood of each point;

[0034] Input information entropy into the inverse entropy function to generate point cloud weights;

[0035] With the goal of minimizing the sum of weighted distance squares, the rotation matrix and translation vector are solved.

[0036] As a preferred solution, in step S5, the real-time optimization of annotation includes size adjustment rules:

[0037] When the user's fatigue value exceeds the fatigue threshold, the ratio of the excess fatigue value to the baseline value is calculated, multiplied by the scaling factor and added to one to generate the fatigue size coefficient;

[0038] The light size coefficient is generated by taking the common logarithm of the ratio of the ambient light intensity to the reference light intensity and adding one.

[0039] Multiply the fatigue size factor, lighting size factor and the dimensioned base size to generate the final display size.

[0040] As an optimal solution, point cloud temperature detection adopts the thermal anomaly probability generation method:

[0041] Calculate the square of the difference between the temperature value at each location in the point cloud and the average temperature of the scene, and divide it by twice the temperature variance;

[0042] Take the opposite of the exponential operation of the above result;

[0043] Multiplying by the spatial Gaussian distribution function value and then dividing by the normalization constant generates a thermal anomaly probability map.

[0044] As a preferred solution, it also includes edge-cloud task scheduling:

[0045] Construct a decision function with the goal of optimizing device energy efficiency, and calculate the local execution energy consumption divided by the local energy efficiency coefficient;

[0046] Add the cloud execution flag multiplied by the cloud energy consumption divided by the cloud energy efficiency factor;

[0047] Solve the optimal execution flag while satisfying the task delay constraint.

[0048] As a preferred solution, in the hazardous chemicals transportation scenario:

[0049] Calculate the cosine of the angle between the camera's current heading vector and the preset standard heading vector;

[0050] If the angle is greater than fifteen degrees, a re-collection instruction is triggered.

[0051] As a preferred solution, in power inspection scenarios:

[0052] Traverse all device coordinates and calculate the product of each device's risk value and its coordinate vector;

[0053] Perform vector summation on the product results of all devices to obtain a sum vector; divide the sum vector by the modulus length of the sum vector to generate the next target navigation direction.

[0054] A system for dynamically labeling dangerous states of AR glasses based on multi-source perception, and a method for dynamically labeling dangerous states of AR glasses based on multi-source perception.

[0055] It can be seen from the technical solutions provided by the present invention that the method and system for dynamically labeling dangerous states of AR glasses based on multi-source perception have the following beneficial effects:

[0056] 1. Hazard identification is more comprehensive and accurate, reducing the risk of missed or false detections:

[0057] Advantages of multi-source data fusion: Integrating multimodal data such as depth images, point clouds, gas concentrations, sound spectra, and temperature, it overcomes the limitations of a single sensor. The rotation-invariant point cloud feature extraction algorithm solves the feature instability problem caused by perspective changes in traditional target detection. Combined with the thermal anomaly convolution kernel, it accurately locates high-temperature hazardous areas, improving the reliability of dangerous target identification.

[0058] Dynamic level assessment is more scientific: Quantifying the risk score through a multimodal fusion decision formula rather than relying on static thresholds can reflect dynamic changes in risk conditions (such as increased gas concentration or sudden temperature rise in equipment) in real time, avoiding misjudgments caused by fixed standards.

[0059] 2. Precise alignment of spatial annotations with real-world scenes improves AR interaction reliability:

[0060] High-precision spatial registration: The VIO-NDC mapping algorithm is used to convert the coordinates of dangerous targets from the image system to the world system. Combined with real-time updates of 6DOF poses, this ensures spatial alignment between annotations and real-world targets, resolving the "false" and "incorrect" labeling issues associated with drift in traditional AR annotation.

[0061] Occlusion-aware adaptive display: Dynamically adjusts transparency through occlusion-aware annotation generation formulas. When the target is obscured, transparency is reduced to avoid interference, while clarity is improved when the user is looking directly at the target, ensuring information visibility without affecting normal working field of view.

[0062] 3. Multi-device collaborative synchronization supports efficient teamwork:

[0063] Real-time pose consistency: An entropy-weighted ICP algorithm is used to synchronize the annotated poses of multiple AR glasses. The information entropy weights are used to highlight the matching contribution of stable feature points, ensuring that team members see completely consistent hazard information and avoiding coordination errors caused by information bias.

[0064] Conflict resolution mechanism: By integrating risk assessment results from multiple devices using confidence-weighted methods, it resolves conflicts in multi-source data and improves the consistency of team decision-making.

[0065] 4. Adaptive optimization of annotation styles to improve user experience and security:

[0066] Dynamically adapting to the user and environment: Incorporating a fatigue-light coupling model, the system automatically adjusts annotation size when the user is fatigued and optimizes transparency to avoid glare in bright light environments. This solves the problem of traditional AR annotations being "fixed in size and easily interfered with," ensuring clear and legible information.

[0067] Scenario-based style design: Customized annotation styles are designed for scenarios such as power inspections and hazardous chemical transportation. Combining risk navigation vectors with attitude tolerance recognition, this provides users with intuitive risk avoidance guidance, reducing operational complexity.

[0068] 5. Edge-cloud collaboration improves real-time performance and reduces device energy consumption:

[0069] Efficient load balancing and scheduling: Energy efficiency optimization algorithms dynamically determine whether tasks are executed at the edge or in the cloud, balancing computing loads while meeting latency requirements and extending the battery life of AR glasses.

[0070] Low-latency response: The core algorithm runs in real time at the edge, ensuring that the update frequency of hazard annotations is synchronized with scene changes, avoiding annotation lags caused by delays. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 This is a schematic flow chart of the steps of a method for dynamically labeling dangerous states of AR glasses based on multi-source perception in the present invention. DETAILED DESCRIPTION

[0072] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0073] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0074] like Figure 1As shown, an embodiment of the present invention provides a method for dynamically labeling dangerous states of AR glasses based on multi-source perception, including the following steps:

[0075] S1. Synchronous acquisition of multi-source data: Through the depth camera, gas sensor, inertial measurement unit and low-light night vision module, the environmental depth image, 3D point cloud data, hazardous gas concentration value and six-degree-of-freedom posture information are synchronously acquired;

[0076] S2. Dangerous state fusion identification:

[0077] Input the depth image and 3D point cloud into the pre-trained dangerous target detection model, and output the target bounding box and category label;

[0078] Fusion of gas concentration values, equipment operation sound spectrum gradients, and point cloud temperature data to generate a dynamic hazard level score;

[0079] S3. Dynamic annotation space registration:

[0080] Calculate the three-dimensional coordinates of the virtual annotation in the world coordinate system based on visual inertial odometry technology;

[0081] Adaptively generate annotation styles based on hazard level scores, including color transparency, flashing frequency, and 3D arrow direction;

[0082] S4. Multi-device collaborative synchronization: Real-time synchronization of the position and status of annotations between multiple AR glasses is achieved through the entropy-weighted iterative closest point algorithm;

[0083] S5. Real-time optimization of annotations: Dynamically adjust the transparency and display size of annotations based on the user's eye fatigue value and ambient light intensity.

[0084] In this embodiment, step S1 is to deploy a multimodal sensor array to achieve real-time, synchronous collection of various key physical quantities and scene information in the operating environment, providing comprehensive and accurate raw data support for subsequent hazard identification and dynamic labeling, and ensuring the consistency of multi-source data in time and space dimensions. The following are the detailed steps:

[0085] Step S1-1: Multimodal sensor array deployment and calibration:

[0086] Sensor selection and function allocation: Deploy specific sensor arrays based on the needs of industrial scenarios (such as power inspections and hazardous chemical transportation). Depth cameras are responsible for collecting 2D depth images and 3D point cloud data of the scene to construct the geometric structure of the environment. Gas sensors (such as electrochemical sensors or infrared sensors) are used to detect the concentration of specific gases (such as toxic and flammable gases) in the environment. IMUs (inertial measurement units, including accelerometers and gyroscopes) are used to obtain the motion parameters (acceleration and angular velocity) of AR glasses in real time. Low-light-level night vision modules enhance image acquisition capabilities in low-light environments, ensuring data validity at night or in dimly lit scenes.

[0087] Sensor internal and external parameter calibration: Perform internal parameter calibration on the depth camera to determine its parameters such as focal length and principal point coordinates, and eliminate the impact of lens distortion on the image; complete external parameter calibration of the depth camera and low-light night vision module through Zhang Zhengyou calibration method, obtain the relative pose relationship between the two (rotation matrix and translation vector), and ensure the spatial consistency of visual data; perform zero bias calibration and scale factor calibration on the IMU to reduce its measurement error; at the same time, through multi-sensor joint calibration, determine the spatial position relationship between gas sensors, IMUs, etc. and visual sensors, laying the foundation for subsequent spatial fusion of multi-source data;

[0088] Step S1-2: Multi-source data timestamp synchronization mechanism:

[0089] Hardware trigger synchronization: Use hardware synchronization trigger signals (such as pulse triggers) to connect the depth camera, IMU, and low-light night vision module to ensure that each sensor starts data acquisition under the same trigger signal. This reduces time synchronization errors at the hardware level and ensures that the data collected by different sensors are initially aligned in terms of timestamps.

[0090] Timestamp unification and interpolation compensation: All sensors are equipped with high-precision clock modules (such as GPS timing or local high-precision crystal oscillators) to timestamp each frame of collected data (depth images, point clouds, gas concentration values, IMU data, and night vision images) with a unified time base. For data with slight time deviations (such as uneven time intervals caused by the low sampling frequency of gas sensors), linear interpolation or motion-state-based predictive interpolation methods are used to map asynchronous data to a unified time axis, ensuring strict correspondence between multi-source data in the time dimension and meeting the time consistency requirements of subsequent multimodal fusion.

[0091] 6DOF pose time synchronization: The high-frequency motion data output by the IMU is time-aligned with the image data output by the depth camera. Combined with the timestamp matching mechanism of the VIO (visual inertial odometry) algorithm, this ensures that the calculated 6DOF (six degrees of freedom) pose data is time-synchronized with the corresponding depth image and point cloud data, ensuring the accuracy of spatial registration.

[0092] Step S1-3: Real-time acquisition and format conversion of multi-source data:

[0093] Distributed acquisition control: Design an acquisition control module based on a real-time operating system (such as RT-Thread or ROS2 Real-Time) to dynamically configure the acquisition parameters of each sensor (such as the resolution and frame rate of the depth camera and the sampling frequency of the gas sensor). This module also implements a multi-threaded concurrent processing mechanism to enable parallel acquisition of depth images, point clouds, gas concentrations, IMU data, and night vision images, avoiding data congestion caused by single-threaded processing.

[0094] Data format standardization and conversion: Convert the raw data output by different sensors into a unified standardized format; for example, convert the raw depth map output by the depth camera into point cloud data (calculate the three-dimensional coordinates through the camera intrinsic parameters and depth values); convert the analog signal output by the gas sensor into a digital concentration value (units such as ppm); convert the angular velocity and acceleration data of the IMU into motion vectors in a unified coordinate system; convert the image data of the low-light night vision module into a pixel format and resolution compatible with the depth image, facilitating subsequent multimodal data fusion and storage;

[0095] Step S1-4: Data quality assessment and dynamic adjustment:

[0096] Real-time data quality monitoring: Perform real-time quality assessments on collected multi-source data, including the noise level of depth images, the sparsity of point clouds, the stability of gas concentration data (whether there are jumps), the drift of IMU data, and the signal-to-noise ratio of night vision images. Set quality thresholds for each data type (e.g., valid point cloud percentage ≥ 90%, gas concentration data fluctuation ≤ 5%), and trigger an alert when data quality falls below the threshold.

[0097] Dynamic adjustment of acquisition parameters: Dynamically adjust the sensor acquisition parameters based on the data quality assessment results. For example, when the depth image noise is too large, the exposure time of the depth camera is automatically increased (within the allowed range of motion) or the noise reduction algorithm is enabled. When the signal-to-noise ratio of the night vision image is too low, the gain of the low-light night vision module is increased. When the gas concentration data fluctuates greatly, the sampling frequency of the gas sensor is increased and sliding window filtering is enabled to ensure that the collected data always meets the accuracy requirements of subsequent hazard identification and labeling.

[0098] In this embodiment, in step S2, the dangerous target detection model includes the following processing steps:

[0099] Rotation-invariant point cloud feature extraction:

[0100] For each point in the point cloud, with the origin of its local reference system as the center, calculate the weighted value of the spherical harmonic basis function of the point coordinates and the center point, and multiply it by the Gaussian radial basis function attenuation coefficient;

[0101] Sum the calculation results of all points according to the order of spherical harmonic basis functions to generate rotationally invariant eigenvectors;

[0102] Multimodal fusion decision:

[0103] The normalized gas concentration value is divided by the preset concentration threshold, and then the hyperbolic tangent function is taken and multiplied by the gas weight coefficient;

[0104] Calculate the first-order gradient norm of the device sound spectrum and multiply it by the sound weight coefficient;

[0105] Extract the difference between the maximum point cloud temperature and the preset temperature threshold, activate it through the linear rectification function, and then multiply it by the temperature weight coefficient;

[0106] The above three results were added together to generate a comprehensive risk score;

[0107] Point cloud temperature detection uses a thermal anomaly probability generation method:

[0108] Calculate the square of the difference between the temperature value at each location in the point cloud and the average temperature of the scene, and divide it by twice the temperature variance;

[0109] Take the opposite of the exponential operation of the above result;

[0110] Multiply by the spatial Gaussian distribution function value and then divide by the normalization constant to generate the thermal anomaly probability map;

[0111] Furthermore, step S2 aims to achieve accurate identification and dynamic level assessment of hazardous targets in industrial scenarios by fusing multimodal perception data. On the one hand, this improves the robustness of target detection by extracting rotation-invariant features from point clouds. On the other hand, it fuses physical quantities such as gas, sound, and temperature to generate a quantitative hazard level, providing accurate hazard information input for subsequent spatial annotation. The detailed steps are as follows:

[0112] Step S2-1: Dangerous target detection based on rotation-invariant features:

[0113] Point cloud and depth image preprocessing: De-noise (using statistical filtering to remove outliers) and downsample (voxel grid filtering to retain key features) point cloud data collected from multiple sources, and spatially align the depth image with the point cloud (calculating 3D coordinates using camera intrinsics and depth values). Preset target category labels (e.g., "high-temperature component," "leak source," "obstacle") are provided for power inspection scenarios (e.g., utility poles and transformers) and hazardous chemical transportation scenarios (e.g., storage tanks and valves).

[0114] Rotation-invariant point cloud feature extraction: An improved feature extraction algorithm is used to calculate the rotation-invariant feature vector of the point cloud. The formula is as follows:

[0115] ,in, is the rotation invariant eigenvector (dimension is ); is the order index of the spherical harmonic function; is the highest order of spherical harmonics; is the point index in the point cloud; The total number of points in the point cloud that participate in feature calculation; are the weight parameters that can be learned during model training; is a 3rd order spherical harmonic basis function (satisfying ); Point cloud The three-dimensional coordinate vector of the point; is the three-dimensional coordinate vector of the origin of the local reference system; for point With the origin The Euclidean distance between is the Gaussian kernel width (the default value is 0.1m); is the natural exponential base;

[0116] This feature can resist the change of viewing angle of the point cloud caused by device rotation, ensuring the consistency of features of targets such as "telephone poles" and "valves" at different observation angles;

[0117] Dangerous target bounding box output:

[0118] The rotation invariant eigenvector The image is then fused with visual features (such as edges and textures) from the depth image and fed into an improved Faster R-CNN model. The model generates candidate boxes using a region proposal network (RPN), which are then classified and regressed based on preset target category labels. The model ultimately outputs a 2D bounding box (in the image coordinate system) and a 3D bounding box (in the point cloud coordinate system) for the dangerous target, labeling the target category (e.g., "overheated transformer," "toxic gas leak source").

[0119] Step S2-2: Point cloud temperature detection based on thermal anomaly convolution kernel:

[0120] Temperature data space mapping:

[0121] The temperature data collected by the infrared sensor is spatially associated with the point cloud, and each temperature measurement value is matched to the corresponding three-dimensional coordinate in the point cloud through coordinate transformation ( ), generate point cloud data with temperature attributes ( , is the coordinate of the point cloud on the two-dimensional projection surface);

[0122] Thermal anomaly probability map generation:

[0123] The thermal anomaly convolution kernel is used to calculate the probability distribution of temperature anomaly areas. The formula is as follows:

[0124] ,in, Point cloud position Thermal anomaly probability map at ; is a normalization constant (ensuring that the sum of the probabilities of all positions in the probability map is 1); is the natural exponential function; For point cloud at position The temperature value at is the average temperature of all point clouds in the current scene; is the variance of the temperature of all point clouds in the current scene; is the square difference between the temperature value and the mean; is the spatial Gaussian kernel function ( is the standard deviation of the Gaussian kernel, which is 0.2m); , is the coordinate of the point cloud on the two-dimensional plane;

[0125] For example: When the temperature at a certain point Much higher than the scene average hour, It approaches 1 (high abnormal probability), otherwise it approaches 0;

[0126] High temperature danger zone extraction:

[0127] Set the thermal anomaly probability threshold (such as 0.7) and The area is marked as a high temperature danger area, and the maximum point cloud temperature in the area is extracted , as input parameters for subsequent hazard level assessment;

[0128] Step S2-3: Hazard level assessment of multimodal physical quantity fusion:

[0129] Multimodal physical quantity standardization: standardize the collected gas concentration, sound spectrum, and temperature data:

[0130] Gas concentration: by formula Normalized to the interval [0,1] ( is the original concentration, is the minimum / maximum value of gas concentration in the scene);

[0131] Sound spectrum: Convert the sound signal into a spectrum through Fourier transform and calculate the L1 norm of its first-order gradient (Reflects the degree of spectrum mutation, such as the sharpness of abnormal noise of equipment);

[0132] Temperature deviation:

[0133] Combined with the extraction in step S2-2 ,pass Calculate the positive temperature deviation ( °C, only temperature contributions above a threshold are retained);

[0134] Quantitative calculation of risk score:

[0135] Generate risk score using multimodal fusion decision formula :

[0136] ,in, is the risk score (range 0-3); 、 、 is the weight coefficient of each modal feature (satisfying , such as hazardous chemicals scene , power scene ); is the hyperbolic tangent function (used to suppress extreme values, the output range is [-1,1]); is the normalized gas concentration value; The reference threshold value of gas concentration (such as the upper limit of safe concentration of such gas); is the ratio of the normalized concentration to the reference threshold; is the L1 norm of the first-order gradient of the sound spectrum (reflecting the degree of mutation of the sound); It is a linear rectification function (output is 0 when the input is negative, and output is the original value when the input is positive); is the maximum value of the point cloud temperature; is the temperature threshold (the value is 60°C); is the difference between the maximum temperature and the threshold;

[0137] For example: When toxic gas concentration is detected (near threshold), equipment noise spectrum gradient , When the temperature is normal, if , , ,but

[0138] (medium risk);

[0139] Dynamic hazard level classification:

[0140] Based on risk score The hazard level is divided into 4 levels:

[0141] Level 0 (non-hazardous): ;

[0142] Level 1 (low risk): ;

[0143] Level 2 (intermediate risk): ;

[0144] Level 3 (high risk): ;

[0145] The level is updated dynamically with the real-time collected data (e.g. when the gas concentration increases, Increase, level up);

[0146] Step S2-3: Verify the spatiotemporal correlation between dangerous targets and multimodal information:

[0147] Spatial correlation verification: For the dangerous target (such as the "leak source") detected in step S2-1, its three-dimensional bounding box coordinates are matched with the spatial distribution of multiple physical quantities:

[0148] Gas concentration: Calculate the spatial distance between the target center and the high-concentration area detected by the gas sensor. If the distance is ≤ 0.5m, it is considered as an association.

[0149] Temperature data: The temperature value within the target 3D bounding box is obtained through point cloud coordinates. Bind to the target (e.g. the "transformer" target is only associated with the temperature data of its internal point cloud);

[0150] Time series consistency check: Perform a time consistency check on three consecutive frames (0.1s interval per frame) of hazardous targets and multimodal data. If the change rate of the target category and hazard level is ≤20%, the target is considered to be in a stable state. Otherwise, the target is marked as "dynamically changing" (e.g., the concentration of the leakage source increases rapidly), and the data collection frequency is increased from 10Hz to 20Hz.

[0151] Step S2-4: Confidence evaluation and correction of recognition results:

[0152] Confidence calculation:

[0153] The IOU (intersection over union) of comprehensive target detection and the consistency of multimodal data (such as whether the gas concentration and temperature anomaly are synchronized) are used to calculate the confidence of the recognition results. ( is a consistency coefficient of 0-1); when , it is marked as a low confidence result;

[0154] Abnormal correction mechanism: For low confidence results, correction is performed in the following ways:

[0155] If the target detection bounding box is fuzzy, re-extract the rotation-invariant features and expand the detection range;

[0156] If multimodal data conflict (e.g., high gas concentration but no corresponding target), initiate point cloud resampling or sensor calibration;

[0157] Combined with historical data (such as common hazard types in the area), / / Temporary adjustment of weights to optimize risk scores ;

[0158] The final output: dangerous target category, 3D coordinates, bounding box, dynamic danger level (0-3) and confidence level, providing core input for subsequent spatial registration and dynamic annotation.

[0159] In this embodiment, in step S3, space registration and dynamic annotation include:

[0160] Space coordinate mapping:

[0161] Using the transformation matrix from the camera to the world coordinate system, perform a reverse projection operation on the two-dimensional coordinates of the target in the depth image, and calculate the world three-dimensional coordinates of the annotation point in combination with the depth value;

[0162] Occlusion-aware transparency generation:

[0163] Calculate the absolute value of the dot product of the user's sight direction vector and the target direction vector, divide it by the product of the two vectors' moduli, and then multiply it by the sight weight coefficient to generate the first transparency component;

[0164] The inverse of the hazard score is multiplied by the hazard weight coefficient to generate a second transparency component;

[0165] Adding the first transparency component and the second transparency component to generate a final annotation transparency;

[0166] Furthermore, step S3 accurately maps the identified dangerous target information to the real physical space and dynamically generates annotation styles based on the hazard level and environmental factors. This ensures that the annotations displayed by the AR glasses are spatially aligned with the actual scene in real time, while also improving the readability and accuracy of the annotations through adaptive adjustment. The detailed steps are as follows:

[0167] Step S3-1: World coordinate system positioning based on VIO:

[0168] 6DOF pose real-time calculation:

[0169] By fusing IMU data with the visual features of the depth image, the visual inertial odometry (VIO) algorithm is used to calculate the 6DOF pose (3 translation parameters + 3 rotation parameters) of the AR glasses in real time, and to construct the transformation matrix from the camera coordinate system to the world coordinate system. ;Update the pose once for each frame of image acquisition to ensure that the positioning frequency is synchronized with the data acquisition frequency (such as 10Hz);

[0170] Target coordinate world system transformation:

[0171] The VIO-NDC mapping algorithm is used to convert the two-dimensional image coordinates of the dangerous target into the world coordinate system coordinates. The formula is as follows:

[0172] ,in, is the three-dimensional coordinate of the dangerous target in the world coordinate system; is the transformation matrix from the camera coordinate system to the world coordinate system; is the camera inverse projection function (converting pixel coordinates and depth values ​​into three-dimensional coordinates in the camera coordinate system through the camera intrinsic parameters); is the two-dimensional pixel coordinate of the target in the image ; is the depth value corresponding to the pixel;

[0173] For example: When the target is located in the image , depth value , calculate its coordinates in the camera coordinate system by inverse projection, and then multiply it by Get world coordinates ;

[0174] Step S3-2: Spatial registration optimization of multi-source information fusion:

[0175] Point cloud auxiliary coordinate calibration:

[0176] Extract the point cloud data within the 3D bounding box of the dangerous target, match it with the 3D model of the known scene (such as the CAD model of power equipment; the point cloud map of the hazardous chemical storage tank), and calculate the coordinate deviation ( is the standard coordinate of the target in the model); if , then update through iterative optimization , reduce positioning error;

[0177] Dynamic drift compensation:

[0178] The short-term high-frequency data (100Hz) from the IMU is fused with the low-frequency precise pose of the VIO (10Hz). The pose changes between adjacent frames are predicted through Kalman filtering to compensate for the accumulated drift of the VIO. When the drift exceeds 0.1m, relocalization is triggered (combined with scene keyframe matching) to correct the world coordinate system coordinates.

[0179] Step S3-3: Generate adaptive label styles based on hazard levels:

[0180] Annotation visual element definition:

[0181] Design differentiated labeling styles based on hazard levels (0-3):

[0182] Level 0 (non-hazardous): no label;

[0183] Level 1 (low risk): solid green border, target category name (e.g., “slightly overheated pipe”), font size 12pt;

[0184] Level 2 (medium hazard): Yellow flashing border (frequency 2Hz), additional hazard parameters (such as "Temperature: 70°C"), font size 14pt;

[0185] Level 3 (High Risk): A red explosion-shaped animated border (with zooming animations), a superimposed warning icon (such as a skull or flame), and a hazard avoidance prompt (such as "Stay 5m away"), with a font size of 16pt.

[0186] The annotation anchor point is bound to the coordinate system:

[0187] The anchor point of the annotation is set to the geometric center of the dangerous target (calculated by a 3D bounding box) to ensure that the annotation always moves with the target; when the target moves (such as the shaking of a hazardous chemical container during transportation), the annotation is based on the real-time updated Dynamically adjust the anchor point position to maintain the spatial binding between the annotation and the target;

[0188] Step S3-4: Dynamic adjustment of occlusion-aware annotation transparency:

[0189] Calculation of sight and target direction vectors:

[0190] Collect user eye tracking data in real time (or estimate the gaze vector through head posture) (unit vector from the user's eyes to the gaze point), calculate the normal vector of the dangerous target surface (as the target direction vector);

[0191] Transparency adaptive formula applied:

[0192] Adjust the transparency of the annotation based on the relative angle between the sight line and the target. The formula is as follows:

[0193] ,in, The transparency of the annotation (value range is 0.2-0.9); is the sight weight (fixed at 0.6); is the dot product of two vectors; 、 is the vector modulus (all are 1 because they are unit vectors);

[0194] For example: When the user looks directly at the target (dot product close to 1) and hour, (higher transparency to avoid blocking target details); when the target is blocked (dot product close to 0) and hour, (Reduce transparency and reduce interference);

[0195] Step S3-5: Spatial registration accuracy evaluation and feedback correction:

[0196] Registration error is calculated in real time:

[0197] The spatial registration accuracy is evaluated by the following metrics:

[0198] The overlap between the annotation and the target (IOU): ≥ 0.8;

[0199] World coordinate deviation: ( is the real coordinate obtained by lidar);

[0200] Dynamic following error: When the target moves, the marking lag time is ≤50ms;

[0201] Error correction mechanism:

[0202] When the accuracy does not meet the standard, take the following corrective measures:

[0203] If the deviation is caused by VIO drift, use artificial landmarks in the scene (such as QR codes and reflectors) for relocalization;

[0204] If the target is lost due to occlusion, it is predicted based on the historical trajectory. , and a "Predicting" prompt will be displayed next to the annotation;

[0205] If the device's computing power is insufficient and causes delays, reduce the complexity of the annotation animation (for example, stop flickering and keep a static border);

[0206] Final output: Dynamic annotations precisely aligned with the real scene in the AR glasses' field of view, including target category, hazard level, real-time parameters, and risk avoidance prompts. The annotation style and transparency are adaptively adjusted according to the environment and user status, providing users with intuitive and accurate guidance on dangerous spaces.

[0207] In this embodiment, in step S4, multiple devices are collaboratively synchronized using an entropy-weighted registration method:

[0208] For the matching point clouds between the server and the client, calculate the information entropy of the neighborhood of each point;

[0209] Input information entropy into the inverse entropy function to generate point cloud weights;

[0210] With the goal of minimizing the sum of squared weighted distances, solve the rotation matrix and translation vector;

[0211] Furthermore, step S4 is used to achieve real-time synchronization and pose consistency calibration of hazard annotation information between multiple AR glasses devices, ensuring that the hazard annotations seen by all users in a team operation are completely consistent in spatial position and status parameters, avoiding coordination errors caused by information deviation between devices. The detailed steps are as follows:

[0212] Step S4-1: Collaborative communication network construction and data specification definition:

[0213] Low-latency communication module deployment:

[0214] Each AR glasses terminal is integrated with an industrial-grade wireless communication module (such as 5GSub-6GHz or Wi-Fi6E), supporting two communication modes: point-to-point direct connection between devices and edge server relay. For mobile scenarios (such as hazardous chemical transport fleets), a mesh self-organizing network topology is adopted to ensure automatic routing when any device goes offline, maintaining communication continuity.

[0215] Synchronous data format standardization:

[0216] Define a unified annotation synchronization data packet structure, including: annotation unique ID, world coordinate system coordinates , hazard level , timestamp , device ID, data confidence and checksum; among them, the coordinate data adopts floating point type (retain 6 decimal places), and the time stamp is accurate to millisecond level to ensure the integrity and traceability of data transmission;

[0217] Step S4-2: Multi-device pose benchmark unification and initial alignment:

[0218] World coordinate system anchor point sharing:

[0219] Select fixed feature points in the scene (such as the base of the tower in the power inspection and the corner of the hazardous chemical warehouse) as global anchor points, and pre-store their world coordinates through the edge server When each AR device is started, it uses the VIO algorithm to identify the anchor point, calculates the relative transformation between its own pose and the anchor point, and converts the local coordinate system to the global world coordinate system to eliminate the initial pose deviation of the device (the initial alignment error is required to be ≤0.1m);

[0220] Dynamic benchmark calibration mechanism:

[0221] For dynamic scenarios without fixed anchor points (such as a mobile transport fleet), the initial pose of the first device to be started is used as the reference, and its world coordinate system parameters are broadcast via Bluetooth. After receiving it, other devices convert their local coordinates to this reference system to achieve temporary reference unification. The reference parameters are rebroadcast every 30 seconds to compensate for accumulated drift.

[0222] Step S4-3: Entropy-weighted ICP algorithm to achieve real-time synchronization of annotation poses:

[0223] Point cloud feature matching data collection:

[0224] Each device collects point cloud data within the local field of view in real time, and extracts key feature points around the dangerous annotation (such as the vertices of the annotation bounding box and the geometric feature points of the target surface) as the matching source for posture synchronization; the feature point coordinates are converted into (Server / Master) and (Client / slave device) Package and upload to edge server or direct point-to-point transmission;

[0225] Entropy weighted pose optimization calculation:

[0226] The entropy-weighted ICP algorithm is used to solve the optimal rotation matrix With translation vector , to achieve spatial alignment of the feature points of the master and slave devices, the formula is as follows: , ,in, The coordinates of the main device feature points, The coordinates of the feature points corresponding to the slave device; For the The weight of a feature point is determined by its neighborhood point set Information entropy Decision - the lower the entropy value (the more stable the feature), the greater the weight; The square of the Euclidean distance is calculated iteratively (the convergence threshold of each iteration is ≤ 0.01m) to output the posture correction of the slave device relative to the master device. ; To solve the problem of minimizing the objective function and ;

[0227] Real-time correction of annotation coordinates:

[0228] After receiving the pose correction from the device, all the world coordinates of the local annotations are Performing the transformation , achieving spatial alignment with the main device annotation; the synchronization frequency is consistent with the VIO pose update frequency (10Hz), ensuring dynamic annotation without lag;

[0229] Step S4-4: Incremental synchronization and conflict resolution mechanism:

[0230] Incremental data transfer optimization:

[0231] Only when the annotation information changes (such as the danger level increases from 2 to 3, or the coordinate offset is ≥0.05m) is the synchronous data packet sent, avoiding bandwidth usage caused by full data transmission. The changes are transmitted through differential coding compression, reducing the data volume by more than 60%.

[0232] Rules for arbitration of multi-source conflicts:

[0233] When multiple devices update the same annotation at the same time (for example, device A detects , device B detects ), the edge server starts conflict resolution:

[0234] Prioritize confidence Higher device data (such as , then adopt B's );

[0235] If the confidence level is close (difference ≤ 0.1), the formula Calculate the weighted fusion result, For equipment The detected danger level score, For equipment Detected hazard level score;

[0236] Conflict results are synchronized to all devices and marked with the "Fusion Update" label to ensure that users know the source of the data;

[0237] Step S4-5: Collaborative accuracy assessment and dynamic compensation:

[0238] Real-time monitoring of synchronization error:

[0239] The edge server periodically (every 2 seconds) extracts the same labeled coordinates of more than 3 devices and calculates the standard deviation ,in, is the number of devices, is the average coordinate; when When , it is determined that the synchronization accuracy is insufficient and the compensation mechanism is triggered;

[0240] Dynamic compensation strategy execution:

[0241] If the error is caused by communication delay, automatically promote high priority mark ( ) transmission priority, occupying 70% of the communication bandwidth;

[0242] If the error is due to insufficient feature matching, expand the point cloud feature extraction range (from 1m around the annotation to 2m) and increase the number of matching points;

[0243] In extreme cases (such as ), force all devices to re-perform anchor point alignment and reset the pose benchmark;

[0244] Final output: Danger labels on all AR glasses are synchronized at the sub-meter level (≤0.1m) in the world coordinate system, with the labeled hazard levels and update timestamps completely consistent, providing a unified spatial information benchmark for team collaborative inspections and emergency response.

[0245] In this embodiment, in step S5, the real-time optimization of annotations includes size adjustment rules:

[0246] When the user's fatigue value exceeds the fatigue threshold, the ratio of the excess fatigue value to the baseline value is calculated, multiplied by the scaling factor and added to one to generate the fatigue size coefficient;

[0247] The light size coefficient is generated by taking the common logarithm of the ratio of the ambient light intensity to the reference light intensity and adding one.

[0248] Multiply the fatigue size factor, the illumination size factor and the marked basic size to generate the final display size;

[0249] Also includes edge-cloud task scheduling:

[0250] Construct a decision function with the goal of optimizing device energy efficiency, and calculate the local execution energy consumption divided by the local energy efficiency coefficient;

[0251] Add the cloud execution flag multiplied by the cloud energy consumption divided by the cloud energy efficiency factor;

[0252] Solve the optimal execution flag under the condition of meeting the task delay constraint;

[0253] Furthermore, step S5 is to optimize the display properties (size, transparency, style) and computing resource allocation of AR annotations in real time, taking into account the user's physiological state, environmental factors, and scenario requirements. This ensures that the annotations are clearly legible while avoiding interference with user operations, ultimately improving the efficiency and safety of hazard information transmission. The following are the detailed steps:

[0254] Step S5-1: Real-time monitoring and quantification of user fatigue status:

[0255] Fatigue feature collection:

[0256] The AR glasses' integrated eye tracking module (sampling frequency: 30Hz) collects the user's blink rate (BlinkRate), pupil diameter change (Pupil Dilation), and fixation duration (Fixation Duration). The IMU sensor also captures head tremor amplitude (Head Tremor). These features are transmitted in real time to the local edge computing unit via Bluetooth.

[0257] Quantitative calculation of fatigue value:

[0258] Use weighted fusion algorithm to map multiple features into fatigue values (range 0-100%), the formula is:

[0259] ,in, is the normalized blink frequency (resting state is 0, fatigue threshold is 1); is the normalized pupil constriction rate; is the normalized fixation duration; is the normalized head jitter amplitude); when When the fatigue level reaches 0.05, it is judged as a significant fatigue state;

[0260] Step S5-2: Ambient light intensity assessment and standardization:

[0261] Light data collection: Use the light sensor on the front of the AR glasses (range 0-10000 lux) to collect the ambient light intensity in real time ,The sampling frequency is 10 Hz, and the data is filtered by sliding window (window size 5 frames) to remove instantaneous fluctuations;

[0262] Light intensity normalization:

[0263] Converts light values ​​to relative to the base light The standardized coefficient of lux is: (when lux is treated as 1 lux to avoid negative logarithmic values); for example: in strong light environment Lux, ;Dim environment Lux, ;

[0264] Step S5-3: Dynamic adjustment of dimension based on fatigue-illumination coupling model:

[0265] Basic dimension settings:

[0266] Preset basic dimensions Pixels (corresponding to the standard size under medium light and no fatigue), according to the hazard level Adjusted base value: low risk ( )hour Pixel, high risk ( )hour pixels, ensuring that the hazard level is positively correlated with the base size;

[0267] The coupled model is applied to calculate the final dimensions:

[0268] Calculate the final dimensions using the fatigue-light coupling model :

[0269] ,in, is the final size of the annotation (pixels); is the scaling factor; is a fatigue modifier (only if (Effective when); This is the lighting correction item (enlarge the size in dim environments and reduce it moderately in strong light);

[0270] Example: When 、 lux lux), ( pixels):

[0271] ; , so

[0272] Pixels (although tired in dim environments, the light is weak, so the size is slightly reduced to avoid occlusion);

[0273] Step S5-4: Secondary optimization of annotation transparency integrating multiple factors:

[0274] Basic transparency settings:

[0275] Occlusion-aware transparency based on step S3-4 (calculated by the angle between the sight line and the target), combined with the danger level Set base range: low risk 0.3-0.5, high risk 0.6-0.8, ensuring that high-risk targets are marked more prominently;

[0276] Fatigue and Lighting Transparency Correction:

[0277] Introducing correction factors :When the user is tired( )Reduce transparency( , i.e. more opaque), to avoid blur; strong light environment ( lux) improves transparency , avoid glare); ultimate transparency , and limited to the range of 0.2-0.9;

[0278] Step S5-5: Edge-cloud offload scheduling supports real-time optimization:

[0279] Task energy consumption and latency evaluation:

[0280] Evaluate the execution cost of edge and cloud for annotation optimization tasks (such as fatigue value calculation, size adjustment, and transparency correction): Edge energy consumption Low but limited computing power, cloud energy consumption High but strong computing power; measure task latency at the same time , ensuring that the maximum delay is not exceeded ms;

[0281] Energy efficiency optimization algorithm decision:

[0282] The edge-cloud load sharing scheduling formula is used to determine the task execution location:

[0283] , ,in, Indicates edge execution. Indicates cloud execution; is the energy efficiency coefficient of the edge device (e.g. 0.8), is the cloud energy efficiency coefficient (e.g. 0.6); Indicates a constraint; for example: when the edge can complete the task within 150ms and hour, ,The tasks are executed locally to ensure real-time performance;

[0284] Step S5-6: Dynamic adaptation of scenario-based annotation style:

[0285] Power inspection scenario optimization:

[0286] Combined with risk navigation vector ( ,in, is the direction vector of the next target; The number of devices to be inspected; For the Risk value of each device; is the coordinate of the current position in the world coordinate system), an arrow is superimposed next to the mark (pointing to the direction of high-risk equipment), and the size of the arrow changes with Synchronous adjustment, and the arrow flashes (frequency 1Hz) to strengthen the reminder when the user is tired;

[0287] Optimization of hazardous chemicals transportation scenarios:

[0288] When the camera optical axis deviates from the angle ( ,in, is the arccosine function; is a vector With vector The inner product of is the current direction vector of the camera; is the preset standard orientation vector; is a vector Length of the module; is a vector When the module length is greater than 100), the annotation automatically adds a "viewing angle offset" warning box, enlarges the size by 20%, and reduces the transparency to 0.5 to ensure that users can quickly detect the annotation misalignment caused by viewing angle deviation;

[0289] Step S5-7: Real-time feedback of optimization effect and parameter iteration:

[0290] User interaction data collection:

[0291] Record the user's adjustments to the annotation (such as manual zooming in / out), the duration of gaze (the percentage of gaze on the annotated area), and evaluate the effectiveness of the annotation (if the gaze percentage is ≥ 60%, the annotation is considered effective);

[0292] Dynamic iteration of model parameters:

[0293] If the user is detected to have manually enlarged the annotation for 5 consecutive times, it means that the current size is too small, and the zoom factor is automatically increased. Temporarily increase to 0.6; if the user squints frequently under strong light, reduce the strong light transparency correction coefficient to , continuously optimize model parameters through feedback;

[0294] The final output: dynamically adapted hazard labels in the AR field of view—their size intelligently scales with fatigue and lighting, and their transparency adaptively adjusts to the scene and user status. Edge-cloud collaboration ensures low latency during the optimization process, providing users with "just right" hazard information prompts, balancing safety and operational convenience.

[0295] A multi-source sensing-based dynamic labeling system for AR glasses with dangerous conditions. This system is based on a multi-source sensing method for dynamically labeling dangerous conditions in AR glasses. By integrating multimodal sensor data with edge computing technology, it achieves real-time identification of dangerous targets in industrial scenarios, precise spatial labeling, and multi-device collaboration, providing immersive safety guidance for workers. The following is a detailed description of the system:

[0296] System architecture design:

[0297] The system adopts a three-layer architecture of "terminal perception-edge computing-cloud collaboration" to achieve rapid processing and intelligent labeling of dangerous information:

[0298] Terminal layer: AR glasses hardware that integrates depth cameras, infrared thermal imagers, gas sensors, and microphone arrays, responsible for multi-source data collection and preliminary preprocessing;

[0299] Edge layer: A lightweight AI computing unit deployed locally on AR glasses, which performs algorithms for dangerous target identification, spatial registration, and annotation optimization.

[0300] Cloud layer: Provides historical data storage, deep learning model training, and multi-device collaborative scheduling services, interacting with terminals via 5G / Wi-Fi;

[0301] Multi-source sensor fusion design:

[0302] The system integrates five types of heterogeneous sensor data to form complementary sensing capabilities:

[0303] Visual sensor: RGB camera (1920×1080@30fps) + depth camera (1280×720@30fps), providing scene geometry and texture information;

[0304] Temperature sensor: Infrared thermal imager (640×480@15fps), detects scene heat distribution and high temperature anomalies;

[0305] Gas sensor: Distributed electrochemical sensor array to detect the concentration of hazardous gases such as H2S, CO, and CH4 (response time <100ms);

[0306] Acoustic sensor: 4-microphone array, collects device operating sound (frequency response 20Hz-20kHz) and identifies abnormal noise;

[0307] IMU sensor: 6-axis inertial measurement unit (accelerometer + gyroscope), providing device posture information (sampling rate 100Hz);

[0308] Core algorithm module:

[0309] The system includes five core algorithm modules to achieve the perception, analysis and visualization of dangerous information:

[0310] Multimodal dangerous target recognition module:

[0311] A point cloud processing algorithm based on rotation-invariant features is used to extract the geometric features of the target;

[0312] A temperature detection algorithm that integrates thermal anomaly convolution kernels to identify high-temperature hazardous areas;

[0313] Multimodal physical quantity fusion decision-making algorithm, integrating gas concentration, sound spectrum and temperature data to assess hazard level;

[0314] Dynamic annotation space registration module:

[0315] VIO-NDC mapping algorithm, which realizes the precise conversion from image coordinates to world coordinates;

[0316] Point cloud-assisted spatial calibration algorithm improves the spatial accuracy of annotation;

[0317] Occlusion-aware transparency adjustment algorithm to optimize annotation visibility;

[0318] Multi-device collaborative synchronization module:

[0319] Entropy-weighted ICP algorithm to achieve precise alignment of annotation poses between multiple devices;

[0320] Incremental synchronization and conflict resolution mechanisms ensure data consistency and low latency;

[0321] Annotation real-time optimization module:

[0322] Fatigue-illumination coupling model, dynamic adjustment of annotation size;

[0323] Edge-cloud load-sharing scheduling algorithm balances computing efficiency and real-time performance;

[0324] Risk prediction and navigation module:

[0325] Risk evolution prediction algorithm based on spatiotemporal sequence analysis to provide early warning of potential dangers;

[0326] Risk navigation vector generation algorithm provides optimal inspection path guidance;

[0327] System workflow:

[0328] Data collection phase:

[0329] AR glasses simultaneously collect RGB images, depth maps, temperature distribution, gas concentration and sound data of the scene, and the IMU records the device's position;

[0330] Hazard identification stage:

[0331] The edge computing unit fuses and processes multi-source data: extracting rotation-invariant features from point clouds to identify target categories, detecting high-temperature areas using thermal anomaly convolution kernels, and assessing hazard levels using a combination of multiple physical quantities.

[0332] Space registration stage:

[0333] Based on VIO-NDC mapping, dangerous targets are mapped to the world coordinate system, the spatial position is calibrated through point cloud matching, and the annotation transparency is adjusted according to the occlusion situation;

[0334] Multi-device collaboration stage:

[0335] If multiple AR glasses are working together, the poses are synchronously annotated using the entropy-weighted ICP algorithm, and an incremental update mechanism is used to reduce communication overhead. In the event of conflict, the data is fused based on the confidence level.

[0336] Annotation optimization phase:

[0337] Real-time monitoring of user fatigue and ambient lighting, application of a fatigue-lighting coupling model to adjust annotation dimensions, and efficient optimization through edge-cloud offloading.

[0338] Visualization output stage:

[0339] AR glasses superimpose optimized hazard labels (including target category, hazard level, and dynamic parameters) onto the real scene, and users observe the enhanced environment through binocular lenses;

[0340] System innovations:

[0341] Multimodal hazard perception fusion:

[0342] Breaking through the limitations of traditional single-modality detection, it integrates visual, thermal imaging, gas, and acoustic data to achieve comprehensive identification of complex hazards in industrial scenarios.

[0343] Adaptive dynamic annotation:

[0344] Not only can the annotation style be adjusted according to the danger level, but it can also be optimized in real time based on the user's fatigue status and ambient lighting to ensure efficient information transmission;

[0345] High-precision spatial synchronization:

[0346] Through the entropy-weighted ICP algorithm and incremental update mechanism, sub-meter synchronization of annotations between multiple devices is achieved to meet the needs of team collaboration.

[0347] Edge-cloud intelligent offloading:

[0348] Design an energy-efficiency optimization scheduling algorithm to reduce energy consumption while ensuring system real-time performance, thereby extending the battery life of AR glasses;

[0349] Risk prediction and navigation:

[0350] Based on historical data and real-time monitoring, it can predict the evolution trend of dangers in advance and generate navigation paths, turning passive defense into active warning.

[0351] Application scenarios:

[0352] This system is suitable for high-risk industrial fields, including but not limited to:

[0353] Power inspection: Real-time marking of dangers such as transformer overheating and line damage;

[0354] Hazardous chemicals storage: Dynamically display the location of the leakage source and the spread of toxic gas;

[0355] Mining operations: marking risk areas such as roof falls and gas exceeding limits;

[0356] Fire rescue: guide firefighters to avoid high-temperature and toxic areas and plan safe routes;

[0357] Through this system, industrial site workers can obtain real-time and accurate danger warnings, significantly improving safety protection capabilities and work efficiency.

[0358] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for dynamically labeling dangerous states of AR glasses based on multi-source perception, characterized by: The following steps are involved: S1. Synchronous acquisition of multi-source data: Through the depth camera, gas sensor, inertial measurement unit and low-light night vision module, the environmental depth image, 3D point cloud data, hazardous gas concentration value and six-degree-of-freedom posture information are synchronously acquired; S2. Dangerous state fusion identification: Input the depth image and 3D point cloud data into the pre-trained dangerous target detection model, and output the dangerous target bounding box and category label; Fusion of gas concentration values, equipment operation sound spectrum gradients, and point cloud temperature data to generate a dynamic hazard level score; In step S2, the dangerous target detection model includes the following processing steps: Rotation-invariant point cloud feature extraction: For each point in the point cloud, with the origin of its local reference system as the center, calculate the weighted value of the spherical harmonic basis function of the point coordinates and the center point, and multiply it by the Gaussian radial basis function attenuation coefficient; The calculation results of all points are summed according to the order of the spherical harmonic basis function to generate a rotation-invariant feature vector. The rotation-invariant feature vector is fused with the visual features of the depth image and input into the improved Faster R-CNN model. The final output is the 2D bounding box and 3D bounding box of the dangerous target, and the target category is marked. Multimodal fusion decision: The normalized gas concentration value is divided by the preset concentration threshold, and then the hyperbolic tangent function is taken and multiplied by the gas weight coefficient; Calculate the first-order gradient norm of the device sound spectrum and multiply it by the sound weight coefficient; The temperature data collected by the infrared sensor is spatially associated with the point cloud. Each temperature measurement value is matched to the corresponding 3D coordinate in the point cloud through coordinate transformation to generate point cloud data with temperature attributes. The difference between the maximum temperature of the point cloud data and the preset temperature threshold is extracted and activated by the linear rectification function and then multiplied by the temperature weight coefficient. The above three results are added together to generate a hazard level score; S3. Dynamic annotation space registration: Calculate the three-dimensional coordinates of the virtual hazard marker in the world coordinate system based on visual inertial odometry technology; Adaptively generate annotation styles based on hazard level scores, including color transparency, flashing frequency, and 3D arrow direction; S4. Multi-device collaborative synchronization: Real-time synchronization of the location and status of hazard annotations between multiple AR glasses is achieved through the entropy-weighted iterative closest point algorithm; S5. Real-time optimization of annotations: Dynamically adjust the transparency and display size of annotations based on the user's eye fatigue value and ambient light intensity.

2. The method for dynamically labeling dangerous states of AR glasses based on multi-source perception according to claim 1 is characterized by: In step S3, space registration and dynamic annotation include: Space coordinate mapping: Using the transformation matrix from the camera to the world coordinate system, perform a reverse projection operation on the two-dimensional coordinates of the target in the depth image, and calculate the world three-dimensional coordinates of the annotation point in combination with the depth value; Occlusion-aware transparency generation: Calculate the absolute value of the dot product of the user's gaze direction vector and the target direction vector, divide it by the product of the moduli of the two vectors, and then multiply it by the gaze weight coefficient to generate a first transparency component, where the user's gaze direction vector is the gaze vector obtained by the user's eye tracking data, and the target direction vector is the normal vector of the dangerous target surface. The inverse of the hazard level score is multiplied by the hazard weight coefficient to generate a second transparency component; The first transparency component is added to the second transparency component to generate the final annotation transparency.

3. The method for dynamically labeling dangerous states of AR glasses based on multi-source perception according to claim 1 is characterized by: In step S4, the multi-device collaborative synchronization adopts the entropy weighted registration method: For the matching point clouds between the server and the client, calculate the information entropy of the neighborhood of each point; Input information entropy into the inverse entropy function to generate point cloud weights; With the goal of minimizing the sum of weighted distance squares, the rotation matrix and translation vector are solved.

4. The method for dynamically labeling dangerous states of AR glasses based on multi-source perception according to claim 1, characterized in that: In step S5, the real-time optimization of annotation includes size adjustment rules: When the user's fatigue value exceeds the fatigue threshold, the ratio of the excess fatigue value to the baseline value is calculated, multiplied by the scaling factor and added to one to generate the fatigue size coefficient; The light size coefficient is generated by taking the common logarithm of the ratio of the ambient light intensity to the reference light intensity and adding one. Multiply the fatigue size factor, lighting size factor and the dimensioned base size to generate the final display size.

5. The method for dynamically labeling dangerous states of AR glasses based on multi-source perception according to claim 1, characterized in that: The point cloud temperature detection adopts the thermal anomaly probability generation method: Calculate the square of the difference between the temperature value at each location in the point cloud and the average temperature of the scene, and divide it by twice the temperature variance; Take the opposite of the exponential operation of the above result; Multiplying by the spatial Gaussian distribution function value and then dividing by the normalization constant generates a thermal anomaly probability map.

6. The method for dynamically labeling dangerous states of AR glasses based on multi-source perception according to claim 1, characterized in that: Also includes edge-cloud task scheduling: Construct a decision function with the goal of optimizing device energy efficiency, and calculate the local execution energy consumption divided by the local energy efficiency coefficient; Add the cloud execution flag multiplied by the cloud energy consumption divided by the cloud energy efficiency factor; Solve the optimal execution flag while satisfying the task delay constraint.

7. The method for dynamically labeling dangerous states of AR glasses based on multi-source perception according to claim 1, characterized in that: In hazardous chemicals transportation scenarios: Calculate the cosine of the angle between the camera's current heading vector and the preset standard heading vector; If the angle is greater than fifteen degrees, a re-collection instruction is triggered.

8. The method for dynamically labeling dangerous states of AR glasses based on multi-source perception according to claim 1, characterized in that: In the power inspection scenario: Traverse all device coordinates and calculate the product of each device's risk value and its coordinate vector; Perform vector summation on the product results of all devices to obtain a sum vector; divide the sum vector by the modulus length of the sum vector to generate the next target navigation direction.

9. A dynamic labeling system for AR glasses dangerous states based on multi-source perception, characterized by: The system executes the method for dynamically labeling dangerous states of AR glasses based on multi-source perception as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Automatic labeling system and method for electric transmission line entering equipotential dangerous point

    CN119763116A

  • Assistant decision-making platform for water conservancy project operation and maintenance based on AI unmanned aerial vehicle

    CN119990627A