Safety helmet wearing and inspection track checking method based on fusion of AI multi-mode recognition
By integrating a panoramic camera and a real-time positioning module into the safety helmet, a spatial coordinate system is constructed. Combined with AI multi-modal recognition technology, the helmet wearing status and safety hazards are automatically identified, and a panoramic inspection trajectory report is generated. This solves the problems of information omission and task lag in existing technologies, and realizes automated and integrated management of safety inspections.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG JIUJIAN CONSTR GRP CO LTD
- Filing Date
- 2026-03-26
- Publication Date
- 2026-07-31
AI Technical Summary
In existing on-site safety inspections, single-vision recognition cannot simultaneously detect the wearing status of safety helmets and safety hazards. Inspection trajectories and safety hazard data are stored separately, leading to information omissions and coordinate parameter deviations. The authenticity of inspection trajectories lacks multi-source data verification, and rectification tasks rely on manual operation, resulting in problems such as delayed task assignment and information matching errors.
By employing a multi-modal AI-integrated recognition method, environmental images and location data are collected through a panoramic camera and real-time positioning module on the safety helmet. A spatial coordinate system is constructed, and the helmet wearing status and safety hazards are automatically identified. A panoramic inspection trajectory report is generated, and a rectification task sheet is automatically generated, realizing the spatiotemporal correlation and authenticity verification of the trajectory data.
It enables the simultaneous output of information on helmet wearing status, safety hazards, and rectification points, reducing manual annotation bias, forming a multi-dimensional identification result system, eliminating task assignment delays, improving the authenticity and integrated nature of inspection trajectories, and reducing information mismatch.
Smart Images

Figure CN122493381A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of engineering safety inspection technology, and in particular to a method for verifying the wearing of safety helmets and inspection trajectories by integrating AI multi-modal recognition. Background Technology
[0002] Current on-site safety inspections at construction sites primarily employ a combination of manual inspections and single-mode visual recognition equipment. These systems collect personnel location information using conventional positioning modules, relying solely on single-mode visual detection to identify helmet-wearing status or individual safety hazards. Inspection trajectories and hazard data are stored independently, requiring manual review of inspection records before issuance of rectification tasks. Trajectory reconstruction relies solely on basic positioning data, failing to integrate panoramic images to construct a spatial coordinate system. Furthermore, verification of trajectory coverage across floor-level inspection areas is entirely manual.
[0003] Single-mode visual recognition cannot simultaneously detect helmet wearing status, safety hazards, and areas requiring rectification. The coordinate marking of hazard locations and rectification points relies on manual operation, which is prone to information omissions and coordinate parameter deviations. Inspection trajectory data and safety recognition results cannot be integrated in the spatiotemporal dimension to form a unified panoramic inspection trajectory report. The authenticity of the inspection trajectory lacks cross-validation from multi-source data. The generation process of rectification task orders relies on manual operation, resulting in delayed task assignment and information matching errors. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a method for verifying helmet wearing and inspection trajectories by integrating AI multi-modal recognition.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a method for verifying helmet wearing and inspection trajectory by integrating AI multi-modal recognition, comprising: The on-site environment images and location data are collected by the panoramic camera and real-time positioning module integrated on the safety helmet. The on-site environment images and location data are processed to construct a spatial coordinate system to obtain spatial coordinates. Based on the spatial coordinates, the real-time movement data of the inspection personnel is processed to reconstruct the trajectory and obtain the inspection trajectory data. The inspection trajectory data is then compared with the preset floor plan inspection area data to generate the trajectory coverage result. AI multimodal recognition processing is performed on the on-site environment image to automatically identify the helmet wearing status, safety hazards, and rectification parts in the image, and to obtain the recognition result including status indicators, hazard locations, and rectification point coordinates; The identification results are spatiotemporally correlated with the inspection trajectory data to generate a panoramic inspection trajectory report. The authenticity of the inspection trajectory data is verified, and a rectification task sheet is automatically generated based on the verified panoramic inspection trajectory report.
[0006] As a further aspect of the present invention, the step of acquiring on-site environmental images and location data through a panoramic camera and real-time positioning module integrated on the safety helmet, and performing spatial coordinate system construction processing on the on-site environmental images and location data to obtain spatial coordinates includes: The panoramic camera integrated on the safety helmet is controlled to acquire a panoramic video stream of the scene at a preset frequency. The panoramic video stream of the scene is then subjected to image frame extraction and distortion removal processing to obtain a standardized environmental image. The original positioning data of the positioning module is synchronously acquired by the real-time positioning module integrated on the safety helmet. The original positioning data is then filtered and error compensation is calculated to obtain the corrected positioning data. Visual feature points are extracted and matched based on the standardized environmental image to construct a visual feature point map. The corrected positioning data is then fused and aligned with the visual feature point map to establish a global three-dimensional spatial coordinate system for the site. The coordinate transformation calculation is performed on the pixel positions in the standardized environmental image and the corrected positioning data using the on-site global three-dimensional spatial coordinate system to obtain the spatial coordinates corresponding to each frame of the image.
[0007] As a further aspect of the present invention, the real-time movement data of the inspection personnel is processed based on the spatial coordinates to perform trajectory reconstruction calculations to obtain inspection trajectory data. The inspection trajectory data is then compared with the preset floor plan inspection area data to generate a trajectory coverage result, including: The spatial coordinates are sorted and interpolated according to the time series to obtain a continuous position sequence of the inspection personnel. Path fitting calculation is performed based on the position sequence to obtain continuous and smooth inspection trajectory data. Call the preset floor plan electronic map, extract the boundary coordinates of the key areas to be inspected from the floor plan electronic map, and generate preset floor plan inspection area data; Perform geometric set operations on the set of spatial points covered by the inspection trajectory data and the preset floor plan inspection area data to calculate the area ratio of the inspection trajectory data within the preset floor plan inspection area data. The trajectory coverage rate is calculated based on the area ratio. When the trajectory coverage rate reaches a preset threshold, a result of complete trajectory coverage is generated; otherwise, a result of incomplete trajectory coverage is generated, and the result is used as the trajectory coverage result.
[0008] As a further aspect of the present invention, the on-site environmental image is subjected to AI multi-modal recognition processing to automatically identify the helmet wearing status, safety hazards, and rectification areas in the image, obtaining a recognition result including status indicators, hazard locations, and rectification point coordinates, including: The on-site environment image is preprocessed, including image size normalization, color space conversion and enhancement, to obtain the preprocessed image to be identified; A safety helmet wearing recognition model is constructed. The image to be recognized is input into the safety helmet wearing recognition model, and the safety helmet wearing recognition model outputs the bounding box coordinates of the person's head in the image and the corresponding wearing status classification result. The wearing status classification result includes wearing, not wearing, and incorrect wearing. The safety helmet wearing recognition model is built on a convolutional neural network framework. It uses standardized environmental images of engineering site inspection scenarios as training samples, marks the bounding box of the head area of the personnel in the samples and associates it with the wearing status label, and presets anchor boxes in combination with the pixel distribution characteristics of the personnel's head at the engineering site. After model training, a recognition model adapted to the on-site scenario is obtained. A safety hazard identification model is constructed. The image to be identified is input into the safety hazard identification model, and the safety hazard identification model outputs the category, bounding box coordinates, and confidence level of the safety hazard in the image. The categories of safety hazards include improper placement of items, obstruction of fire lanes, lack of protective facilities, and electrical violations. The safety hazard identification model is built on a deep learning framework. It uses standardized environmental images containing various safety hazards at the engineering site as training samples. The hazard areas in the samples are labeled with bounding boxes and associated with category labels. Anchor boxes are preset based on the pixel size and morphological features of the hazards. The model is trained to obtain an identification model that is adapted to the on-site scene. Based on the bounding box coordinates of the aforementioned safety hazard, the specific rectification area is located in the image to be identified, and the rectification point coordinates of the rectification area in the spatial coordinate system are calculated. The data of the wearing status classification results, the category and bounding box coordinates of the safety hazard, and the coordinates of the rectification point are integrated to generate an identification result that includes status identifier, hazard location and rectification point coordinates.
[0009] As a further aspect of the present invention, the image to be identified is input into a safety hazard identification model, and the safety hazard identification model outputs the category, bounding box coordinates, and confidence level of various safety hazards in the image, including: The convolutional neural network of the safety hazard identification model performs multi-layer convolution and pooling calculations on the input image to be identified to extract multi-scale feature maps of the image. Based on the multi-scale feature map, feature fusion calculation is performed in the feature pyramid structure to generate a fused feature map for target detection. Anchor boxes of a preset size are applied to the fused feature map, and regression and classification calculations are performed on the features within each anchor box to predict the bounding box coordinate adjustment parameters, category probability, and confidence score of potential hazard targets. Non-maximum suppression is applied to all predicted bounding boxes to filter out the final hazard detection results. The hazard detection results include the hazard category, bounding box coordinates, and corresponding confidence scores.
[0010] As a further aspect of the present invention, based on the bounding box coordinates of the safety hazard, the specific rectification area is located in the image to be identified, and the rectification point coordinates of the rectification area in the spatial coordinate system are calculated, including: Obtain the bounding box coordinates of a specific safety hazard from the identification results, and extract the corresponding hazard area image block from the image to be identified based on the bounding box coordinates; Key point detection processing is performed on the image block of the hidden danger area to identify the specific parts in the image block of the hidden danger area that need to be rectified, and the pixel coordinates of the parts to be rectified in the image block of the hidden danger area are obtained. Based on the bounding box coordinates of the hazard and the pixel coordinates of the rectification part in the image block of the hazard area, the overall pixel coordinates of the rectification part in the original image to be identified are calculated through coordinate transformation. By utilizing the mapping relationship between image pixel coordinates and three-dimensional spatial coordinates in the spatial coordinate system, the overall pixel coordinates of the rectification area are converted into three-dimensional coordinates in the global three-dimensional spatial coordinate system of the site, thus obtaining the coordinates of the rectification point.
[0011] As a further aspect of the present invention, the identification result and the inspection trajectory data are spatiotemporally correlated to generate a panoramic inspection trajectory report, including: Establish a time-aligned index to match and associate each frame of identification data in the identification results with the trajectory points of the same timestamp in the inspection trajectory data; The matched and associated data are integrated to form a spatiotemporal associated data chain that includes timestamps, spatial coordinates, helmet wearing status, identified safety hazard information, and coordinates of rectification points; Key information is extracted from the spatiotemporal correlation data chain, including the time, location, type, corresponding on-site environmental image, and coordinates of the rectification point where the hazard occurred, and then formatted and filled in according to the preset report template. The report template embeds keyframes of the on-site environment image and a visualization of the inspection trajectory data on the floor plan to generate a panoramic inspection trajectory report.
[0012] As a further aspect of the present invention, the inspection trajectory data is subjected to authenticity verification processing, including: The inspection personnel's dwell time and movement speed at each trajectory point are obtained from the inspection trajectory data, and the theoretical movement time between adjacent trajectory points is calculated. The calculated theoretical movement time is compared with the actual time interval of the data collection to obtain a time consistency verification value. If the time consistency verification value exceeds the error range threshold, the corresponding trajectory segment is marked as abnormal. A standard inspection path model is constructed, and the morphological characteristics of the inspection trajectory data are analyzed, including the smoothness of the movement path and the rationality of the direction change. The matching degree is calculated with the standard inspection path model to obtain the path rationality verification value. The standard inspection path model is based on a preset floor plan electronic map. It extracts the boundary coordinates of the key areas to be inspected and the coordinates of the necessary points. It plans the inspection route according to the standard operating procedures of the engineering site and combines the inspection sequence and movement constraints. By combining the time consistency check value and the path rationality check value, a trajectory authenticity check result is generated. Only when all check values are within the preset normal threshold range is the inspection trajectory data determined to be authentic and valid.
[0013] As a further aspect of the present invention, the automatic generation of a rectification task sheet based on the verified panoramic inspection trajectory report includes: The safety hazard information identified in the panoramic inspection trajectory report is analyzed, and the category, rectification point coordinates, discovery time and associated environmental images of each hazard are extracted. Based on the pre-defined mapping relationship between hazard categories and responsible departments for rectification, a corresponding responsible department for rectification is determined for each hazard, and the pre-defined rectification time limit rules are invoked to assign a required rectification completion time for each hazard; Fill in the corresponding fields of the standard rectification task sheet template with the category of the hazard, the coordinates of the rectification point, the time of discovery, the associated environmental image, the corresponding department responsible for rectification, and the required rectification completion time. Each generated rectification task order is assigned a unique task order number, and the rectification task order is sent to the corresponding rectification responsible department.
[0014] As a further aspect of the present invention, it also includes closed-loop management of rectification: Track and record the status changes of each rectification task order, including issued, received, under rectification, pending acceptance, and completed; When the department responsible for rectification uploads on-site images after rectification is completed, the AI recognition model is automatically invoked to review and identify the on-site images, verify whether the original hidden danger has been eliminated, and generate a review result. Update the status of the corresponding rectification task sheet according to the review results. If the review is passed, update the status to "completed" and generate a rectification completion record. Based on the rectification completion record and panoramic inspection trajectory report, a closed-loop management report is generated that includes the entire process of hazard discovery, rectification issuance, process tracking, and result verification.
[0015] Compared with the prior art, the advantages and positive effects of the present invention are as follows: AI multi-modal recognition processing is applied to on-site environmental images to simultaneously identify the helmet wearing status, safety hazards, and rectification areas in the images. The system generates recognition results that include status indicators, hazard locations, and rectification point coordinates. This eliminates the need for manual labeling of safety hazards and rectification points, avoids information omissions caused by manual labeling, and reduces deviations caused by coordinate parameter labeling. It allows for the simultaneous output of information on helmet wearing status, safety hazards, and rectification points, forming a multi-dimensional safety recognition result system. This makes the presentation of recognition information more aligned with the actual needs of on-site inspections.
[0016] The identification results, including status indicators, hazard locations, and rectification point coordinates, are spatiotemporally correlated with inspection trajectory data to generate a panoramic inspection trajectory report. The inspection trajectory data undergoes authenticity verification. Based on the verified panoramic inspection trajectory report, rectification task sheets are automatically generated. This eliminates the need for manual summarization of inspection records and manual compilation and issuance of rectification tasks, thus eliminating the lag in task assignment caused by manual operation and reducing mismatch issues in information transmission. The authenticity verification of the inspection trajectory is completed through cross-validation of multi-source data, forming an integrated form of inspection trajectory data and safety identification information, and realizing the linkage generation of inspection reports and rectification tasks. Attached Figure Description
[0017] Figure 1 This is a flowchart of the safety helmet wearing and inspection trajectory verification method integrating AI multi-modal recognition described in this invention; Figure 2 A flowchart for obtaining spatial coordinates; Figure 3 A flowchart for obtaining the recognition results; Figure 4 Time series analysis of helmet wearing status; Figure 5 Analysis diagram of the helmet wearing status of inspection personnel. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0019] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0020] See Figure 1 The system synchronously collects images and location data of the on-site environment through a panoramic camera integrated on the safety helmet and a real-time positioning module. The collected images and location data are processed to construct a spatial coordinate system, resulting in coordinate information with three-dimensional spatial meaning. Based on this spatial coordinate system, the real-time movement data of the inspection personnel is processed to reconstruct the trajectory, generating continuous inspection trajectory data. This trajectory data is then compared with pre-defined floor plan inspection area data to calculate the coverage ratio, generating a trajectory coverage result to assess whether the inspection range meets the standards. Simultaneously, the collected on-site environmental images undergo AI multi-modal recognition processing. A trained deep learning model automatically identifies the personnel's wearing status, the type of safety hazard, and the specific rectification area in the images, outputting recognition results including status indicators, hazard locations, and rectification point coordinates. These recognition results are then spatiotemporally correlated with the inspection trajectory data, matching and fusing the recognition information and location information at the same moment to generate a panoramic inspection trajectory report integrating time, space, status, and hazard information. The authenticity of the inspection trajectory data is verified by analyzing indicators such as the consistency of movement time and the rationality of path shape. The authenticity and validity of the trajectory are verified. Based on the panoramic inspection trajectory report after verification, the hidden danger information is automatically parsed to generate a rectification task sheet containing elements such as responsible department, rectification point, and time limit. This realizes an automated closed loop from inspection, discovery, recording to task issuance.
[0021] In one embodiment of the present invention, see [reference] Figure 2An integrated panoramic camera on a safety helmet captures a panoramic video stream from the inspector's perspective at a preset frequency of one frame per second. The raw video stream captured by the panoramic camera includes radial distortion introduced by the fisheye lens. In the image frame extraction and distortion correction process, a correction algorithm based on lens intrinsic parameters and distortion coefficients is applied to each frame. The lens intrinsic parameters and distortion coefficients are obtained through pre-calibration of the integrated panoramic camera. The calibration process uses a standard camera calibration method, calculating the panoramic camera's intrinsic parameter matrix and distortion coefficient vector by capturing multiple images of a calibration board with known spatial coordinate distributions. These calibration parameters are the prior parameters relied upon by the image frame extraction and distortion correction process and are called in the correction algorithm to achieve geometric correction of pixel coordinates. In specific implementation, image frame extraction and distortion correction process solves the pre-calibrated camera model, mapping each pixel in the original image to a distortion-free standardized image plane, resulting in a standardized environmental image free of barrel or pincushion distortion. This standardized environmental image serves as the foundational input for subsequent visual positioning and AI recognition. In practical implementation, the real-time positioning module employs a positioning system combining an inertial measurement unit (IMU) and an ultra-wideband (UWB) beacon. The acquired raw positioning data includes triaxial acceleration and triaxial angular velocity output from the IMU, as well as raw distance information measured by the UWB base station. These raw positioning data undergo filtering and error compensation calculations. A Kalman filter fuses the dead reckoning results from the IMU with the distance information measured by the UWB, compensating for errors caused by inertial device drift and environmental multipath effects. The final output is corrected positioning data with centimeter-level accuracy.
[0022] AI multi-modal recognition processing refers to the use of two parallel deep neural network models—a helmet-wearing recognition model and a safety hazard recognition model—to analyze and identify the same pre-processed frame of an on-site environment image. The helmet-wearing recognition model specifically analyzes the head region of individuals in the image, outputting a classification result indicating whether a helmet is worn: "worn," "not worn," or "incorrectly worn." The safety hazard recognition model is specifically used to detect and classify various preset categories of safety hazards in the image, such as improper placement of items, obstruction of fire exits, lack of protective facilities, and electrical violations, outputting the bounding box coordinates and confidence scores for each hazard. After recognition, the system further locates the specific parts requiring rectification based on the bounding box coordinates of the hazards and calculates their three-dimensional coordinates in the global spatial coordinate system—the rectification point coordinates—through coordinate transformation. Finally, the outputs of the two models are integrated with the calculated rectification point coordinates to generate a unified recognition result. This process achieves simultaneous and parallel recognition and localization of helmet-wearing status and multiple safety hazards in a single frame image.
[0023] Based on the preprocessed standardized environment image, visual feature point extraction and matching are performed. In the specific implementation, the visual feature point extraction and matching process adopts the scale-invariant feature transformation algorithm to extract key point feature descriptors with scale and rotation invariance from each frame of the standardized environment image. To address issues such as varying light levels, equipment occlusion, and repetitive building textures in engineering site environmental images, the scale-invariant feature transform algorithm was adapted to the engineering scene. The Gaussian difference pyramid was set to 6 groups with 3 layers per group, the initial Gaussian blur σ was set to 1.6, the contrast threshold for feature point detection was set to 0.03, and the edge threshold was set to 10. Low-contrast invalid feature points caused by metallic reflections and solid-color areas on walls were removed. The nearest neighbor distance ratio method was used for feature point matching, with a distance ratio threshold of 0.75, to select highly matched feature point pairs in the engineering site images. Simultaneously, considering the temporal sequence relationship of continuous frame images captured by moving safety helmets, erroneous matching point pairs caused by moving objects across frames were removed, resulting in unique and stable visual feature point matching results. This provides accurate visual feature point foundation data for subsequent fusion and alignment with the corrected positioning data to establish a global 3D spatial coordinate system for the site. In practice, the process of constructing a visual feature point map involves performing descriptor matching on feature points extracted from consecutive frames of images, using epipolar geometry constraints to filter out correctly matched point pairs, and calculating the initial position of the matched feature points in three-dimensional space using triangulation. Ultimately, a visual feature point map is formed, consisting of a series of three-dimensional spatial points and their associated image feature descriptors.
[0024] The baseline distance for triangulation is determined by the actual spatial distance traveled by the inspection personnel during adjacent frame acquisitions from the panoramic camera integrated on the safety helmet. This distance is obtained by extracting the spatial coordinate difference between adjacent frames from the synchronously acquired corrected positioning data. The camera pose is solved by essential matrix decomposition of adjacent frame images. First, the essential matrix is obtained through epipolar geometric constraints, then SVD decomposition is performed on the essential matrix. Combined with the initial pose values from the corrected positioning data, the camera rotation matrix and translation vector that conform to the spatial characteristics of the engineering site are selected to determine the camera pose in the global space of the site. The three-dimensional coordinates are solved using the spatial point projection formula of triangulation. Let the pixel coordinates of spatial point P in two adjacent frames be... , The camera intrinsic parameter matrix is K, and the camera pose in the first frame is... The camera pose in the second frame is By solving , By solving the system of simultaneous equations, we obtain the initial three-dimensional coordinates of point P in space. .
[0025] It is understandable that a visual feature point map is the foundation of visual positioning. The corrected positioning data is fused and aligned with the visual feature point map to establish a global 3D spatial coordinate system for the site. In practice, this fusion and alignment is achieved through graph optimization methods. The pose graph, with the corrected positioning data as nodes, and the map point cloud, with visual feature points as nodes, are jointly optimized to minimize reprojection and ranging errors, thereby establishing a unified global 3D spatial coordinate system for the site.
[0026] The objective function for graph optimization is the weighted sum of the reprojection error term and the ranging error term, as shown in the formula:
[0027] in: This is the reprojection error term for visual feature points. For the first The actual pixel coordinates of each visual feature point in the image These are the projected pixel coordinates of the feature point in three-dimensional space. The weighting coefficient for visual reprojection error; This is the ranging error term in the positioning data. For the first Corrected positioning coordinates of each trajectory point The coordinates of the trajectory point in the visual feature point map are the matching coordinates. Here is the weighting coefficient for the positioning and ranging error, and Combining the accuracy characteristics of on-site engineering positioning and visual acquisition, the following settings are configured. , The reprojection error term is measured by pixel distance and defined as the Euclidean distance between the actual pixel coordinates and the projected pixel coordinates. The ranging error term is measured by spatial distance and defined as the three-dimensional Euclidean distance between the corrected positioning coordinates and the visual matching coordinates. The optimization solver uses the Gauss-Newton method, and the iterative convergence threshold is set to... The maximum number of iterations is 50. During the iteration process, outliers that exceed the reasonable error range are removed in real time until the objective function value converges to within the threshold or the maximum number of iterations is reached, thus completing the joint optimization of the pose graph and the map point cloud.
[0028] The coordinate transformation calculation is performed on the pixel positions in the standardized environmental image and the corrected positioning data using the global three-dimensional spatial coordinate system. The coordinate transformation calculation is based on the pinhole camera model and uses the pre-calibrated extrinsic parameter matrix to convert the two-dimensional coordinates of any pixel in the image into three-dimensional spatial coordinates in the global three-dimensional spatial coordinate system, and finally obtains the spatial coordinates corresponding to each frame of the standardized environmental image. Depth information is provided by the 3D point cloud data of the visual feature point map in the constructed on-site global 3D spatial coordinate system. The depth value in the 3D coordinates of the visual feature point that matches the target pixel is extracted as the depth reference for that pixel. For pixels that do not match visual feature points, the depth value is determined based on the spatial coordinates of the corrected positioning data, combined with the height constraints of the floor plan and the spatial size constraints of the building structure on the engineering site. At the same time, the imaging formula of the pinhole camera model is used as the core constraint. Equivalent constraint equations are established by combining the pre-calibrated camera intrinsic parameter matrix and the extrinsic parameter matrix of the on-site global 3D spatial coordinate system. The intrinsic parameter matrix includes calibration parameters such as camera focal length and principal point coordinates. The extrinsic parameter matrix includes the camera's rotation matrix and translation vector in the global 3D space. The constraint equations are used to correlate and calculate the 2D pixel coordinates, depth information and 3D spatial coordinates to achieve accurate conversion from 2D pixels to 3D coordinates.
[0029] After obtaining a series of timestamped spatial coordinates, the spatial coordinates are sorted and interpolated according to the time sequence. In the specific implementation, the spatial coordinate acquisition frequency is 1 Hz. In the trajectory reconstruction calculation, a cubic spline interpolation algorithm is used to insert new coordinate points between adjacent spatial coordinate points with a time interval greater than 1 second, so that the time interval of the total position sequence points is uniformly reduced to 0.1 seconds, resulting in a high-time-resolution continuous position sequence. Path fitting calculation is performed based on the continuous position sequence. The path fitting calculation uses the B-spline curve fitting method, which smoothly connects the position sequence points through control points to generate a continuous and smooth inspection trajectory data. In the specific implementation, the inspection trajectory data is represented in three-dimensional space as a series of three-dimensional coordinate points sorted by time. A preset floor plan electronic map is called. The preset floor plan electronic map is a CAD vector map containing information on all walls and room boundaries. The boundary coordinates of the key areas to be inspected are extracted from the floor plan electronic map to generate preset floor plan inspection area data. In practice, the preset floor plan inspection area data is a collection of multiple polygonal regions. Each polygon is defined by a series of sequentially connected two-dimensional plane coordinate points, representing the physical space range that needs to be inspected and covered.
[0030] The spatial point set covered by the inspection trajectory data is subjected to geometric set operations with the preset floor plan inspection area data. In specific implementation, the geometric set operation involves projecting the inspection trajectory data onto a horizontal plane to obtain a two-dimensional covered point set of the trajectory, and then calculating the area of the minimum convex hull polygon formed by this two-dimensional covered point set. The area ratio of the inspection trajectory data within the preset floor plan inspection area data is calculated, and the area ratio is given by the formula:
[0031] in: This represents the area covered by the projection of the inspection trajectory data onto the preset inspection area plane. This represents the total area of the preset floor plan inspection area data, with the symbol... This indicates the calculation of the geometric intersection area of two regions. The trajectory coverage rate is calculated based on the area percentage, expressed as a percentage of the area. When the trajectory coverage rate reaches a preset threshold, a result indicating complete trajectory coverage is generated; otherwise, a result indicating incomplete trajectory coverage is generated. In this implementation, the preset threshold is set to 90%, and this result is used as the final trajectory coverage result.
[0032] In one embodiment of the present invention, see [reference] Figure 3 The system preprocesses the on-site environmental images, which are raw images captured by a panoramic camera integrated into the safety helmet. Image preprocessing includes image size normalization, color space conversion, and enhancement. Specifically, image size normalization scales the long side of the original image to 640 pixels, while the short side is scaled proportionally to ensure the input image size meets the requirements of the safety helmet wearing recognition model and the safety hazard recognition model. Color space conversion and enhancement convert the image from the standard RGB color space to the HSV color space. In the HSV color space, the saturation and luminance components are adjusted to enhance the contrast between the target object and the background, resulting in the preprocessed image to be recognized. This preprocessed image serves as the input for the subsequent accurate recognition by the AI model.
[0033] In practical implementation, the image to be identified is input into the safety helmet wearing recognition model. This model is a convolutional neural network model trained using the YOLOv5 framework. Standardized environmental images of engineering site inspection scenarios are used as training samples, with a total of 50,000 images divided into training, validation, and test sets in an 8:1:1 ratio. Pixel-level bounding box annotation is adopted to accurately annotate the head region of personnel and associate it with three categories: wearing, not wearing, and incorrectly wearing. The model's loss function is a weighted sum of classification and regression losses, with a classification loss weight of 0.4 and a regression loss weight of 0.6. Three sets of anchor boxes of different sizes are preset based on the pixel distribution characteristics of personnel heads at the engineering site to adapt to head detection at near, medium, and far distances. During training, data augmentation is performed based on the image distribution at the engineering site to ensure the model's adaptability to complex on-site environments. The safety helmet wearing recognition model receives the image to be identified and, through its internal feature extraction network and detection head, outputs prediction information for all detected target boxes in the image. The prediction information includes the bounding box coordinates of the target boxes and the classification result of the wearing status of the personnel within the target boxes. Bounding box coordinates are typically represented by the coordinates of the top-left and bottom-right corner vertices or the center point coordinates and width and height of the bounding box. The helmet wearing status classification result is the final output of the model classifying each detected person's head target. In practice, the helmet wearing status classification result includes three categories: worn, not worn, and incorrectly worn. Incorrectly worn helmets include situations such as the helmet not being fastened or the helmet being worn at an improper angle.
[0034] In practical implementation, the image to be identified is input into the safety hazard identification model. This model is a deep learning model trained on the Faster R-CNN framework. Standardized environmental images of engineering sites containing various safety hazards are used as training samples, with a total of 80,000 images divided into training, validation, and test sets in an 8:1:1 ratio. Accurate bounding box annotations and category label associations are completed for four types of hazards: improper placement of items, obstruction of fire exits, lack of protective facilities, and electrical violations. Multiple hazards in the same image are labeled independently. The model's loss function is a weighted sum of classification loss, regression loss, and region proposal loss, with weights of 0.3, 0.5, and 0.2 respectively. Five sets of anchor boxes with different aspect ratios are preset based on the pixel size and morphological features of the hazards to adapt to different types and angles of safety hazard detection. During training, engineering scene adaptation data augmentation is performed to improve the model's on-site recognition accuracy and robustness. The safety hazard identification model receives the image to be identified, processes it through a region generation network and a region proposal network, and finally outputs the detection results of various safety hazards identified in the image. The detection results are provided in the form of structured data, including the category of safety hazard, the bounding box coordinates of the hazard in the image to be identified, and the model's confidence score for the identification result. In practice, the categories of safety hazards include, but are not limited to, improper placement of items, obstruction of fire exits, lack of protective facilities, and electrical violations. For example, improper placement of items can include situations such as materials being stacked too high or tools being placed haphazardly; obstruction of fire exits refers to passageways being blocked by debris or equipment; lack of protective facilities includes missing railings or uncovered covers; and electrical violations include situations such as unauthorized wiring or unclosed distribution box doors. The bounding box coordinates of the safety hazard are used to locate the specific area of the hazard in the image to be identified, and the confidence score is used to measure the model's certainty regarding the current identification result.
[0035] Based on the bounding box coordinates output by the safety hazard identification model, specific rectification areas can be located in the image to be identified. In practice, locating specific rectification areas involves cropping the corresponding hazard area image block from the image to be identified according to the bounding box coordinates, and then further analyzing the hazard area image block. In some embodiments, for the hazard of "fire lane obstruction" indicated by the bounding box coordinates, the cropped hazard area image block is the area where the passage is obstructed. The rectification point coordinates in the spatial coordinate system are calculated. The calculation of the rectification point coordinates is based on the image area determined by the bounding box coordinates and the pre-established mapping relationship from the image pixel coordinate system to the spatial three-dimensional coordinate system. In practice, for some hazard categories, the rectification point coordinates can be directly taken as the spatial coordinates corresponding to the center point of the bounding box coordinates; for hazards such as "missing protective facilities," the rectification point coordinates need to be specifically calculated as the pixel coordinates of the installation hole or fixing point in the image, and then converted into spatial three-dimensional coordinates.
[0036] The calculation of pixel coordinates for rectification points is performed in two categories based on the type of hazard. The first category includes hazards such as improper placement of items, obstruction of fire exits, and electrical violations. First, the pixel coordinates of the hazard's bounding box are extracted. Then, the midpoint coordinates of the horizontal and vertical pixels of the bounding box are calculated using the following formula:
[0037]
[0038] in: , These are the minimum and maximum x-coordinates of the bounding box pixels. , These are the minimum and maximum values of the ordinate of the bounding box pixels. The coordinates of this midpoint are the pixel coordinates of the rectification point.
[0039] The second category is hazards caused by missing protective facilities. First, key point detection processing is performed on the cropped hazard area image block to identify the local pixel coordinates of key rectification parts such as installation holes and fixing points in the image block. Then, the starting pixel coordinates of the hazard boundary box in the image to be identified are combined. The global pixel coordinates of the key rectification areas in the image to be identified are calculated using the coordinate superposition formula, which are the pixel coordinates of the rectification points for this type of hidden danger. After the pixel coordinates of the rectification points are determined, the mapping relationship between pixel coordinates and three-dimensional coordinates in the pre-established global three-dimensional spatial coordinate system of the site is called. Substituting the camera intrinsic parameter matrix, extrinsic parameter matrix, and distortion correction parameters, the two-dimensional global pixel coordinates of the rectification points are converted into three-dimensional spatial coordinates using the inverse projection formula of the pinhole camera model. During the conversion process, the coordinate values are calibrated in combination with spatial constraints such as floor height and building structure dimensions at the engineering site, and finally the accurate three-dimensional spatial coordinates of the rectification points are obtained, thus completing the calculation of the rectification point coordinates.
[0040] The system integrates the helmet wearing status classification results output by the safety helmet wearing recognition model, the hazard category and bounding box coordinates output by the safety hazard recognition model, and the calculated rectification point coordinates to generate structured recognition results. In practice, the data integration operation organizes the output information from different models and calculation steps according to a pre-defined data structure. The recognition result can be in the format of a JSON object, containing the following fields: image timestamp index, a list of identified personnel (each person includes bounding box coordinates and wearing status), and a list of identified hazards (each hazard includes hazard category, bounding box coordinates, confidence level, and rectification point coordinates). In essence, the recognition result is the key output that transforms unstructured image information into structured, quantifiable data.
[0041] In one embodiment of the present invention, the safety hazard identification model employs a convolutional neural network with a feature pyramid structure. After the image to be identified is input into the safety hazard identification model, the convolutional neural network of the safety hazard identification model performs multi-layer convolution and pooling calculations on the input image. In a specific implementation, the convolutional neural network uses VGG16 as the backbone feature extraction network. Through its multiple convolutional and pooling layers, it progressively extracts information from the image to be identified, from low-level edge textures to high-level semantic features, generating multi-scale feature maps. These multi-scale feature maps are feature maps output by different convolutional layers, possessing different spatial resolutions and receptive fields. Based on the multi-scale feature maps, the safety hazard identification model performs feature fusion calculations within the feature pyramid structure. In a specific implementation, the feature pyramid structure upsamples the high-level semantically rich feature maps from top to bottom and adds them element-wise to the shallower, spatially detailed feature maps. This process is repeated until the shallowest layer, generating a series of fused feature maps that integrate multi-scale information. Each fused feature map simultaneously possesses accurate localization information and rich semantic information, used for subsequent detection of hazard targets of different sizes. The top-down path of the feature pyramid structure has four layers, corresponding to four feature layers with different spatial resolutions in the multi-scale feature map. The top-down upsampling operation uses 2x nearest neighbor upsampling to enlarge the high-level low-resolution feature map to the same spatial resolution as the next shallow feature map. Before element-wise addition, the shallow feature map with rich spatial details is first processed by a 1×1 convolutional layer to reduce the number of channels, so that the number of channels in the shallow feature map is completely aligned with the number of channels in the upsampled high-level feature map. At the same time, lateral connections are set in the feature fusion process. The lateral connection layer performs a 1×1 convolution transformation on the original shallow feature map, extracts effective features, and directly maps them to the top-down feature fusion path of the corresponding layer. Element-wise addition is performed with the upsampled high-level feature map. After each layer is fused, a 3×1 convolutional layer is used to refine the feature fusion and ensure the consistency of the feature representation of the fused feature map.
[0042] Anchor boxes of a preset size are applied to the fused feature map. In practice, these anchor boxes are a series of predefined reference boxes with different aspect ratios at each feature map pixel. The size of the anchor boxes is set according to the statistical distribution of actual safety hazard targets in the training dataset. Regression and classification calculations are performed on the features within each anchor box. Regression calculates the adjustment parameters of the predicted bounding box of a potential hazard target relative to the corresponding anchor box. These adjustment parameters include the center point coordinate offset and the aspect ratio scaling factor. Classification calculates the probability that the target within the anchor box belongs to each predefined hazard category and outputs a confidence score indicating the presence of a target in the predicted box. Non-maximum suppression (NMS) is applied to all predicted bounding boxes. In practice, NMS uses the confidence score as the sorting criterion. Predicted boxes belonging to the same category and with an overlap exceeding a preset threshold are filtered, and only the predicted box with the highest confidence score is retained to obtain the final hazard target detection result. The hazard target detection result is a data list containing multiple entries. Each entry explicitly includes the hazard category, the bounding box coordinates describing the hazard target's location in the image, and a confidence level characterizing the certainty of the model's prediction. In some embodiments, the overlap judgment threshold used in the non-maximum suppression processing is set to 0.5.
[0043] The bounding box coordinates of specific safety hazards are obtained from the identification results, which are structured data output by the safety hazard identification model. In practice, the bounding box coordinates of specific safety hazards are represented as normalized quadruples, which include the x and y coordinates of the bounding box center point, as well as the width and height of the bounding box. Based on the bounding box coordinates, the corresponding hazard area image block is extracted from the image to be identified. The extraction operation involves restoring the normalized coordinates to the actual pixel coordinates in the image to be identified, and cropping the image sub-block according to the rectangular area defined by the coordinates. Key point detection processing is performed on the hazard area image block. In practice, key point detection processing adopts a key point detection model based on a convolutional neural network. The input of the key point detection model is the hazard area image block, and the output is the pixel coordinates of the specific part of the image block that needs rectification in the image block coordinate system. The pixel coordinates are represented in the form of coordinate pairs. In practical implementation, for hazards such as "missing guardrails," the key point coordinates output by the key point detection model correspond to the fixed points where the missing guardrails should be connected. For hazards such as "debris piled up in front of fire extinguisher boxes," the key point coordinates output by the key point detection model can correspond to the corner points of the designated area where the debris needs to be removed. The key point detection model adopts a lightweight convolutional neural network structure with a backbone feature extraction network and a key point prediction head. The backbone feature extraction network is adapted to the small-size feature extraction requirements of the hazard area image blocks. The key point prediction head includes a feature fusion layer and a coordinate regression layer. The feature fusion layer fuses the multi-scale features output by the backbone network, and the coordinate regression layer directly regresses the pixel coordinates of the key points of the rectification area in the image block coordinate system. The training dataset of this model is constructed based on various hazard area image blocks collected from engineering sites. Hazards are classified into four categories: improper placement of items, obstruction of fire lanes, lack of protective facilities, and electrical violations. The sample size for each category is no less than 10,000 images. The pixel coordinates of the key points of the corresponding rectification area are labeled for each image block. The number of key points for each category of hazard is set according to the actual rectification needs. The dataset is divided into training, validation, and test sets in an 8:1:1 ratio. Data augmentation is performed using image rotation, brightness adjustment, and scaling. Dynamic adaptation logic for different hazard categories is triggered by hazard category labels. When inputting an image block of a hazard area, the hazard category label output by the safety hazard identification model is simultaneously passed in. The model automatically matches the corresponding key point detection parameters and prediction logic according to the preset hazard category-key point detection rule library, and adjusts the output dimension of the coordinate regression layer to adapt to the key point quantity requirements of different hazard categories, thereby achieving accurate and universal detection of key points in various hazard rectification areas.
[0044] Based on the bounding box coordinates of the safety hazard and the pixel coordinates of the rectification area within the hazard region image block, the overall pixel coordinates of the rectification area in the original image to be identified are calculated through coordinate transformation. In practice, the coordinate transformation calculation combines the local pixel coordinates of the rectification area in the hazard region image block coordinate system with the starting coordinates of the hazard region image block in the original image to be identified, and uses a formula to transform them to obtain the overall pixel coordinates of the rectification area in the original image to be identified.
[0045] in: This represents the overall pixel coordinates of the area to be rectified in the original image to be identified, after transformation. The vertical coordinate represents the overall pixel coordinate of the area to be rectified in the original image to be identified, after transformation. The minimum x-coordinate of the bounding box representing the safety hazard in the original image to be identified. The minimum ordinate of the bounding box representing the safety hazard in the original image to be identified. This represents the local x-coordinate pixel coordinates of the rectified area in the image block coordinate system of the hidden danger area, obtained through key point detection processing. This represents the local vertical pixel coordinate of the rectified part in the image block coordinate system of the hidden danger area, obtained through key point detection processing.
[0046] By utilizing the mapping relationship between image pixel coordinates in the spatial coordinate system and three-dimensional spatial coordinates, the overall pixel coordinates of the rectification area are converted into three-dimensional coordinates in the global three-dimensional spatial coordinate system at the site, thus obtaining the rectification point coordinates. In specific implementations, the mapping relationship between image pixel coordinates in the spatial coordinate system and three-dimensional spatial coordinates is defined by the camera's intrinsic parameter matrix, extrinsic parameter matrix, and distortion coefficients. In specific implementations, this conversion is achieved by solving a system of linear equations containing camera parameters, projecting the two-dimensional pixel coordinates back onto a ray in three-dimensional space. Since a dense three-dimensional point cloud constructed from a visual feature point map already exists in the global three-dimensional spatial coordinate system at the site, the intersection point of this ray and the nearest three-dimensional point located on a known horizontal plane can be calculated; the coordinates of this intersection point are the rectification point coordinates. In some embodiments, assuming that the known hidden danger target is roughly located on the ground or a workbench plane, and its vertical coordinates are known, the pixel coordinates can be directly converted into three-dimensional planar coordinates through homography transformation using the camera's extrinsic parameters, thereby obtaining the rectification point coordinates.
[0047] The system of linear equations for the transformation between pixel coordinates and 3D spatial coordinates is constructed based on a pinhole camera model. Let the overall pixel coordinates of the rectification area be... The three-dimensional spatial coordinates are The camera intrinsic parameter matrix is:
[0048] in: The horizontal and vertical focal lengths of the camera. (Primary pixel coordinates), extrinsic parameter matrix is It is a 3×3 rotation matrix. It is a 3×1 translation vector with distortion coefficients of . First, perform distortion correction on the pixel coordinates to obtain distortion-free pixel coordinates. Then through the system of equations Coordinate transformation is implemented; depth values are preferentially obtained from the dense 3D point cloud of the visual feature point map. Neighborhood interpolation is performed on the point cloud data corresponding to the pixel coordinates of the rectification point to generate the accurate depth value of that pixel. For sparse point cloud areas, the spatial coordinate depth value is combined with the corrected positioning data as a benchmark, and the dimensional constraints of the building structure at the engineering site are superimposed to complete the depth value. At the same time, the interference of point cloud noise points on the depth value is eliminated; the determination of known planes is achieved through a preset plane feature library. The visual features of the images around the rectification point and the geometric features of the spatial point cloud are extracted and compared with the ground and work areas in the library. The feature templates of preset planes such as the tabletop and equipment operating surface are matched, and the plane with the highest matching degree is the corresponding known plane. Its vertical coordinates are retrieved from the building structure parameter library of the global three-dimensional spatial coordinate system on site. The intrinsic and extrinsic parameters and distortion coefficients of the camera are calibrated using Zhang's calibration method. The intrinsic and extrinsic parameters and distortion coefficients are the inherent parameters of the safety helmet panoramic camera. They are calibrated once before leaving the factory and then checked and calibrated every 6 months thereafter. The extrinsic parameters are calibrated in real time in conjunction with the construction process of the global three-dimensional spatial coordinate system on site. The extrinsic parameters are calibrated through the on-site calibration board before each inspection to ensure the accuracy and reliability of coordinate transformation.
[0049] In one embodiment of the present invention, a time-aligned index is established. Both the recognition results and the inspection trajectory data are recorded with timestamps using a high-precision synchronized clock. The data structure of the recognition results includes a timestamp field, identified safety hazard information, helmet wearing status, bounding box coordinates, and rectification point coordinates. The data structure of the inspection trajectory data includes a timestamp field and spatial coordinates. In a specific implementation, each frame of recognition data in the recognition results is matched and associated with trajectory points at the same timestamp in the inspection trajectory data. The matching and association operation is performed in a memory database or data processing module, using the nanosecond-level precision timestamp as the primary key, to perform an equi-join query on the recognition result data table and the inspection trajectory data table, thereby binding the visual recognition information collected at the same time with the personnel location information. The matched and associated data is then integrated. The integration operation creates a new spatiotemporal associated data chain structure. Each data node in the spatiotemporal associated data chain includes the original timestamp, the X component, Y component, and Z component of the spatial coordinates, the helmet wearing status identified at that location point, the identified safety hazard category, the bounding box coordinates of the hazard, the rectification point coordinates, and the file index path of the associated on-site environmental image. In some embodiments, the spatiotemporal correlated data chain is represented in memory as an ordered dictionary or linked list, with its nodes arranged strictly in ascending order of timestamps.
[0050] Key information is extracted from the spatiotemporal correlation data chain. In specific implementation, this key information includes the timestamp of the hazard occurrence, the spatial coordinates of the hazard occurrence, the identified hazard type, the index of the on-site environmental image file associated with the hazard occurrence time, and the coordinates of the rectification point calculated for the hazard. The report is formatted and filled according to a preset report template, which is a structured document template containing chapter titles, fixed tables, and chart placeholders. In specific implementation, formatting involves filling the key information extracted from the spatiotemporal correlation data chain into the corresponding fields of the report template. For example, the timestamp is converted to a standard date and time format and entered into the "Discovery Time" field; the spatial coordinates are mapped to the floor plan and entered into the "Location Description" field; and the hazard type name is entered into the "Hazard Category" field. The report template embeds keyframes of on-site environmental images and visualizations of inspection trajectory data on floor plans. In practice, embedding keyframes of on-site environmental images involves selecting the highest-quality image from image files linked by a spatiotemporal data chain node, within a few frames before and after the time of the hazard occurrence, and inserting it at a specified size into the corresponding section of the report document. Visualizing the inspection trajectory data on the floor plan involves plotting the X and Y coordinate sequences from the inspection trajectory data as continuous trajectory lines on the floor plan CAD base map, and then inserting this trajectory-integrated floor plan as an image into the report document. The generated panoramic inspection trajectory report is a comprehensive document integrating text, tables, and images. Organized chronologically, the report meticulously records the location movements, status identification, and hazard discovery during the inspection process.
[0051] The process involves acquiring the dwell time and movement speed of inspection personnel at each trajectory point from the inspection trajectory data. Dwell time is calculated by analyzing the timestamp sequence of personnel near the same coordinate point, while movement speed is calculated by dividing the spatial straight-line distance between adjacent trajectory points by the time interval. The theoretical movement time between adjacent trajectory points is then calculated. In practice, this calculation is based on the spatial coordinates and movement speed of the trajectory points recorded in the inspection trajectory data. In the ideal uniform linear motion model, the distance between two adjacent points... and Theoretical travel time between It can be calculated using a formula:
[0052] in: Indicates from the trajectory point Move to The theoretical travel time is measured in seconds. Represents trajectory points Coordinates in three-dimensional space Represents trajectory points Coordinates in three-dimensional space. Indicates from the trajectory point To the trajectory point The average moving speed of the inspection personnel along this route. The velocity can be calculated by taking the arithmetic mean of the starting and ending velocities of the path segment, or by dividing the length of the path segment by the actual time interval. This formula describes a theoretical time estimation model based on physical kinematics. The starting and ending velocities of the trajectory points are extracted from the restored continuous and smooth inspection trajectory data. Each trajectory point in the inspection trajectory data is associated with a unique timestamp and three-dimensional spatial coordinates. The starting velocity is the instantaneous movement speed of the trajectory point, obtained by calculating the three-dimensional straight-line distance to the previous adjacent trajectory point and dividing it by the timestamp difference between the two points. The ending velocity is also the instantaneous movement speed of the trajectory point, obtained by calculating the three-dimensional straight-line distance to the next adjacent trajectory point and dividing it by the timestamp difference between the two points. If a trajectory point is the starting or ending point of an inspection trajectory, its instantaneous velocity is taken as the average movement speed of an adjacent single trajectory segment as a substitute value, ensuring that both the starting and ending velocities can be effectively obtained from the inspection trajectory data.
[0053] The calculated theoretical movement time is compared with the actual time interval of data collection, where the actual time interval is the time interval between trajectory points. and The difference between the corresponding timestamps. In practice, a time consistency check value is obtained by comparison, and the time consistency check value can be defined as the theoretical shift time. Time interval between actual data collection The absolute value of the difference between them, i.e. If the time consistency check value exceeds the error range threshold, the corresponding trajectory segment is marked as abnormal. The error range threshold is a preset threshold based on factors such as positioning system error, personnel gait changes, and model simplification; in practice, it can be set to 3 seconds. The morphological characteristics of the inspection trajectory data are analyzed, including the smoothness of the movement path and the rationality of directional changes. The smoothness of the movement path is evaluated by calculating the curvature change of the trajectory point sequence, and the rationality of directional changes is judged by analyzing whether the angle changes of continuous trajectory points are within the allowable range of human kinematics. A matching degree calculation is performed with the standard inspection path model. The standard inspection path model is a data model pre-constructed based on standard operating procedures and floor plans, containing recommended inspection paths and key inspection point locations. The standard inspection path model is based on a preset floor plan electronic map, which is drawn based on the actual building structure at the engineering site and includes the geographic information and coordinate data of walls, room boundaries, and key areas to be inspected. The process involves extracting the boundary coordinates of key areas to be inspected and the coordinates of necessary inspection points. Based on the standard operating procedures for the engineering site, the optimal inspection route is planned. Route nodes are sorted according to the inspection operation sequence and assigned standard dwell time and direction constraints, forming a structured data model containing a sequence of standard points, standard path segments, key checkpoints, inspection direction, and standard speed range. This completes the construction of the standard inspection path model. In practice, the matching degree calculation can include calculating the Friesian distance or dynamic time warping distance between the actual inspection trajectory and the standard inspection path, obtaining a path rationality verification value that characterizes the similarity between the two. The path rationality verification value is a scalar value; the smaller the value, the higher the matching degree and the more reasonable the trajectory.
[0054] The trajectory authenticity verification result is generated by combining the time consistency verification value and the path rationality verification value. In specific implementation, the trajectory authenticity verification result is a binary judgment or comprehensive score. Only when both the time consistency verification value and the path rationality verification value are within their respective preset normal threshold ranges is the inspection trajectory data deemed authentic and valid. It can be understood that the normal threshold range is a pre-set numerical range used to distinguish between normal and abnormal behavior. For example, the normal threshold range for the time consistency verification value is 0 to 3 seconds, and the normal threshold range for the path rationality verification value is 0 to 5 meters. In some embodiments, the trajectory authenticity verification result is presented in the form of a structured verification record table, which records the verification result of each segment of the inspected trajectory (see Table 1).
[0055] Table 1: Inspection Track Authenticity Verification Record Table
[0056] See Figure 4This is a time-series analysis chart of helmet wearing status, showing the changes in the number of personnel wearing helmets at different times during the inspection period. It is a visualization output from the AI multi-modal recognition stage. The number of personnel wearing helmets is stable between 75-90, representing the absolute mainstream state. The fluctuation range is small, maintaining a high level of compliance overall, with peaks occurring at times such as 10:01:30 and 10:02:15. The number of personnel not wearing helmets remains between 0-10, representing a minority of abnormal states. High-incidence periods are concentrated at 10:00:15, 10:01:15, 10:02:00, and 10:04:00, requiring close attention to on-site management during these periods. The number of personnel incorrectly wearing helmets is similar to that of "not wearing helmets," also fluctuating between 0-10, and even slightly higher than the number of "not wearing helmets" at some times. The time distribution of this state highly overlaps with that of "not wearing helmets," reflecting inadequate enforcement of helmet wearing regulations at certain times.
[0057] In one embodiment of the present invention, safety hazard information identified in the panoramic inspection trajectory report is parsed. The panoramic inspection trajectory report is a data file or database record containing structured hazard information. The parsing operation extracts detailed content for each independent hazard entry from the data structure or specified fields of the panoramic inspection trajectory report. The extracted information includes a hazard category description, the coordinates of the rectification point in the three-dimensional spatial coordinate system, the specific timestamp of the hazard's discovery, and an index or thumbnail of the associated on-site environmental image file. In a specific implementation, the hazard category is described in string form, the rectification point coordinates are stored in the form of three-dimensional coordinate pairs, the discovery time is recorded in the ISO8601 standard date and time format, and the associated environmental image is stored in the form of a file path or a database binary large object.
[0058] Based on a pre-defined mapping relationship between hazard categories and responsible departments, a corresponding responsible department is assigned to each hazard. In practice, this mapping relationship is a key-value pair mapping table stored in a database or configuration file. The mapping table defines unique identifiers for one or more responsible departments corresponding to a specific hazard category string. For example, the hazard category "Occupied Fire Exit" maps to the responsible department "Safety Management Department," and the hazard category "Electrical Violation" maps to the responsible department "Equipment Support Department." A required rectification completion time is assigned to each hazard. In practice, this allocation is achieved by calling pre-defined rectification time limit rules. These pre-defined rectification time limit rules can be a rule engine containing multiple rules, calculated based on factors such as the hazard's category, level, location, and historical records. In some embodiments, the rectification time limit rules are defined as formulas:
[0059] in: This indicates the calculated required completion time for rectification. This indicates the time when the potential hazard was discovered. This represents the baseline time coefficient related to the hazard category; different hazard categories correspond to different coefficient values. The severity level coefficient represents the severity of a hazard. It is determined by a combination of factors, including the confidence level output by the hazard identification model and the importance of the hazard category. It is usually a scalar with a value greater than 0. This represents a base buffer time constant. A fixed value, uniformly set at 4 hours, serves as the basic preparation time for rectifying various potential hazards; a benchmark time coefficient related to the hazard category. Based on the four categories of hazards specified in the original document—improper placement of items, obstruction of fire exits, lack of protective facilities, and electrical violations—a quantitative value was assigned, with improper placement of items being a key hazard. Take 8 hours, fire lane blocked 12-hour sampling, lack of protective equipment Take 24 hours, electrical violation The timeframe is 16 hours, and the assignment is based on the complexity of rectification operations for various hazards at the engineering site, the preparation cycle for required materials and manpower, and the severity level coefficient of the hazard. The confidence level is calculated quantitatively from the output of the safety hazard identification model. When the confidence level is ≥90%, When the confidence level is 70% ≤ confidence level < 90%, the value is 1.5. When the confidence level is 1.0, it is less than 70%. Take 0.8, and make corrections based on the importance of the hazard category. Hazards such as fire lane obstruction and electrical violations are multiplied by an additional correction factor of 1.2 on the original calculation results to ensure that the severity level coefficient matches the actual risk level of the hazard. All parameter calculation results are in hours. The final rectification completion time is calculated according to the formula and the hourly time is taken as the final required time limit.
[0060] The categories of hazards, coordinates of rectification points, discovery time, associated environmental images, corresponding responsible departments for rectification, and required rectification completion time are filled into the corresponding fields of the standard rectification task sheet template. The standard rectification task sheet template is a structured electronic form containing predefined fields. In practice, the filling operation is a programmatic data writing process. The corresponding values are read from the parsed and calculated data structure and written into the designated locations in the standard rectification task sheet template. For example, the coordinates of the rectification point are written into the "Hazard Location" field, and the required rectification completion time is written into the "Rectification Deadline" field. A unique task sheet number is assigned to each generated rectification task sheet. The unique task sheet number can be generated in the form of "Year + Month + Date + Sequence Number", such as "20260317001", to ensure the uniqueness of each task sheet number within the system. The rectification task sheet is sent to the corresponding responsible department. Sending methods include, but are not limited to, pushing through the company's internal workflow system, sending to the email address designated by the responsible department, or notifying via the company's instant messaging robot.
[0061] The system tracks and records the status changes of each rectification task order. In practice, the status of each rectification task order is defined as a finite state machine in the system, with typical statuses including "issued," "received," "in rectification," "pending acceptance," and "completed." When a task order is generated and sent by the system, the status is initialized to "issued." When the responsible department confirms receipt of the task in the system, the status changes to "received." After the responsible department begins processing, the status can be manually or automatically changed to "in rectification." After the responsible department submits a rectification completion report, the status changes to "pending acceptance." After review and approval, the status finally changes to "completed." Each status change is accompanied by a timestamp, the operator, and optional remarks. These records are stored in the system database.
[0062] When the department responsible for rectification uploads on-site images after rectification is completed, these images are re-taken at the location of the potential hazard after the department claims completion. The system automatically uses an AI recognition model to verify the uploaded images. In practice, this AI model can be the same as the safety hazard identification model or a version specifically optimized for rectification verification tasks. Verification involves inputting the uploaded images into the AI model, which outputs a judgment on whether the image contains a hazard of the same type as in the original report, along with a confidence score. The system verifies whether the original hazard has been eliminated, generating a verification result. This result is a binary judgment and may include a description of the current image status as identified by the model. Based on the verification result, the system updates the status of the corresponding rectification task. In practice, if the verification result is "passed," the system automatically updates the status of the corresponding rectification task from "pending acceptance" to "completed," and records the verification time and a thumbnail of the verification basis in the database. If the verification result is "failed," the status can revert to "rectifying," and a feedback message containing the reason for the failure is sent to the responsible department. If the review is successful, the status will be updated to "completed" and a rectification completion record will be generated. The rectification completion record is a structured record containing detailed information such as task number, hidden danger information, before rectification photos, after rectification photos, review results, and completion time.
[0063] Based on the rectification completion record and the panoramic inspection trajectory report, a closed-loop management report is generated, encompassing the entire process from hazard discovery, rectification order issuance, process tracking, and result verification. In implementation, the closed-loop management report is generated by querying the original hazard records from the rectification completion record and the panoramic inspection trajectory report in the system database, and integrating the complete process log of task order status changes. Presented in a timeline format, the closed-loop management report clearly shows the entire process from the initial hazard discovery during inspection, the system's generation and issuance of the rectification task order, the responsible department's handling and status update, to the final upload of images and AI verification, forming a traceable management loop.
[0064] See Figure 5This is a chart analyzing the helmet-wearing status of inspection personnel, showing the trend of helmet-wearing compliance rate during the inspection period from 8:00 to 18:00. It is a core visualization output of the AI multi-modal recognition stage. The compliance rate remained stable between 92% and 98% throughout the day, demonstrating excellent overall performance and meeting safety production requirements. It exhibits a daily fluctuation pattern of "high-low-high," with the compliance rate at its lowest point around noon (12:00), possibly related to personnel movement such as shift changes and mealtimes. The two peaks at 10:00 and 18:00 indicate that helmet-wearing compliance was more effectively enforced at the end of the morning shift and the beginning of the evening shift. The compliance rate dropped to 92% at 12:00, the lowest point of the day, requiring close monitoring of on-site supervision and AI recognition verification during this period. The overall compliance rate remained above 90% throughout the day, with no extreme violations, indicating good basic management effectiveness. The higher compliance rates in the morning and evening, followed by a decline at noon, suggest the need for differentiated management strategies for different time periods.
[0065] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A method for verifying helmet wearing and inspection trajectory by integrating AI multi-modal recognition, characterized in that, The method includes: The on-site environment images and location data are collected by the panoramic camera and real-time positioning module integrated on the safety helmet. The on-site environment images and location data are processed to construct a spatial coordinate system to obtain spatial coordinates. Based on the spatial coordinates, the real-time movement data of the inspection personnel is processed to reconstruct the trajectory and obtain the inspection trajectory data. The inspection trajectory data is then compared with the preset floor plan inspection area data to generate the trajectory coverage result. AI multimodal recognition processing is performed on the on-site environment image to automatically identify the helmet wearing status, safety hazards, and rectification parts in the image, and to obtain the recognition result including status indicators, hazard locations, and rectification point coordinates; The identification results are spatiotemporally correlated with the inspection trajectory data to generate a panoramic inspection trajectory report. The authenticity of the inspection trajectory data is verified, and a rectification task sheet is automatically generated based on the verified panoramic inspection trajectory report.
2. The method for verifying helmet wearing and inspection trajectory by integrating AI multi-modal recognition according to claim 1, wherein the method involves acquiring on-site environmental images and location data through a panoramic camera and real-time positioning module integrated on the helmet, and processing the on-site environmental images and location data to construct a spatial coordinate system to obtain spatial coordinates, characterized in that... include: The panoramic camera integrated on the safety helmet is controlled to acquire a panoramic video stream of the scene at a preset frequency. The panoramic video stream of the scene is then subjected to image frame extraction and distortion correction to obtain a standardized environmental image. The original positioning data of the positioning module is synchronously acquired by the real-time positioning module integrated on the safety helmet. The original positioning data is then filtered and error compensation is calculated to obtain the corrected positioning data. Visual feature points are extracted and matched based on the standardized environmental image to construct a visual feature point map. The corrected positioning data is then fused and aligned with the visual feature point map to establish a global three-dimensional spatial coordinate system for the site. The coordinate transformation calculation is performed on the pixel positions in the standardized environmental image and the corrected positioning data using the on-site global three-dimensional spatial coordinate system to obtain the spatial coordinates corresponding to each frame of the image.
3. The method for verifying helmet wearing and inspection trajectory by integrating AI multi-modal recognition according to claim 1, wherein the real-time movement data of the inspection personnel is processed by trajectory reconstruction calculation based on the spatial coordinates to obtain inspection trajectory data, and the inspection trajectory data is compared with the coverage data of the preset floor plan inspection area to generate a trajectory coverage result, characterized in that, include: The spatial coordinates are sorted and interpolated according to the time series to obtain a continuous position sequence of the inspection personnel. Path fitting calculation is performed based on the position sequence to obtain continuous and smooth inspection trajectory data. Call the preset floor plan electronic map, extract the boundary coordinates of the key areas to be inspected from the floor plan electronic map, and generate preset floor plan inspection area data; Perform geometric set operations on the set of spatial points covered by the inspection trajectory data and the preset floor plan inspection area data to calculate the area ratio of the inspection trajectory data within the preset floor plan inspection area data. The trajectory coverage rate is calculated based on the area ratio. When the trajectory coverage rate reaches a preset threshold, a result of complete trajectory coverage is generated; otherwise, a result of incomplete trajectory coverage is generated, and the result is used as the trajectory coverage result.
4. The method for verifying helmet wearing and inspection trajectory by integrating AI multi-modal recognition as described in claim 1, wherein the on-site environment image is processed by AI multi-modal recognition to automatically identify the helmet wearing status, safety hazards, and rectification areas in the image, and to obtain an identification result including status indicators, hazard locations, and rectification point coordinates, characterized in that, include: The on-site environment image is preprocessed, including image size normalization, color space conversion and enhancement, to obtain the preprocessed image to be identified; A safety helmet wearing recognition model is constructed. The image to be recognized is input into the safety helmet wearing recognition model, and the safety helmet wearing recognition model outputs the bounding box coordinates of the person's head in the image and the corresponding wearing status classification result. The wearing status classification result includes wearing, not wearing, and incorrect wearing. The safety helmet wearing recognition model is built on a convolutional neural network framework. It uses standardized environmental images of engineering site inspection scenarios as training samples, marks the bounding box of the head area of the personnel in the samples and associates it with the wearing status label, and presets anchor boxes in combination with the pixel distribution characteristics of the personnel's head at the engineering site. After model training, a recognition model adapted to the on-site scenario is obtained. A safety hazard identification model is constructed. The image to be identified is input into the safety hazard identification model, and the safety hazard identification model outputs the category, bounding box coordinates, and confidence level of the safety hazard in the image. The categories of safety hazards include improper placement of items, obstruction of fire lanes, lack of protective facilities, and electrical violations. The safety hazard identification model is built on a deep learning framework. It uses standardized environmental images containing various safety hazards at the engineering site as training samples. The hazard areas in the samples are labeled with bounding boxes and associated with category labels. Anchor boxes are preset based on the pixel size and morphological features of the hazards. The model is trained to obtain an identification model that is adapted to the on-site scene. Based on the bounding box coordinates of the aforementioned safety hazard, the specific rectification area is located in the image to be identified, and the rectification point coordinates of the rectification area in the spatial coordinate system are calculated. The data of the wearing status classification results, the category of safety hazard, the bounding box coordinates, and the rectification point coordinates are integrated to generate an identification result that includes status identifier, hazard location, and rectification point coordinates.
5. The method for verifying helmet wearing and inspection trajectory by integrating AI multi-modal recognition according to claim 4, wherein the image to be identified is input into a safety hazard identification model, and the safety hazard identification model outputs the category, bounding box coordinates, and confidence level of various safety hazards in the image, characterized in that, include: The convolutional neural network of the safety hazard identification model performs multi-layer convolution and pooling calculations on the input image to be identified to extract multi-scale feature maps of the image. Based on the multi-scale feature map, feature fusion calculation is performed in the feature pyramid structure to generate a fused feature map for target detection. Anchor boxes of a preset size are applied to the fused feature map, and regression and classification calculations are performed on the features within each anchor box to predict the bounding box coordinate adjustment parameters, category probability, and confidence score of potential hazard targets. Non-maximum suppression is applied to all predicted bounding boxes to filter out the final hazard detection results. The hazard detection results include the hazard category, bounding box coordinates, and corresponding confidence scores.
6. The method for verifying helmet wearing and inspection trajectory by integrating AI multi-modal recognition according to claim 4, based on the bounding box coordinates of the safety hazard, locates the specific rectification part in the image to be identified, and calculates the rectification point coordinates of the rectification part in the spatial coordinate system, characterized in that, include: Obtain the bounding box coordinates of a specific safety hazard from the identification results, and extract the corresponding hazard area image block from the image to be identified based on the bounding box coordinates; Key point detection processing is performed on the image block of the hidden danger area to identify the specific parts in the image block of the hidden danger area that need to be rectified, and the pixel coordinates of the parts to be rectified in the image block of the hidden danger area are obtained. Based on the bounding box coordinates of the hazard and the pixel coordinates of the rectification part in the image block of the hazard area, the overall pixel coordinates of the rectification part in the original image to be identified are calculated through coordinate transformation. By utilizing the mapping relationship between image pixel coordinates and three-dimensional spatial coordinates in the spatial coordinate system, the overall pixel coordinates of the rectification area are converted into three-dimensional coordinates in the global three-dimensional spatial coordinate system of the site, thus obtaining the coordinates of the rectification point.
7. The method for verifying helmet wearing and inspection trajectory by integrating AI multi-modal recognition according to claim 1, wherein the recognition result and the inspection trajectory data are spatiotemporally correlated to generate a panoramic inspection trajectory report, characterized in that, include: Establish a time-aligned index to match and associate each frame of identification data in the identification results with the trajectory points of the same timestamp in the inspection trajectory data; The matched and associated data are integrated to form a spatiotemporal associated data chain that includes timestamps, spatial coordinates, helmet wearing status, identified safety hazard information, and coordinates of rectification points; Key information is extracted from the spatiotemporal correlation data chain, including the time, location, type, corresponding on-site environmental image, and coordinates of the rectification point where the hazard occurred, and then formatted and filled in according to the preset report template. The report template embeds keyframes of the on-site environment image and a visualization of the inspection trajectory data on the floor plan to generate a panoramic inspection trajectory report.
8. The method for verifying helmet wearing and inspection trajectory by integrating AI multi-modal recognition according to claim 1, wherein the inspection trajectory data is subjected to authenticity verification processing, characterized in that, include: The inspection personnel's dwell time and movement speed at each trajectory point are obtained from the inspection trajectory data, and the theoretical movement time between adjacent trajectory points is calculated. The calculated theoretical movement time is compared with the actual time interval of the data collection to obtain a time consistency verification value. If the time consistency verification value exceeds the error range threshold, the corresponding trajectory segment is marked as abnormal. A standard inspection path model is constructed, and the morphological characteristics of the inspection trajectory data are analyzed, including the smoothness of the movement path and the rationality of the direction change. The matching degree is calculated with the standard inspection path model to obtain the path rationality verification value. The standard inspection path model is based on a preset floor plan electronic map. It extracts the boundary coordinates of the key areas to be inspected and the coordinates of the necessary points. It plans the inspection route according to the standard operating procedures of the engineering site and combines the inspection sequence and movement constraints. By combining the time consistency check value and the path rationality check value, a trajectory authenticity check result is generated. Only when all check values are within the preset normal threshold range is the inspection trajectory data determined to be authentic and valid.
9. The method for verifying helmet wearing and inspection trajectory by integrating AI multi-modal recognition according to claim 1, wherein the automatic generation of a rectification task sheet based on the panoramic inspection trajectory report after successful verification is characterized in that, include: The safety hazard information identified in the panoramic inspection trajectory report is analyzed, and the category, rectification point coordinates, discovery time and associated environmental images of each hazard are extracted. Based on the pre-defined mapping relationship between hazard categories and responsible departments for rectification, a corresponding responsible department for rectification is determined for each hazard, and the pre-defined rectification time limit rules are invoked to assign a required rectification completion time for each hazard; Fill in the corresponding fields of the standard rectification task sheet template with the category of the hazard, the coordinates of the rectification point, the time of discovery, the associated environmental image, the corresponding department responsible for rectification, and the required rectification completion time. Each generated rectification task order is assigned a unique task order number, and the rectification task order is sent to the corresponding rectification responsible department.
10. The method for verifying helmet wearing and inspection trajectory by integrating AI multi-modal recognition according to claim 9, characterized in that, This also includes closed-loop management and handling of rectification: Track and record the status changes of each rectification task order, including issued, received, under rectification, pending acceptance, and completed; When the department responsible for rectification uploads on-site images after rectification is completed, the AI recognition model is automatically invoked to review and identify the on-site images, verify whether the original hidden danger has been eliminated, and generate a review result. Update the status of the corresponding rectification task sheet according to the review results. If the review is passed, update the status to "completed" and generate a rectification completion record. Based on the rectification completion record and panoramic inspection trajectory report, a closed-loop management report is generated that includes the entire process of hazard discovery, rectification issuance, process tracking, and result verification.