Image recognition method and system for intelligent video safety helmet

By using multiple cameras in a smart video safety helmet to work collaboratively, three-dimensional reconstruction and identification of dynamic and static hazards at construction sites are achieved, solving the problems of real-time performance and accuracy in risk assessment in traditional safety management, and providing efficient risk warning and management support.

CN121482686APending Publication Date: 2026-02-06山东港源管道物流有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511746960.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Construction sites are complex and dynamic, making it difficult for traditional safety management methods to track the dynamic operation of machinery and equipment and the movement paths of personnel in real time. This results in blurred boundaries of dangerous areas, a lack of awareness of hazards among new employees, an inability to provide accurate safety guidance, and difficulty in achieving comprehensive monitoring and timely intervention.

Method used

The system uses an intelligent video safety helmet equipped with multiple cameras facing different directions to collect real-time video of the construction site, generate 3D reconstruction and positioning, identify static and dynamic hazards, calculate the comprehensive risk value by fixed and variable distances, determine whether the preset risk value is exceeded, and issue an alarm.

Benefits of technology

It enables real-time risk assessment and timely early warning at construction sites, improving construction safety and management efficiency. It provides centimeter-level positioning and dynamic risk assessment, and supports multi-level early warning and remote monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121482686A_ABST
    Figure CN121482686A_ABST
Patent Text Reader

Abstract

The invention relates to the field of building construction, in particular to an image recognition method and system for an intelligent video safety helmet. A real-time construction site is generated through multiple construction site videos collected in real time, multiple danger sources in the real-time construction site are obtained according to the multiple construction site videos, the danger sources comprise multiple static operation units and multiple dynamic operation units, fixed distance data are obtained according to actual coordinates and the multiple static operation units, and the fixed distance data are stored in the real-time construction site. Acquiring conversion distance data according to the actual coordinates and the plurality of dynamic operation units, acquiring a comprehensive risk value according to the fixed distance data and the conversion distance data, judging whether the comprehensive risk value is greater than a preset risk value, if so, judging that the user is in a dangerous area, and giving an alarm. On-site three-dimensional reconstruction and positioning are realized through cooperation of multiple cameras, and comprehensive and real-time risk assessment, timely early warning and construction safety improvement are realized in combination with dynamic and static hazard source identification, distance calculation and risk data normalization and fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of building construction, and in particular to an image recognition method and system for an intelligent video safety helmet. Background Technology

[0002] In the construction industry, worker safety has always been a core concern, and accurate identification and early warning of hazardous areas on construction sites are crucial for building a strong safety barrier. Currently, the construction environment exhibits significant complexity and dynamism, posing serious challenges to traditional safety management methods.

[0003] From a spatial perspective, construction sites are characterized by a high density of personnel and machinery, with manual construction areas and mechanized operation areas intertwined to form a complex operational network. Traditional methods of area demarcation rely heavily on manual marking and static blueprint planning, making it difficult to track the dynamic operating trajectories of machinery and the random movement paths of personnel in real time, resulting in unclear boundaries of hazardous areas. Once changes in construction plans or temporary equipment relocation occur, existing area demarcations become outdated, failing to provide workers with accurate safety guidance and leaving them constantly exposed to potential dangers.

[0004] From a personnel awareness perspective, the construction industry experiences high staff turnover, with many newly hired workers or those temporarily assigned to unfamiliar construction sites lacking a clear understanding of the distribution of hazards on the construction site. Due to the lack of systematic pre-job safety training and visually marked hazardous areas, these workers are highly susceptible to misjudging and entering dangerous areas during operations. For example, newly arrived general laborers may be unaware of the hazard warnings around the excavation pit and approach the edge while handling materials, potentially leading to falls. Furthermore, on-site safety officers, limited by manpower and supervisory scope, cannot achieve real-time, comprehensive monitoring of all workers or intervene immediately when workers approach hazardous areas. Therefore, an image recognition method and system for intelligent video safety helmets is needed to address these issues. Summary of the Invention

[0005] The main objective of this invention is to provide an image recognition method and system for smart video safety helmets, aiming to solve the technical problems in the prior art.

[0006] This invention proposes an image recognition method and system for a smart video safety helmet, applicable to a smart video safety helmet equipped with multiple cameras facing different directions, including: Based on real-time video of multiple construction sites collected by the smart video safety helmet worn by the user, a real-time construction site is generated, and a coordinate system is generated by dividing the real-time construction site into regions using a building grid. The actual coordinates of the user wearing the smart video safety helmet at the real-time construction site are obtained based on the coordinate system. Multiple hazardous sources within the real-time construction site were obtained from multiple construction site videos. These multiple hazardous sources include multiple static work units and multiple dynamic work units. Fixed distance data is obtained based on actual coordinates and multiple static work units, and variable distance data is obtained based on actual coordinates and multiple dynamic work units; A comprehensive risk value is obtained based on fixed distance data and variable distance data; Determine whether the overall risk value is greater than the preset risk value; If the risk level exceeds the preset risk value, the user is determined to be in a dangerous area, and an alarm is issued.

[0007] Preferably, the step of generating an actual real-time construction scene based on multiple construction site videos collected in real time from the smart video safety helmet worn by the user includes: The orientation of multiple cameras mounted on each smart video safety helmet is obtained, the corresponding acquisition angle is obtained according to the orientation of each camera, and multiple real-time construction site videos are obtained according to the multiple acquisition angles; Align the initial frames of multiple real-time construction site videos with the same timestamp to extract multiple construction site photos; Based on the collection angle, obtain the original left and right images of the construction site within the same timestamp, and then align the polar lines of the original left and right images of the construction site after correction. Depth maps are obtained by extracting the projection points of the same object from the corrected construction site photos, and a real-time construction site is generated based on multiple depth maps.

[0008] Preferably, the step of obtaining fixed distance data based on actual coordinates and multiple static work units includes: The construction site video is split into multiple video frame sequence lengths according to a preset video frame sequence length. Multiple static work units are obtained based on the multiple video frame sequence lengths. The static work units include foundation pits, material piles, electrical components, and structures. The corresponding contour coordinates of the foundation pit, material pile, electrical components and structures are obtained. Multiple distance data are obtained by comparing each contour coordinate with the actual coordinates. A location distance map is generated based on the multiple distance data. The shortest distance data in the location distance map is extracted and used as the fixed distance data.

[0009] Preferably, after the step of obtaining fixed distance data based on actual coordinates and multiple static work units, the method includes: Based on the corresponding video frames, sequentially obtain images of the foundation pit, material pile, electrical components, and structures; The photos of the foundation pit and the building structure are simultaneously grayed out and the edge features of the corresponding photos are extracted. The edge pixels are obtained based on the edge features and the edge outlines of the foundation pit and the building structure are generated based on the edge pixels. The material pile photo is decomposed into blocks with similar textures, the material pile connected component contour point sequences are merged, and the contour point sequences are connected sequentially to obtain the material pile edge contour line. The electrical component image is automatically binarized to obtain the component body and the interfering background. The component body is extracted and the edges of the broken component body are connected to obtain the component edge contour line.

[0010] Preferably, the step of obtaining the transformed distance data based on the actual coordinates and multiple dynamic work units includes: The construction site video is split into multiple video frame sequence lengths according to a preset video frame sequence length. Multiple dynamic work units are obtained based on the multiple video frame sequence lengths. The dynamic work units include construction vehicles and construction equipment. The system acquires the reference position and construction trajectory of the construction equipment, obtains multiple construction change coordinates based on the construction trajectory, acquires the construction coverage edge line of the construction equipment based on the initial coordinates and multiple construction change coordinates, acquires multiple edge coordinates based on the construction coverage edge line and the real-time construction site, and acquires the first distance data based on the multiple edge coordinates and the actual coordinates. Obtain the initial coordinates and movement path of the construction vehicle. Within a preset time, based on the initial coordinates of the construction vehicle, move along the movement path and sequentially obtain the coordinates of multiple changes in the position of the construction vehicle. Based on the initial coordinates of the construction vehicle and the coordinates of multiple changes in the position of the construction vehicle, obtain the movement speed of the construction vehicle. The second change distance data is obtained based on the actual coordinates, the initial coordinates of the construction vehicles, and the coordinates of the changing positions of multiple construction vehicles. Transformed distance data is generated based on the first distance data and the second distance data.

[0011] Preferably, the step of obtaining the comprehensive risk value based on fixed distance data and transformed distance data includes: A fixed distance data vector is obtained based on the fixed distance data, and the fixed distance data vector is normalized to obtain a normalized value of the fixed distance data vector. Based on the transformed distance data, a transformed distance data vector is obtained, and the transformed distance data vector is normalized to obtain a normalized value of the transformed distance data vector. Based on the weights corresponding to the normalized values ​​of the fixed-distance data vector and the transformed-distance data vector, a comprehensive risk value is obtained.

[0012] Preferably, after the step of obtaining multiple hazard sources in the real-time construction site based on multiple construction site videos, the method includes: The real-time video information of the actual construction site is obtained by the smart video safety helmet. The real-time video information is split into multiple video frame sequence lengths according to the preset video frame sequence length. Multiple diffusion visibilitys are obtained according to the multiple video frame sequence lengths. The visibility dissipation distance is obtained from multiple diffuse visibility values, and the visibility dissipation speed and multiple real-time location coordinates of visibility dissipation are obtained from the visibility dissipation distance. Environmental risk characteristics are calculated based on multiple actual coordinates, multiple diffusion visibility, and multiple real-time location coordinates of visibility dissipation.

[0013] This application also provides an image recognition system for a smart video safety helmet, which is equipped with multiple cameras facing different directions, including: The generation module generates a real-time construction site based on multiple construction site videos collected in real time from the smart video safety helmet worn by the user. It then uses a building grid to divide the real-time construction site into regions and generate a coordinate system. The first acquisition module acquires the actual coordinates of the user wearing the smart video safety helmet at the real-time construction site based on the coordinate system. The second acquisition module acquires multiple hazardous sources in the real-time construction site based on multiple construction site videos. The multiple hazardous sources include multiple static work units and multiple dynamic work units. The third acquisition module acquires fixed distance data based on actual coordinates and multiple static work units, and acquires transformed distance data based on actual coordinates and multiple dynamic work units; The fourth acquisition module obtains a comprehensive risk value based on fixed distance data and variable distance data; The judgment module determines whether the overall risk value is greater than the preset risk value; If the risk level exceeds the preset risk value, the user is determined to be in a dangerous area, and an alarm is issued.

[0014] Preferably, the generation module includes: The first acquisition unit acquires the orientation of multiple cameras mounted on each smart video safety helmet, acquires the corresponding acquisition angle based on the orientation of each camera, and acquires multiple real-time construction site videos based on the multiple acquisition angles. The first extraction unit aligns the initial frames of multiple real-time construction site videos with the same timestamp and extracts multiple construction site photos. The second acquisition unit acquires the original left and right images of the construction site photos within the same time stamp based on the acquisition angle, and aligns the polar lines of the original left and right images of the construction site photos after correction. The second extraction unit extracts the projection points of the same object from the corrected construction site photos to obtain depth maps, and generates a real-time construction site based on multiple depth maps.

[0015] Preferably, the third acquisition module includes: The third acquisition unit divides the construction site video into multiple video frame sequence lengths according to a preset video frame sequence length, and acquires multiple static work units based on the multiple video frame sequence lengths. The static work units include foundation pits, material piles, electrical components, and structures. The fourth acquisition unit acquires the corresponding contour coordinates of the foundation pit, material pile, electrical components and structures, acquires multiple distance data corresponding to each contour coordinate by comparing it with the actual coordinates, generates a location distance map based on the multiple distance data, extracts the shortest distance data in the location distance map and uses the shortest distance data as fixed distance data.

[0016] The beneficial effects of this invention are as follows: This invention achieves on-site three-dimensional reconstruction and positioning through multi-camera collaboration, combined with the identification and distance calculation of dynamic and static hazard sources, as well as the normalization and fusion of risk data, which can comprehensively and in real time assess risks, provide timely early warnings, and improve construction safety and management efficiency. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of a method flow according to an embodiment of the present invention.

[0018] Figure 2 This is a schematic diagram of the system structure according to an embodiment of the present invention.

[0019] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0020] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0021] like Figure 1 As shown, this application provides an image recognition method and system for a smart video safety helmet. The smart video safety helmet is equipped with multiple cameras facing different directions, including: S1. Based on the real-time collection of multiple construction site videos from the smart video safety helmet worn by the user, a real-time construction site is generated based on the multiple actual construction site videos. A coordinate system is generated by dividing the real-time construction site into regions using a building grid. S2. Obtain the actual coordinates of the user wearing a smart video safety helmet at the real-time construction site according to the coordinate system; specifically: at key locations on the construction site, such as tower crane base, foundation pit corner, and site benchmark pile, deploy highly reflective QR code / AR markers with built-in absolute coordinate information. After the worker enters the site wearing a safety helmet, he stops at the benchmark point to confirm the completion of the initial calibration, thereby achieving the binding of the relative model and the absolute coordinate system. During the operation, the markers are automatically scanned every 30 seconds to correct the accumulated error. If no marker is detected for 2 consecutive minutes, a recalibration prompt will be triggered to avoid the error from exceeding 1 meter. If the marker is obscured or damaged, temporary correction can also be achieved through feature matching of pre-stored key landmarks to control the error. S3. Based on multiple construction site videos, obtain multiple hazardous sources within the real-time construction site, including multiple static work units and multiple dynamic work units; S4. Obtain fixed distance data based on actual coordinates and multiple static work units, and obtain variable distance data based on actual coordinates and multiple dynamic work units; S5. Obtain a comprehensive risk value based on fixed distance data and variable distance data; S6. Determine whether the overall risk value is greater than the preset risk value; S7. If the risk level is greater than the preset risk value, the user is determined to be in a dangerous area and an alarm is issued.

[0022] As described in steps S1-S7 above, this invention generates a real-time three-dimensional or two-dimensional digital model of the construction site by collecting multi-view videos from an intelligent video safety helmet. Combined with building grid segmentation technology, the site is divided into precise coordinate grids (e.g., 1m×1m squares), achieving accurate mapping from physical space to digital space. The building grid coordinate system provides a unified benchmark for subsequent personnel and equipment positioning, avoiding positional deviations caused by differences in coordinate systems of different positioning technologies (e.g., GPS, UWB). Combined with video image recognition (e.g., safety helmet feature point matching) and auxiliary positioning technologies (e.g., UWB, BeiDou), this invention enables workers to... Centimeter-level positioning in digital construction sites, with positioning data updated synchronously with video acquisition frequency (e.g., 30Hz), can capture workers' rapid movements (e.g., climbing, running), providing real-time location data for dynamic risk assessment. Based on computer vision algorithms (e.g., YOLOv8, SSD), it can identify static hazards (e.g., unprotected foundation pits or piled construction materials) and dynamic hazards (e.g., moving tower crane hooks, moving engineering vehicles) in real time in the video. Fixed distance data (distance between personnel and static hazards) and variable distance data (relative speed and distance change rate between personnel and dynamic hazards) are combined to construct a dynamic risk model.For example, the risk is low when a worker is 2 meters away from a stationary pile of materials (safe distance 1 meter), but high when the worker is 5 meters away from a moving forklift at a relative speed of 3 meters per second. When the overall risk value exceeds the preset threshold, the worker is immediately notified through safety helmet audible and visual alarms, vibration alerts, etc., and the information is simultaneously pushed to the management platform to achieve multi-level early warning. Taking the hoisting operation of a large factory steel structure, where a tower crane lifts a steel truss to a high altitude for installation, and there are multiple overlapping operations such as welding and bolt tightening on site as an example: First, a digital model of the construction site is generated through video data. The site is divided into areas using a 1m×1m precision grid coordinate system. At the same time, the operation of the tower crane at a height of 20 meters is also included. The radius is marked as a red warning zone. The welder's precise location on the steel column (X=50, Y=60, Z=15) is determined in real-time using the safety helmet worn by the worker. Then, multiple construction site videos are used to identify several potential hazards in real-time. Static hazards include the unprotected edge of the steel column and paint buckets piled on the ground, with corresponding safety distances of 1 meter and 2 meters respectively. Dynamic hazards include the tower crane hook (X=45, Y=65, Z=25) at a speed of 1 m / s and coordinates (X=45, Y=65, Z=25), and the moving welding robot (X=55, Y=58, Z=0) at a speed of 0.5 m / s. Then, based on the actual coordinates and... Multiple static work units acquire fixed distance data, and multiple dynamic work units acquire variable distance data based on actual coordinates. For example, calculations show that the actual distance between the welder and the edge of the steel column is 0.8 meters, which is less than the basic safety distance of 1 meter. However, considering a high-altitude work coefficient of 1.5, this safety distance is adjusted to 1.5 meters. Meanwhile, the horizontal distance between the welder and the paint bucket is 8 meters, far exceeding the 2-meter safety distance. Regarding variable distance data, the vertical distance between the tower crane hook and the welder is 10 meters, and it is descending at a speed of 1 meter / second. It is expected that the vertical distance will shorten to 5 meters after 5 seconds. Therefore, combining the relevant data of fixed and variable distances, a preset... The weighted comprehensive risk value is calculated using the formula 0.7×(1-0.8 / 1.5)+0.3×(1 / 0.5), which yields 0.93. Since this value exceeds the preset threshold of 0.8, a Level I warning is triggered. An alarm is then issued, and the welder's safety helmet immediately emits an audible and visual alarm, displaying the message "Tower crane hook is approaching, evacuate immediately." The management platform also simultaneously displays the welder's location and the movement trajectory of the tower crane hook in real time. On-duty personnel remotely command the tower crane to pause its descent via the platform, and the welder, following the warning and instructions, promptly moves to a safe area in the center of the steel column, effectively mitigating potential risks.

[0023] In one embodiment, the step of generating an actual real-time construction scene based on multiple construction site videos collected in real time from a smart video safety helmet worn by the user includes: S201. Obtain the orientation of multiple cameras mounted on each smart video safety helmet, obtain the corresponding acquisition angle based on the orientation of each camera, and obtain multiple real-time construction site videos based on the multiple acquisition angles; S202. Align the initial frames of multiple real-time construction site videos with the same timestamp and extract multiple construction site photos. S203. Based on the acquisition angle, obtain the original left and right images of the construction site within the same timestamp, and align the polar lines of the original left and right images of the construction site after correction. S204. Extract projection points of the same object from the corrected construction site photos to obtain depth maps, and generate a real-time construction site based on multiple depth maps. To ensure accurate identification of projection points of the same object, this solution combines the geometric features (such as rectangular frames and linear extension directions), color features (such as green protective netting and yellow warning paint on equipment), and texture details (such as rust on steel pipes and seams in templates) of the extracted objects to construct a multi-dimensional feature vector. Projection points are searched only within the range of objects with similar features. For areas with highly repetitive textures, histogram equalization is used to highlight subtle differences, morphological dilation is used to strengthen edges, and the local search range is limited to ±10 pixels to reduce cross-regional mismatches. Simultaneously, a lightweight YOLOv8-tiny optimization model is integrated for real-time... The system identifies object categories and verifies matching results, eliminating mismatched pairs of different categories. To ensure battery life, the helmet only handles basic computations such as video capture, marker recognition, simple feature extraction, and local early warning. It uses a low-power processor (such as NVIDIA Jetson Orin NX16GB, with a computing power of 100 TOPS and a weight of 45g) and four miniature wide-angle cameras (12g per camera, 48g total). The core complex computations (stereo reconstruction, depth map generation, and dynamic tracking) are migrated to the cloud / edge server, and the compressed feature data is transmitted via 5G / Wi-Fi 6.

[0024] As described in steps S201-S203 above, this invention achieves 360° coverage of the construction site without blind spots by acquiring the orientation of multiple cameras on the smart video safety helmet (e.g., front, rear, left, and right four-way cameras) and combining them with the collected angle parameters (e.g., horizontal viewing angle 120°, vertical viewing angle 90°). The initial frames of multiple video streams are aligned to the same timestamp (e.g., UTC time accurate to milliseconds) to ensure that the image data processed subsequently has temporal consistency and avoids 3D reconstruction errors caused by time deviation. Left and right views are selected from multi-view photos with the same timestamp (e.g., left image from the front camera, right image from the right camera), and lens distortion (e.g., radial distortion, tangential distortion) is corrected by camera calibration algorithms (e.g., Zhang Zhengyou calibration) to keep straight lines in the image straight and improve the accuracy of subsequent matching. The projection points of the same object (e.g., bolts, rebar endpoints) are extracted from the corrected left and right images, and pixel-level depth values ​​are calculated using the parallax formula (depth = baseline × focal length / parallax) to generate a dense depth map with centimeter-level depth accuracy. The depth maps from multiple perspectives are fused to generate a complete real-time construction site.

[0025] Taking the installation of curtain walls in a super high-rise office building, where workers operate in suspended pods at high altitudes, surrounded by tower cranes, construction elevators, and other equipment, and requiring real-time monitoring of the working environment and potential hazards, as an example: A smart video safety helmet is equipped with three cameras facing different directions: front, rear, and left. The front camera faces the curtain wall surface, capturing angles covering a horizontal range of ±60° and a vertical range of ±45°. The rear camera faces behind the suspended pod to monitor the tower crane's operation in real time. The left camera is specifically positioned facing the adjacent suspended pod. Simultaneously, the system acquires the specific orientation information of each camera in real time, clearly identifying the positive X-axis direction (i.e., the direction of the curtain wall surface) corresponding to the front camera. The rear camera corresponds to the negative X-axis, and the left camera corresponds to the positive Y-axis. The video streams from all cameras are then synchronized to UTC time, ensuring a time error within 1 millisecond. Photos of the construction site from three perspectives (front, rear, and left) are extracted every 2 seconds, and their timestamps are recorded. The installation status of the curtain wall keel (from the front view), the specific position of the tower crane hook (from the rear view), and the spacing between adjacent suspended platforms (from the left view) are observed from the extracted photos, providing complete and synchronized visual data for subsequent image processing. Next, for the front and left views, fisheye distortion is corrected using a camera calibration algorithm. Then, the epipolar lines of both images are adjusted to the horizontal direction. Taking the curtain wall vertical keel as an example, the epipolar line deviation in the left and right views before correction is corrected. After correction, the railings of adjacent suspended platforms in the left view and the corresponding railings in the front view are precisely on the same horizontal epipolar line, thus reducing the difficulty of subsequent parallax calculations and improving data accuracy. Finally, based on the front and left views that have been calibrated and aligned with the epipolar lines, a depth map of the curtain wall is generated by calculating the parallax between the two images. With the help of this depth map, the three-dimensional coordinates of components such as keel and bolts can be accurately obtained, and a complete three-dimensional point cloud model including the suspended platform, curtain wall and tower crane can be constructed, providing high-precision support for the digital presentation of the construction site.

[0026] In one embodiment, the step of obtaining fixed-distance data based on actual coordinates and multiple static work units includes: S301. The construction site video is split into multiple video frame sequence lengths according to a preset video frame sequence length. Multiple static work units are obtained according to the multiple video frame sequence lengths. The static work units include foundation pits, material piles, electrical components and structures. S302. Obtain the corresponding contour coordinates of the foundation pit, material pile, electrical components and structures. Obtain multiple distance data corresponding to each contour coordinate by comparing it with the actual coordinates. Generate a location distance map based on the multiple distance data. Extract the shortest distance data in the location distance map and use the shortest distance data as the fixed distance data.

[0027] As described in steps S301-S302 above, this invention can capture the spatial location and shape of static work units (foundation pits, material piles, etc.) at high frequency by splitting the construction site video according to a preset frame sequence (e.g., every 10 frames as a sequence), avoiding omissions in element recognition due to video continuity. It combines color features (e.g., the green protective netting of foundation pit support), texture features (the granular feel of material piles), and geometric features (the rectangular outline of structures) for comprehensive recognition, improving the recognition accuracy in complex scenes. It extracts the outline coordinates of static work units (e.g., the coordinates of the polygon vertices of the foundation pit edge) and calculates the spatial distance with the actual coordinates of personnel (e.g., the GPS coordinates of safety helmets). It supports multiple measurement methods such as Euclidean distance and Manhattan distance, with an accuracy of up to 0.3 meters. For example, when calculating the shortest distance between a worker and the edge of the foundation pit, it can be accurate to the vertex of a specific section of the support structure, generating a location distance map (e.g., a heat map) to visualize the distance distribution between personnel and each static unit, and intuitively displaying risk areas (e.g., red indicates distance < safety threshold, green indicates safety), facilitating the rapid location of high-risk points.

[0028] In one embodiment, after the step of obtaining fixed distance data based on actual coordinates and multiple static work units, the method includes: S401. Sequentially acquire images of the foundation pit, material pile, electrical components, and structure based on the corresponding video frames. S402. Simultaneously grayscale the foundation pit photo and the structure image and extract the edge features of the corresponding photos. Obtain edge pixels based on the edge features and generate the foundation pit edge outline and the structure edge outline based on the edge pixels. S403. Decompose the material pile photo into blocks with similar textures, merge the material pile connected component contour point sequences, and connect the contour point sequences in sequence to obtain the material pile edge contour line. S404. Automatically binarize the image of the electrical component to obtain the electrical component body and the interference background, extract the electrical component body and connect the broken edges of the electrical component body to obtain the edge contour line of the electrical component.

[0029] As described in steps S401-S404 above, this invention accurately extracts corresponding images from video frames according to preset categories (foundation pit, material pile, electrical components, and structures), avoiding interference from mixed scenes. For example, in complex construction areas, the system can automatically filter video frames containing only foundation pits, excluding dynamic objects such as tower cranes and vehicles, improving the accuracy of subsequent contour extraction. It supports custom acquisition strategies to adapt to the needs of different construction stages. Grayscale processing eliminates color and reduces RGB three-channel data to a single channel, improving the robustness of edge detection. Edge feature extraction combined with morphological operations (dilation, erosion) can filter noise and connect broken edges, generating continuous contour lines of foundation pits and structures with a contour integrity of over 90%. The contour lines generated based on edge pixels can accurately describe the geometric shape of objects (such as the polygonal contour of foundation pits and the rectangular frame of structures), providing geometric parameters (such as side length, angle, and area) for subsequent distance calculation and risk assessment. For example, the outline of a structure can be directly used to calculate its occupied space and determine whether it obstructs safety passages. Based on texture similarity, material pile images are segmented into several blocks, with consistent textures within the same block (e.g., linear texture for steel pipe piles, rectangular texture for template piles), avoiding outline confusion caused by mixing different materials. For instance, when steel pipes and templates are mixed, the system can segment them into two blocks based on texture differences, generating outlines separately. By merging blocks with similar textures through connected component analysis, discrete points caused by unevenness on the material pile surface are eliminated, generating a continuous sequence of outline points with an outline closure rate of over 95%. Based on the bimodal characteristics of pixel grayscale histograms, electrical component images are divided into body and background without the need for manual threshold setting, adapting to different lighting conditions. For breaks in the edges of electrical components caused by stains or occlusion (e.g., dust covering the surface of a distribution box), morphological analysis is used... The operation (skeletonization, edge tracking) connects the broken points and generates closed contour lines with an edge continuity rate of over 95%. Taking a substation expansion project as an example, where there are distribution cabinets, cable trenches (foundation pits), steel piles (material piles), and temporary scaffolding (structures) on site, and the contours of each static unit need to be accurately extracted to assess safe distances: First, images of different types of static work units are extracted from the construction video. The foundation pit image mainly shows the cable trench support structure (i.e., steel sheet piles), the material pile image focuses on the steel pile with linear texture as the core content, with a clear hardened ground background, the electrical component image focuses on the distribution cabinet with a rectangular outline and yellow warning signs, and the structure image shows the scaffolding with a grid structure. Similarly, there is no mechanical interference, which provides clear and clean image materials for the subsequent contour processing of various types of static work units.

[0030] For cable trench images, grayscale processing is first performed to eliminate color interference, and then the Canny operator is used to accurately extract image edges. For edge breaks, morphological closing operations are used to repair and connect them, ultimately generating a regular trapezoidal outline. After grayscale processing, scaffolding images are directly extracted using edge extraction to quickly generate rectangular outlines that conform to their structural characteristics, effectively ensuring accurate restoration of the spatial form of the foundation pit and structures. For steel stack images, GLCM (Gray-Level Co-occurrence Matrix) texture segmentation technology is first used to divide them into multiple linear texture blocks based on texture similarity. Then, connected component merging is performed on these blocks to remove discrete noise points. Subsequently, the merged edge point sequence is extracted, and an irregular polygonal outline conforming to the actual stacking shape of the steel stack is generated based on this sequence. For electrical cabinet images, the Otsu binarization algorithm is first used to automatically segment the cabinet body and interfering background. Then, morphological closing operations are used to repair edge breaks caused by latch obstruction, ultimately generating a complete rectangular outline. After completing the outline extraction of each static work unit, the risk assessment application phase begins. The system converts image coordinates into real-world physical coordinates (unit: meters) through camera calibration or a spatial mapping model. Taking the world coordinates of the cable trench outline vertices as (X=10m, Y=10m), (X=50m, Y=10m), (X=55m, Y=30m), and (X=5m, Y=30m), and the worker's current position as (X=30m, Y=20m), with a safety threshold set at 1.5 meters, the system uses a point-to-polygon shortest distance algorithm (e.g., calculating the perpendicular / endpoint distance from the point to each edge segment and taking the minimum value) to find the shortest distance between the worker and the cable trench outline to be approximately 2.24 meters. This distance is greater than the 1.5-meter safety threshold, indicating that the worker is within the safe zone of the cable trench. Simultaneously, the shortest distance between the worker and the distribution cabinet outline is calculated to be 3.8 meters (example value, needs to be calculated based on actual calibration coordinates), which is also significantly greater than the preset safety distance, further confirming that there is currently no safety risk caused by proximity to a static hazard source.

[0031] In traditional methods, the calculated distance is 10 meters (actually 14.14 meters) because the cable trench outline is missing a vertex. Although no danger is misjudged, the insufficient outline accuracy leads to errors in the distance redundancy calculation. This application ensures the reliability of risk assessment through accurate outline extraction.

[0032] In one embodiment, the step of obtaining transformed distance data based on actual coordinates and multiple dynamic work units includes: S501. The construction site video is split into multiple video frame sequence lengths according to a preset video frame sequence length, and multiple dynamic work units are obtained according to the multiple video frame sequence lengths. The dynamic work units include construction vehicles and construction equipment. S502. Obtain the reference position and construction trajectory of the construction equipment; obtain multiple construction change coordinates based on the construction trajectory; obtain the construction coverage edge line of the construction equipment based on the initial coordinates and multiple construction change coordinates; obtain multiple edge coordinates based on the construction coverage edge line and the real-time construction site; obtain the first distance data based on the multiple edge coordinates and the actual coordinates. S503. Obtain the initial coordinates and movement path of the construction vehicle. Within a preset time, obtain multiple change position coordinates of the construction vehicle according to the initial coordinates of the construction vehicle and the multiple change position coordinates of the construction vehicle. Obtain the movement speed of the construction vehicle according to the initial coordinates of the construction vehicle and the multiple change position coordinates of the construction vehicle. S504. Obtain the second change distance data based on the actual coordinates, the initial coordinates of the construction vehicles, and the coordinates of the changing positions of multiple construction vehicles; S505. Generate transformed distance data based on the first distance data and the second distance data.

[0033] As described in steps S501-S505 above, this invention, through video frame sequence splitting and target detection algorithms (such as DeepSORT), synchronously tracks the spatial positions of construction vehicles (such as concrete mixer trucks) and construction machinery (such as tower cranes and excavators). Based on motion features, it distinguishes different types of dynamic units to provide typified data support for subsequent trajectory analysis. The position coordinates of dynamic units are obtained according to video frame sequences (e.g., 200 milliseconds / time), capturing the instantaneous state of fast-moving objects (such as sudden stops of tower crane hooks or sharp turns of vehicles). The trajectory point density reaches 5 points / second, meeting the tracking requirements of high-speed motion scenarios. The invention also uses a reference position (such as the initial parking point of the excavator) and the construction track... The system calculates the construction coverage edge line based on traces (such as excavation paths), which can accurately describe the working range of machinery (such as the excavation radius of an excavator or the lifting range of a tower crane). The edge line positioning accuracy is within 0.5 meters. The construction coverage edge line is updated in real time with the movement of machinery (such as the coverage edge rotating synchronously when the excavator rotates), providing a dynamic benchmark for assessing the safe distance between personnel and machinery, avoiding misjudgments caused by static boundaries. By using the initial coordinates (such as the position of a vehicle when it enters the site) and high-frequency coordinate points on the movement path (such as one point every 0.2 seconds), the system uses the difference method to calculate the real-time speed (such as v=ΔtΔs), with a speed accuracy of 0.1 meters / second. It can distinguish the vehicle's idling, driving, and emergency stop states. For example, when a concrete mixer truck turns, its speed decreases from 5 m / s to 2 m / s. The system can capture this speed change in real time. Vehicle speed data can be used to predict future positions (e.g., predicting the vehicle's position 10 seconds later based on the current speed and direction), providing a time window for collision warnings (e.g., a 5-second advance warning). The first distance data, based on the geometric relationship between the personnel's current position and the edge of the equipment's coverage, calculates the minimum Euclidean distance, reflecting static spatial risk. The second change distance data, through the relative motion state of personnel and vehicles (e.g., speed difference, change of direction), calculates dynamic risk indicators. The static and dynamic data are linearly combined, and the formula is: d s =d1+d2·K; Wherein, the d s The distance data is represented by d1, which represents the first distance data, d2, which represents the second distance data, and K represents the fixed scaling factor.

[0034] When applying this formula, a preset threshold needs to be set according to different application scenarios. For example, when d1 < the safety threshold, the second change distance data d2 should be used as the output first, and conditional judgments should be made (such as "if the vehicle is approaching, then d1 < d2 ... s =min(d1,d2)) merges the two types of data to avoid subjective weighting.

[0035] In one embodiment, the step of obtaining a comprehensive risk value based on fixed distance data and transformed distance data includes: S601. Obtain a fixed distance data vector based on the fixed distance data, and normalize the fixed distance data vector to obtain a normalized value of the fixed distance data vector. S602. Obtain a transformation distance data vector based on the transformation distance data, and normalize the transformation distance data vector to obtain a normalized value of the transformation distance data vector. S603. Based on the weights corresponding to the normalized values ​​of the fixed-distance data vector and the normalized values ​​of the transformed-distance data vector, obtain the comprehensive risk value according to the corresponding weights.

[0036] As described in steps S601-S603 above, this invention converts fixed distance data of different dimensions (such as the distance between personnel and the foundation pit of 1.2 meters and the distance between personnel and the material pile of 5 meters) into dimensionless vector normalized values ​​(range 0-1), thereby eliminating the interference of dimensional differences on risk assessment. The normalization process converts absolute distance into relative risk level (such as the smaller the normalized value when the distance is closer, the higher the risk), which makes it easier to set a unified risk threshold (such as a normalized value < 0.3 indicating high risk). It also converts variable distance data (such as relative speed of 3 m / s and distance change rate of -2 m / s) into normalized values ​​to reflect the dynamic trend of risk change. For example, a higher relative speed results in a smaller normalized value and a higher risk; a faster rate of distance reduction also results in a smaller normalized value and a higher risk. The system supports multi-dimensional dynamic data fusion (e.g., simultaneously considering vehicle speed and the rate of change of equipment coverage edges) to construct a dynamic risk vector, capturing the time-sensitive characteristics of risk. Normalized transformed distance data and fixed distance data reside in the same metric space (0-1), facilitating subsequent weighted fusion and achieving a comprehensive risk assessment of "spatial distance + time change." Based on the normalized vectors of fixed and transformed distances, a comprehensive risk value is calculated through weight allocation (e.g., a fixed risk weight of 0.6 and a dynamic risk weight of 0.4), balancing static spatial risk and dynamic temporal risk. For example, when personnel are close to the pit (high fixed risk) but no dynamic objects are nearby (low dynamic risk), the comprehensive risk value is dominated by the fixed risk. The system supports adaptive weight adjustment (e.g., dynamically adjusting weights based on construction stage and hazard type) to improve the model's adaptability to complex scenarios. For example, the dynamic risk weight is increased to 0.6 for high-altitude operations and 0.7 for ground operations.

[0037] In the scenario of risk assessment at a construction site, taking a fixed distance of d=10 meters between a worker and a static hazard source (such as a material pile) and a safety threshold of S=15 meters for such a hazard source as an example, based on the logic that the closer the distance, the higher the risk, the static risk normalization formula Rd=1-d / S (applicable only when d≤S, and the risk is 0 when d>S) is used to calculate the static risk normalization value as 1-10 / 15≈0.33, which accurately reflects the current static risk level. For dynamic hazards (such as construction vehicles), their relative speed to workers is u = 4 m / s. Considering the characteristic that higher speed means higher risk, the dynamic risk normalization formula Ru = min(1, u / umax) (umax = 5 m / s is the safe speed threshold, and the value should not exceed 1 to avoid risk value overflow) is used. The calculated speed risk value is 4 / 5 = 0.80. At the same time, the distance change rate between the vehicle and the worker is -4 m / s (the negative sign indicates that the distance is continuously shortening), and the absolute value of the safe change rate threshold is 3 m / s. Similarly, the approach rate risk value is calculated as min(1, 4 / 3) = 1.00, indicating that the current distance shortening rate has exceeded the safe threshold, and the dynamic risk is extremely high. Taking the average of the speed risk value and the approach rate risk value, the normalized value of the transformed distance vector is obtained as: (0.80 + 1.00) / 2 = 0.90. Based on historical accident data and risk distribution characteristics at the construction site, and considering that dynamic hazards are prone to triggering sudden accidents, a fixed risk weight of wd=0.4 and a variable risk weight of wu=0.6 (dynamic risk is dominant) are set. Substituting these values ​​into the comprehensive risk value formula R=wd×Rd+wu×Rtransformation, the comprehensive risk value R=0.4×0.33+0.6×0.90=0.132+0.54=0.672 is calculated. According to the risk level classification standard (0-0.3 is low risk, 0.3-0.7 is medium risk, and >0.7 is high risk), 0.672 falls within the medium-high risk range. The system immediately triggers a yellow warning via the smart safety helmet and simultaneously issues an audible and visual alert: "Construction vehicle ahead is approaching at high speed, please take emergency evasive action!" After receiving the warning, if the worker evacuates to a distance of 15 meters from the static hazard source within 2 seconds (d=15 meters, satisfying d=S, the static risk normalization value drops to 0), and the construction vehicle decelerates and adjusts its path in time, reducing the relative speed to 1 m / s and the approach rate to close to 0, the dynamic risk normalization value is recalculated as (1 / 5+0) / 2=0.10. At this time, the comprehensive risk value is updated to R=0.4×0+0.6×0.10=0.06, successfully dropping to the low-risk range, and the warning is automatically lifted.Compared to traditional risk assessment methods that rely solely on a fixed distance of 10 meters (less than the safety threshold of 15 meters) to determine "mild static risk," this approach ignores the dynamic threat of vehicles approaching at high speeds. If not intervened in time, the distance between the worker and the vehicle will shrink to 2 meters in about 2.5 seconds, and the worker will be within the tower crane's operating area, which could easily lead to a double high-risk accident of "vehicle collision + tower crane operation." This solution, through deep integration of dynamic and static risk data, identifies compound risks in advance and triggers warnings, effectively achieving proactive risk avoidance and ensuring the safety of construction personnel.

[0038] In one embodiment, after the step of obtaining multiple hazard sources within the real-time construction site based on multiple construction site videos, the following steps are included: S701. Obtain real-time video information of the actual construction site based on the smart video safety helmet, split the real-time video information into multiple video frame sequence lengths according to the preset video frame sequence length, and obtain multiple diffusion visibilitys based on the multiple video frame sequence lengths. S702. Obtain the visibility dissipation distance based on multiple diffuse visibility values, and obtain the visibility dissipation speed and multiple real-time location coordinates of visibility dissipation based on the visibility dissipation distance; S703. Calculate environmental risk characteristics based on multiple actual coordinates, multiple diffusion visibility, and multiple real-time location coordinates of visibility dissipation.

[0039] As described in steps S701-S703 above, this invention uses an intelligent video safety helmet to capture real-time video of the construction site. Combined with preset frame sequence splitting technology, it can acquire environmental images at high frequency (e.g., 25 frames per second), ensuring the temporal continuity of the data. Based on the video frames, diffusion visibility is calculated. Diffusion visibility is mainly the visibility diffusion characteristics caused by dust and smoke. It can simultaneously analyze the visibility changes in multiple areas of the image (e.g., tower crane operation area, material stacking area). Compared with single-point sensor detection, it has a wider coverage and can capture the spatial distribution characteristics of environmental risks. Based on multiple diffusion visibility data, a visibility dissipation curve can be fitted through time series analysis (e.g., Kalman filtering) to calculate the dissipation distance (e.g., the boundary distance from dust diffusion to visibility ≤ 10 meters) and dissipation speed (e.g., 2 meters per second diffusion). Furthermore, based on the calculation of the real-time location coordinates of visibility dissipation (e.g., using GPS and image matching technology), the spatial movement trajectory of pollution clouds can be dynamically tracked, achieving accurate positioning of risk sources. Risk assessment and risk warning are constructed based on the actual coordinates of the workers, diffusion visibility, and dissipation location coordinates.

[0040] Specifically, dynamic risk characteristics are calculated based on multiple actual coordinates, visibility dissipation distances, and multiple real-time visibility dissipation locations, wherein the calculation formula is: ; Wherein, the Rhj The environmental risk characteristics are represented by: n represents the total number of video frame sequence samples; vh represents diffusion visibility; vmax represents the maximum theoretical visibility of the site; ph represents the real-time location coordinates; psafe represents the real-time location coordinates of visibility dissipation; Dmax represents the maximum spatial scale of the construction site; and α and β represent weighting coefficients.

[0041] Example: Steel structure welding is underway at a construction site, and welding fumes are causing localized visibility reduction. The smart safety helmet is monitoring and calculating environmental risks in real time.

[0042] Assume Vmax=100m, Dmax=50m, α=0.6, β=0.4; Analyze 3 frames of data at a certain moment (N=3): Frame 1: Diffuse visibility vh=30m, worker's real-time position coordinates ph=(10,20), visibility dissipation safety position coordinates psafe=(12,22); Frame 2: Diffuse visibility vh=50m, worker real-time position coordinates ph=(15, 18), visibility dissipation safety position coordinates psafe=(14, 19); Frame 3: Diffuse visibility vh=10m, worker real-time position coordinates ph=(8,15), visibility dissipation safety position coordinates psafe=(12,22).

[0043] Calculation steps: Frame 1: 1 - 30 / 100 = 0.7 2.83 / 50 = 0.0566; Item value: 0.6 × 0.7 + 0.4 × 0.0566 = 0.4426; Frame 2: 1 - 40 / 100 = 0.5 1.41 / 50 = 0.0282; Term value: 0.6 × 0.5 + 0.4 × 0.0282 = 0.3113; Frame 3: 1 - 10 / 100 = 0.1 8.06 / 50 = 0.1612; Term value: 0.6 × 0.9 + 0.4 × 0.1612 = 0.6045; Final risk value: (0.4426+0.3113+0.6045) / 3 ≈ 0.4528; The risk value falls within the medium-risk range (preset: low risk <0.3, medium risk 0.3–0.6, high risk >0.6), indicating that the current smoke diffusion has significantly impacted operational safety. In particular, with visibility in frame 3 at only 10m, far below the recommended safe visibility threshold for welding operations (typically ≥30m), the system should immediately trigger a local high-risk warning, issuing an audible and visual alert via safety helmet: "Smoke concentration too high, please suspend work, evacuate to a ventilated area, or wear a protective mask." Simultaneously, the warning information should be pushed to the management platform, recommending adjustments to the welding sequence or activation of local ventilation equipment to prevent secondary accidents caused by smoke accumulation leading to suffocation or obstructed vision.

[0044] like Figure 2 As shown, this application also provides an image recognition system for an intelligent video safety helmet, comprising: Module 1 generates a real-time construction site based on multiple construction site videos collected in real time from the smart video safety helmet worn by the user. It then uses a building grid to divide the real-time construction site into regions and generate a coordinate system. The first acquisition module 2 acquires the actual coordinates of the user wearing the smart video safety helmet at the real-time construction site according to the coordinate system. The second acquisition module 3 acquires multiple hazardous sources in the real-time construction site based on multiple construction site videos. The multiple hazardous sources include multiple static work units and multiple dynamic work units. The third acquisition module 4 acquires fixed distance data based on actual coordinates and multiple static work units, and acquires transformed distance data based on actual coordinates and multiple dynamic work units; The fourth acquisition module 5 obtains the comprehensive risk value based on fixed distance data and variable distance data; Module 6 determines whether the overall risk value is greater than the preset risk value; If the risk level exceeds the preset risk value, the user is determined to be in a dangerous area, and an alarm is issued.

[0045] In one embodiment, the generation module includes: The first acquisition unit acquires the orientation of multiple cameras mounted on each smart video safety helmet, acquires the corresponding acquisition angle based on the orientation of each camera, and acquires multiple real-time construction site videos based on the multiple acquisition angles. The first extraction unit aligns the initial frames of multiple real-time construction site videos with the same timestamp and extracts multiple construction site photos. The second acquisition unit acquires the original left and right images of the construction site photos within the same time stamp based on the acquisition angle, and aligns the polar lines of the original left and right images of the construction site photos after correction. The second extraction unit extracts the projection points of the same object from the corrected construction site photos to obtain depth maps, and generates a real-time construction site based on multiple depth maps.

[0046] In one embodiment, the third acquisition module includes: The third acquisition unit divides the construction site video into multiple video frame sequence lengths according to a preset video frame sequence length, and acquires multiple static work units based on the multiple video frame sequence lengths. The static work units include foundation pits, material piles, electrical components, and structures. The fourth acquisition unit acquires the corresponding contour coordinates of the foundation pit, material pile, electrical components and structures, acquires multiple distance data corresponding to each contour coordinate by comparing it with the actual coordinates, generates a location distance map based on the multiple distance data, extracts the shortest distance data in the location distance map and uses the shortest distance data as fixed distance data.

[0047] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in this application and in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0048] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.

[0049] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. An image recognition method for a smart video safety helmet, applied to a smart video safety helmet, wherein the smart video safety helmet is equipped with multiple cameras facing different directions, characterized in that, include: Based on real-time video of multiple construction sites collected by the smart video safety helmet worn by the user, a real-time construction site is generated, and a coordinate system is generated by dividing the real-time construction site into regions using a building grid. The actual coordinates of the user wearing the smart video safety helmet at the real-time construction site are obtained based on the coordinate system. Multiple hazardous sources within the real-time construction site were obtained from multiple construction site videos. These multiple hazardous sources include multiple static work units and multiple dynamic work units. Fixed distance data is obtained based on actual coordinates and multiple static work units, and variable distance data is obtained based on actual coordinates and multiple dynamic work units; A comprehensive risk value is obtained based on fixed distance data and variable distance data; Determine whether the overall risk value is greater than the preset risk value; If the risk level exceeds the preset risk value, the user is determined to be in a dangerous area, and an alarm is issued.

2. The image recognition method for the intelligent video safety helmet according to claim 1, characterized in that, The steps of generating a real-time construction scene based on multiple construction site videos collected in real time using the smart video safety helmet worn by the user include: The orientation of multiple cameras mounted on each smart video safety helmet is obtained, the corresponding acquisition angle is obtained according to the orientation of each camera, and multiple real-time construction site videos are obtained according to the multiple acquisition angles; Align the initial frames of multiple real-time construction site videos with the same timestamp to extract multiple construction site photos; Based on the collection angle, obtain the original left and right images of the construction site within the same timestamp, and then align the polar lines of the original left and right images of the construction site after correction. Depth maps are obtained by extracting the projection points of the same object from the corrected construction site photos, and a real-time construction site is generated based on multiple depth maps.

3. The image recognition method for the intelligent video safety helmet according to claim 1, characterized in that, The step of obtaining fixed distance data based on actual coordinates and multiple static work units includes: The construction site video is split into multiple video frame sequence lengths according to a preset video frame sequence length. Multiple static work units are obtained based on the multiple video frame sequence lengths. The static work units include foundation pits, material piles, electrical components, and structures. Obtain the corresponding contour coordinates of the foundation pit, material pile, electrical components and structures. Obtain multiple distance data corresponding to each contour coordinate by comparing it with the actual coordinates. Generate a location distance map based on the multiple distance data. Extract the shortest distance data in the location distance map and use the shortest distance data as the fixed distance data.

4. The image recognition method for the intelligent video safety helmet according to claim 3, characterized in that, After the step of obtaining fixed distance data based on actual coordinates and multiple static work units, the following steps are included: Based on the corresponding video frames, sequentially obtain images of the foundation pit, material pile, electrical components, and structures; The photos of the foundation pit and the building structure are simultaneously grayed out and the edge features of the corresponding photos are extracted. The edge pixels are obtained based on the edge features and the edge outlines of the foundation pit and the building structure are generated based on the edge pixels. The material pile photo is decomposed into blocks with similar textures, the material pile connected component contour point sequences are merged, and the contour point sequences are connected sequentially to obtain the material pile edge contour line. The electrical component image is automatically binarized to obtain the component body and the interfering background. The component body is extracted and the edges of the broken component body are connected to obtain the component edge contour line.

5. The image recognition method for the intelligent video safety helmet according to claim 1, characterized in that, The step of obtaining transformed distance data based on actual coordinates and multiple dynamic work units includes: The construction site video is split into multiple video frame sequence lengths according to a preset video frame sequence length. Multiple dynamic work units are obtained based on the multiple video frame sequence lengths. The dynamic work units include construction vehicles and construction equipment. The system acquires the reference position and construction trajectory of the construction equipment, obtains multiple construction change coordinates based on the construction trajectory, acquires the construction coverage edge line of the construction equipment based on the initial coordinates and multiple construction change coordinates, acquires multiple edge coordinates based on the construction coverage edge line and the real-time construction site, and acquires the first distance data based on the multiple edge coordinates and the actual coordinates. Obtain the initial coordinates and movement path of the construction vehicle. Within a preset time, based on the initial coordinates of the construction vehicle, move along the movement path and sequentially obtain the coordinates of multiple changes in the position of the construction vehicle. Based on the initial coordinates of the construction vehicle and the coordinates of multiple changes in the position of the construction vehicle, obtain the movement speed of the construction vehicle. The second change distance data is obtained based on the actual coordinates, the initial coordinates of the construction vehicles, and the coordinates of the changing positions of multiple construction vehicles. Transformed distance data is generated based on the first distance data and the second distance data.

6. The image recognition method for the intelligent video safety helmet according to claim 1, characterized in that, The step of obtaining the comprehensive risk value based on fixed distance data and transformed distance data includes: A fixed distance data vector is obtained based on the fixed distance data, and the fixed distance data vector is normalized to obtain a normalized value of the fixed distance data vector. Based on the transformed distance data, a transformed distance data vector is obtained, and the transformed distance data vector is normalized to obtain a normalized value of the transformed distance data vector. Based on the weights corresponding to the normalized values ​​of the fixed-distance data vector and the transformed-distance data vector, a comprehensive risk value is obtained.

7. The image recognition method for the intelligent video safety helmet according to claim 1, characterized in that, After the step of obtaining multiple hazard sources in the real-time construction site based on multiple construction site videos, the following steps are included: The real-time video information of the actual construction site is obtained by the smart video safety helmet. The real-time video information is split into multiple video frame sequence lengths according to the preset video frame sequence length. Multiple diffusion visibilitys are obtained according to the multiple video frame sequence lengths. The visibility dissipation distance is obtained from multiple diffuse visibility values, and the visibility dissipation speed and multiple real-time location coordinates of visibility dissipation are obtained from the visibility dissipation distance. Environmental risk characteristics are calculated based on multiple actual coordinates, multiple diffusion visibility, and multiple real-time location coordinates of visibility dissipation.

8. An image recognition system for an intelligent video safety helmet, characterized in that, include: The generation module generates a real-time construction site based on multiple construction site videos collected in real time from the smart video safety helmet worn by the user. It then uses a building grid to divide the real-time construction site into regions and generate a coordinate system. The first acquisition module acquires the actual coordinates of the user wearing the smart video safety helmet at the real-time construction site based on the coordinate system. The second acquisition module acquires multiple hazardous sources in the real-time construction site based on multiple construction site videos. The multiple hazardous sources include multiple static work units and multiple dynamic work units. The third acquisition module acquires fixed distance data based on actual coordinates and multiple static work units, and acquires transformed distance data based on actual coordinates and multiple dynamic work units; The fourth acquisition module obtains a comprehensive risk value based on fixed distance data and variable distance data; The judgment module determines whether the overall risk value is greater than the preset risk value; If the risk level exceeds the preset risk value, the user is determined to be in a dangerous area, and an alarm is issued.

9. The image recognition system for an intelligent video safety helmet according to claim 8, characterized in that, The generation module includes: The first acquisition unit acquires the orientation of multiple cameras mounted on each smart video safety helmet, acquires the corresponding acquisition angle based on the orientation of each camera, and acquires multiple real-time construction site videos based on the multiple acquisition angles. The first extraction unit aligns the initial frames of multiple real-time construction site videos with the same timestamp and extracts multiple construction site photos. The second acquisition unit acquires the original left and right images of the construction site photos within the same time stamp based on the acquisition angle, and aligns the polar lines of the original left and right images of the construction site photos after correction. The second extraction unit extracts the projection points of the same object from the corrected construction site photos to obtain depth maps, and generates a real-time construction site based on multiple depth maps.

10. The image recognition system for an intelligent video safety helmet according to claim 8, characterized in that, The third acquisition module includes: The third acquisition unit divides the construction site video into multiple video frame sequence lengths according to a preset video frame sequence length, and acquires multiple static work units based on the multiple video frame sequence lengths. The static work units include foundation pits, material piles, electrical components, and structures. The fourth acquisition unit acquires the corresponding contour coordinates of the foundation pit, material pile, electrical components and structures, acquires multiple distance data corresponding to each contour coordinate by comparing it with the actual coordinates, generates a location distance map based on the multiple distance data, extracts the shortest distance data in the location distance map and uses the shortest distance data as fixed distance data.