Limited space communication and security detection method, system and equipment

By constructing an occlusion perception map and dynamic communication topology, and combining optical flow method and BIM model, the problem of identifying abnormal behavior under fixed deployment and occlusion in a limited space was solved, achieving efficient security detection and event localization.

CN120808446AInactive Publication Date: 2025-10-17HAINAN SHENGDE ELECTRIC POWER TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511012290.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-10-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies have problems such as fixed deployment in limited spaces, non-robust communication, inability to identify abnormal behavior under occlusion, and inaccurate event positioning, resulting in poor security detection results.

Method used

An occlusion perception map is constructed, and a 3D semantic map is established by combining ranging and visual SLAM. A communication mesh topology is dynamically generated, and the human posture trajectory under occlusion is restored by optical flow. Spatial alignment is performed by combining BIM model, and a triple event structure is generated to realize edge response and data upload.

Benefits of technology

It improves the continuity and completeness of abnormal behavior detection, realizes the precise binding of behavioral events and spatial components, provides data support, and provides reliable data support for edge response and scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808446A_ABST
    Figure CN120808446A_ABST
Patent Text Reader

Abstract

The invention discloses a finite space communication and security detection method, system and device, and relates to the technical field of image recognition, and the method comprises the steps: building a shielding perception graph, and completing the building of an initial three-dimensional semantic map and a device topological graph through the combination of distance measurement and visual SLAM; a communication Mesh topology is dynamically generated based on sight distance shielding information, path reconstruction is triggered when shielding occurs between nodes, and edge nodes are scheduled to perform image acquisition and transmission on the basis of communication guarantee; and restoring the shielded human body posture track through an optical flow method and a dynamic semantic mask, extracting a behavior tuple, and calibrating an abnormal action. According to the method, the stability and continuity of communication in a limited space are effectively improved, the recognition capability of abnormal behaviors in a shielding environment is enhanced, efficient fusion and intelligent processing of multi-source sensing data are realized, the false alarm rate and missing report rate of safety monitoring are remarkably reduced, and the real-time performance and precision of risk event response are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image recognition, and in particular to a limited space communication and safety detection method, system and device. BACKGROUND

[0002] With the popularization of concepts such as smart construction, safe operation and digital twin, the safety problems of the operation environment in limited spaces such as buildings, tunnels and mines have gradually attracted attention. Limited spaces (Confined Space) usually have characteristics such as narrow space, poor ventilation and serious visual obstruction, which leads to many technical challenges in the deployment and operation of traditional video monitoring systems and wireless communication equipment. On the one hand, due to the limitation of the geometric structure of the space, it is difficult for conventional cameras to cover all dead angle areas, and continuous tracking of personnel behavior cannot be achieved. On the other hand, traditional wireless communication methods are easily disturbed by factors such as obstruction, reflection and attenuation in limited spaces, leading to unstable or even interrupted communication between nodes, which seriously affects the efficiency of event response and personnel safety protection.

[0003] Under this background, the "space perception-communication coordination-safety detection" integrated method that combines multi-source perception, edge intelligence and semantic mapping has gradually become a research hotspot. In recent years, the integration capabilities of technologies such as visual SLAM, UWB ranging and millimeter wave radar have been continuously improved, making it possible to construct a three-dimensional semantic map through multi-sensor fusion. At the same time, the behavior recognition method has also gradually expanded from early two-dimensional image analysis to three-dimensional pose recognition based on key point trajectory and depth estimation, greatly improving the accuracy and robustness of human motion recognition. However, the current mainstream research still has high system coupling, fragmented communication perception, low abnormal positioning accuracy and other bottlenecks, and has not yet formed a general solution that is adaptive to complex occlusion scenarios, dynamically maintains communication, and maps component-level risks. SUMMARY

[0004] In view of the above existing problems, the present application is proposed.

[0005] Therefore, the present application provides a limited space communication and safety detection method, system and device, which solves the problems of "fixed deployment, non-robust communication, inability to identify abnormal behavior under occlusion, and inaccurate event positioning" that exist in existing limited space safety detection technologies.

[0006] To solve the above technical problems, the present application provides the following technical solutions: In a first aspect, the present application provides a limited space communication and safety detection method, which comprises, constructing an occlusion perception map and combining ranging and visual SLAM to complete the establishment of an initial three-dimensional semantic map and a device topology map; A communication Mesh topology is dynamically generated based on the line-of-sight occlusion information, and path reconstruction is triggered when occlusion occurs between nodes, and on the basis of communication guarantee, edge nodes are dispatched for image acquisition and transmission; The human body posture trajectory under occlusion is restored through the optical flow method and dynamic semantic mask, behavior tuples are extracted and abnormal actions are calibrated, the behavior information is spatially aligned with the BIM model, a triple is generated and uploaded as a standard risk event structure; The edge node performs local response according to the abnormal type, and realizes video access, alarm pushing and task scheduling functions on the App side.

[0007] As a preferred scheme of the limited space communication and safety detection method, wherein: an occlusion-aware map is constructed, and the establishment of the initial three-dimensional semantic map and the device topology map is completed by combining ranging and visual SLAM; the portable laser range finder and the depth camera are used to collect space boundary points, and an initial plane point set is constructed; the edge detection is performed on the two-dimensional projection map by using OpenCV, and the high-risk area is manually labeled to obtain a risk space map; According to the dead angle occlusion area, the narrow channel area and the vertical wellhead area marked in the risk mask map, a deployment density weight map is generated, a 3*3 grid field centered on the sliding window center point is constructed on the two-dimensional ground structure map by sliding window traversal with a fixed step, the risk weight sum of all cells in the neighborhood is counted, if the risk weight sum exceeds the preset threshold, the sliding window center point is deployed as a candidate deployment point and added to the candidate point set, and a group of fusion perception nodes composed of UWB locators, millimeter wave radars and depth cameras are deployed on each selected candidate point; Each group of three-mode nodes is fixed on the same calibration station, and a coordinate alignment calibration operation is performed by using a double plane plate structure, a depth map of the current frame of each group of nodes is obtained from the depth camera, a corresponding frame echo intensity map is obtained from the millimeter wave radar, and the three-dimensional positioning coordinates of the nodes are obtained from the UWB; Gradient calculation is performed on each frame of the depth map to obtain an edge intensity map and extract the area with obvious edge changes on site, and a strong threshold judgment is performed on the millimeter wave radar image to identify the area with occlusion and detection blind area, a local occlusion-aware map is generated by pixel-level superposition of the depth edge map and the millimeter wave blind area map, and all node-generated local occlusion-aware maps are mapped to the main coordinate system through coordinate transformation and voxel merging to obtain a global occlusion-aware map; On the basis of obtaining the global occlusion map, a three-dimensional semantic map is further constructed, and on the basis of the constructed semantic map, device class semantic points are extracted and a device topology map is constructed.

[0008] As a preferred scheme of the limited space communication and safety detection method, wherein: the communication Mesh topology is dynamically generated based on the line-of-sight occlusion information, and path reconstruction is triggered when occlusion occurs between nodes; the three-dimensional coordinates of each pair of connected nodes in the device topology graph are obtained from the three-dimensional semantic map and and a spatial path sampling point sequence from to is generated by using a straight line interpolation method; The occlusion state of each sampling point in the path is retrieved from the global occlusion awareness map. If the sampling point is in the occlusion region in the global occlusion awareness map, the corresponding reflection intensity value recorded by the millimeter wave radar is called and added to the occlusion intensity set O. The spatial distance between nodes i and j is measured in real time by calling the UWB sensor array to obtain the ranging value. The link equivalent attenuation value of the node pair is calculated based on the ranging value and the occlusion intensity set. A set V composed of all nodes in the communication Mesh network is defined m For any two nodes in the node set, if the link equivalent attenuation value between the nodes does not exceed the preset maximum communication attenuation threshold I, it is considered that the communication link can maintain communication in the current environment, and is retained in the communication topology structure, otherwise it is excluded. All node pairs that meet the condition are collected as the edge set H m A dynamic communication Mesh topology graph is constructed based on V m and H m . After the communication topology graph is updated, it is immediately determined whether the connection path between the target task node pair is interrupted or the forwarding node fails. If any node communication link fails or the shortest path changes, the dynamic path reconstruction mechanism is triggered, and AODV protocol is used as the optimal path reconstruction strategy.

[0009] As a preferred scheme of the limited space communication and safety detection method, wherein: the human body posture trajectory under occlusion is restored by the optical flow method and dynamic semantic mask, and the behavior tuple is extracted and the abnormal action pointer is calibrated; each received image frame is preprocessed, and the image data after standardization is input into the YOLOv8 model for target detection. The YOLOv8 model outputs K candidate target boxes in each frame, each target includes upper left corner coordinates, width and height dimensions, belonging category and confidence score. The detection result is filtered, only the target box with a category of “human body” and a confidence higher than a set threshold is retained to generate a human body candidate box set. The human body candidate box set is traversed, and image cropping operations are performed on the regions corresponding to each candidate box in adjacent image frames to extract local image blocks of each human body in adjacent frames, and a a set of image pairs as input, each pair corresponds to the local region change image of a human target in the front and back frames; The fine dense optical flow estimation is performed on each pair of human image blocks by the double-branch RAFT network, the main branch is responsible for the basic motion calculation, the auxiliary branch introduces the edge attention and context residual optimization, and finally the dense optical flow map of the human region is output. The depth map of the current frame is generated by using the MiDaS depth estimation network, and the dense optical flow map calculated by RAFT is combined to perform forward propagation on the key point position in the last frame in each human candidate box, to obtain the preliminary predicted key point coordinates of the current frame. Through the depth consistency judgment mechanism, the propagated key points are screened. If the depth change of the key point in the front and back frames is lower than the set threshold, it is considered that the key point is not occluded, and is marked as a "trusted key point", otherwise it is marked as an "occluded key point"; The human region image block where all the occluded key points are located is extracted, the initial coordinates of the occluded points are combined with the human region image block and input into the pose completion network. The completion network performs fine regression on the spatial position of each key point based on the context information, and outputs the accurate coordinates of the occluded key points; The trusted key points and the completed key points are merged to form a complete key point trajectory set of each candidate human in the current frame. The key point trajectory of each candidate human is time-aligned with the trajectory in the previous J frames to construct a continuous time window trajectory sequence, and the velocity vector of each key point is calculated. For the three points constituting the biological motion chain, the joint angle of the current frame is calculated; Based on the continuous time change characteristics of velocity and angle, an action feature vector is constructed. The action feature vectors of the last J frames are constructed into a sliding window sequence and input into a lightweight time sequence model for action recognition. The behavior category label of each individual in the current frame is output, including common actions and uncommon actions; A risk template set is defined, and each behavior is compared with the risk template set. If the action category belongs to the risk template set, it is judged as a potential anomaly; At the same time, the velocity mutation value, the angle change rate and the displacement offset are combined and weighted to calculate the risk score of the behavior. If the risk score exceeds the preset risk threshold Z, the behavior is marked as a high-risk behavior event. All detected abnormal behavior events are packaged as structured records.

[0010] As a preferred scheme of the limited space communication and safety detection method, wherein: the behavior information is spatially aligned with the BIM model, a triple is generated, and a standard risk event structure is uploaded according to the detection result of each abnormal event in the image frame, the corresponding behavior center point coordinates are extracted, and the three-dimensional position coordinates of the behavior center point in the camera space are restored in combination with the depth map estimation value obtained in the previous stage, the conversion from the camera coordinate system to the BIM model world coordinate system is completed through the external parameter matrix, and each event position is projected to the actual component area. The spatial boundary information of all components in the BIM model is traversed, a spatial inclusion query is performed on each mapping point, the component number corresponding to the event is automatically determined, the key fields of the abnormal event are extracted and organized into an event triple, and the event triple is uploaded to the project early warning database in real time to form a standard risk event structured record.

[0011] As a preferred scheme of the limited space communication and safety detection method, wherein: the edge node performs local response according to the abnormal type, and realizes video access, alarm pushing and task scheduling functions on the App side, and the local rule matching is performed according to the behavior type field in the event triple, and the control signal is sent to the local alarm controller to start the on-site sound and light alarm device and link the camera to perform area tracking, the mapped component and space position are used as indexes to drive the controller to lock the camera view angle, focus on the corresponding construction area, and encapsulate the event structure into a unified format and push it to the mobile App of the construction site manager, and the App side automatically triggers the video access module, the alarm pushing module and the task scheduling module after receiving the push.

[0012] As a preferred scheme of the limited space communication and safety detection method, wherein: the safety detection data is stored into the database, and the triple information of each time is generated into a structured record together with the key frame image, the space position and the processing result, and is locally cached in the edge node and uploaded to the center platform database through a safety channel.

[0013] As a preferred scheme of the limited space communication and safety detection method, wherein: the image acquisition and transmission of the edge node are scheduled on the basis of communication guarantee, the image acquisition instruction is issued according to the preset frame rate and resolution parameter, the acquisition node calls the depth camera to complete continuous image acquisition, and the image sequence is generated by combining the lightweight compression algorithm coding, and the data transmission is performed after the time stamp and pose meta information are attached.

[0014] In the second aspect, the application provides a limited space communication and safety detection system, comprising, The semantic mapping module is used for completing space boundary recognition, risk area labeling and global occlusion perception map generation, and constructing a three-dimensional map and a topological structure with device semantic information. A communication topology construction module is configured to construct and dynamically update a communication Mesh topology based on the occlusion state between nodes and the equivalent attenuation value of a link, and guarantee the reachability of data transmission between nodes. A risk identification module is configured to estimate the human key point trajectory under occlusion by means of optical flow estimation and semantic mask completion, identify high-risk actions, and output structured behavior events. A space mapping module is configured to locate abnormal behaviors in three dimensions to BIM components and generate a three-tuple structured record, and synchronously to a pre-warning database in real time. A task scheduling module is configured to perform local linkage response, video focus tracking, and mobile App pushing and task scheduling according to the type of abnormal events.

[0015] A limited space communication and safety detection device suitable for the limited space communication and safety detection method and a limited space communication and safety detection system, comprising a main body, a winding wheel rotatably installed in the main body, a storage battery and a communication host fixedly installed in the main body, a receiving bin movably installed on the main body, a terminal switch and a plurality of gas detection and communication terminals movably installed in the receiving bin, a handheld tablet detachably installed on the top of the main body, a communication cable connected with the terminal switch wound on the winding wheel, and a threading hole for the communication cable to pass through formed on the inner wall of one side of the main body.

[0016] The present application has the following advantages: the present application constructs an occlusion perception map, fuses a depth map and a millimeter wave blind area map, dynamically generates a three-dimensional semantic map and a communication topology, realizes risk guidance of perception deployment, adaptive communication maintenance and visual restoration of abnormal behaviors, adopts a posture restoration method based on optical flow estimation and key point completion in the identification of human actions under occlusion, significantly improves the continuity and integrity of abnormal behavior detection, realizes accurate binding of behavior events and space components by means of BIM model coordinate positioning mechanism and event three-tuple structured storage, and provides data support for subsequent edge response and scheduling. The system as a whole embodies the innovative features of communication perception collaborative enhancement, perception-identification-response closed-loop control and cross-modal information fusion. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0018] Figure 1 The flowchart of the limited space communication and safety detection method in embodiment 1.

[0019] Figure 2 Structure diagram of limited space communication and safety detection system in embodiment 1.

[0020] Figure 3 Flowchart of building occlusion-aware map and device deployment in embodiment 1.

[0021] Figure 4 Overall structure diagram of limited space communication and safety detection device in the application.

[0022] Figure 5 Cross-sectional view of limited space communication and safety detection device in the application Figure 1 .

[0023] Figure 6 Cross-sectional view of limited space communication and safety detection device in the application Figure 2 .

[0024] In the figure: 1, main body; 2, winding wheel; 3, battery; 4, communication host; 5, handheld tablet; 6, storage bin; 7, gas detection and communication terminal; 8, terminal switch; 9, threading hole. DETAILED DESCRIPTION

[0025] In order to make the above objectives, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0026] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, but the present application can also be implemented in other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the scope of the present application, therefore the present application is not limited to the specific embodiments disclosed below.

[0027] Secondly, the "one embodiment" or "embodiment" referred to herein can include specific features, structures or characteristics contained in at least one implementation of the present application. In this specification, "in one embodiment" appearing in different places does not refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments.

[0028] Embodiment 1, refer to Figures 1-3 , the first embodiment of the present application provides a limited space communication and safety detection method, comprising the following steps: S1, build an occlusion-aware map, and complete the establishment of an initial three-dimensional semantic map and device topology map in combination with ranging and visual SLAM; Specifically, a shielding perception map is constructed, and combined with ranging and visual SLAM, the establishment of an initial three-dimensional semantic map and device topology map is completed. A portable laser range finder and a depth camera are used to collect spatial boundary points, and an initial plane point set is constructed. A two-dimensional projection map is edge detected by OpenCV and a high-risk area is manually labeled to obtain a risk space map, which specifically includes: A measuring station is set up every 1 meter inside the limited space, and a portable laser range finder (such as Leica Disto S910) is used to record the three-dimensional coordinates of the boundary corner points. At the same time, a depth camera (such as Intel RealSense D455) is used to obtain the depth information of each frame of image. For each frame of image, the pixel coordinates and depth values are converted into three-dimensional coordinates by using the camera intrinsic matrix to form the preliminary point cloud data. Through the registration process based on the iterative closest point (ICP) algorithm, the laser ranging points and the depth camera point cloud are spatially aligned to generate a unified boundary point set S. The registered three-dimensional point cloud is projected onto the XY plane to generate a two-dimensional ground structure map. The spatial structure boundary information is extracted by the Canny edge detection algorithm in OpenCV, and the edge map is loaded into the manual labeling interface. The typical risk areas, including dead angle shielding, narrow channel and vertical wellhead, are identified and framed by the workers; After labeling, the corresponding risk mask map is generated according to the region category. Each pixel point in the mask map corresponds to a risk label value, which is used to represent the risk type it belongs to. The specific labeling rules are as follows: the dead angle shielding area is assigned a value of 1, the narrow channel area is assigned a value of 2, the vertical wellhead area is assigned a value of 3, and the no-risk area is assigned a value of 0. The risk mask map and the boundary point set S are bound to generate a risk space map, which contains three pieces of information: spatial point set, edge map and mask; According to the dead angle shielding area, the narrow channel area and the vertical wellhead area marked in the risk mask map, a layout density weight map is generated, which is specifically defined as follows in this embodiment: ; In the formula, W(x, y) is the risk weight of the pixel point (x, y), is the risk type number of the region to which the pixel point (x, y) belongs , , and are the weighted values corresponding to the risk levels; Sliding window traversal is performed on the two-dimensional ground structure diagram at a fixed step size, a 3*3 grid field centered on the sliding window center point is constructed, the total risk weight of all cells in the neighborhood is counted (i.e. the weight values of all cells in the field are added to form a local total risk score), if the total risk weight exceeds a preset threshold (based on experience), the sliding window center point is deployed as a candidate deployment point and added to the candidate point set, a set of fusion perception nodes composed of UWB locators, millimeter wave radars and depth cameras are deployed on each selected candidate point; In this embodiment, the UWB locator is installed at the center position of the visible structural beam column or the cabin top panel at the highest position of each closed area, and the specific installation height is 2.5~3.2m above the ground, which is used to provide a three-dimensional global coordinate reference system for the area. In order to avoid signal shielding, the center point of the unshielded top structure is selected, and one UWB node is installed every 8~10m in the long corridor area according to the triangular layout principle, forming an overlapping positioning area; The millimeter wave radar device is fixed to the structural column or wall on both sides of the corridor or working platform, and the installation height is 1.2~1.5m above the ground, which is inclined downward by 10°~15° towards the center of the corridor. This position makes its radar beam can penetrate through the low shielding objects, which is suitable for penetration detection and speed estimation, and is suitable for long and narrow areas such as ventilation pipe corridor and manhole corridor; The depth camera is installed on the bracket near the equipment operation platform or above the corridor entrance, and the installation height is 1.6~2.0m above the ground, the lens is directed towards the key operation area and adjusted to a downward angle of 30°~45°. This installation position can maximize the coverage of personnel behavior trajectory and space deformation, and is used for accurate pose modeling and body tracking; All nodes are assigned a unique number, and a node registration table is constructed based on the physical location, orientation and device type of the device. Each group of nodes is configured with a GNSS timing module, which sends a unified pulse signal (PPS) to the UWB locator, millimeter wave radar and depth camera respectively, and transmits UTC time data through NMEA protocol. After receiving, each device records the time offset and corrects the internal timestamp to the GNSS time standard, thereby realizing time synchronization across devices; Each group of three-mode nodes is fixed on the same calibration station, and a double-plane plate structure is used to perform coordinate alignment calibration operation. Specifically, all input points are converted to a unified unit through linear scaling, and the depth camera view is used as the main coordinate system. The two-dimensional pattern of the calibration plate is observed and collected at N angles, combined with the coordinates of the millimeter wave radar and UWB locator, and through the registration relationship between image coordinates and three-dimensional point cloud, the spatial transformation matrix of the millimeter wave radar and UWB locator relative to the depth camera is estimated, and the mapping relationship between the coordinate systems of different modalities is established; ; In the formula, is a 4x4 transformation matrix from the millimeter wave radar coordinate system M to the depth camera coordinate system D, is a 4x4 transformation matrix from the UWB locator coordinate system U to the depth camera coordinate system D, is a 3x3 rotation matrix representing the spatial pose of the millimeter wave radar coordinate system relative to the depth camera, is a 3x3 rotation matrix representing the spatial pose of the UWB locator coordinate system relative to the depth camera, is a 3x1 translation vector representing the spatial displacement of the millimeter wave radar relative to the depth camera, is a 3x1 translation vector representing the spatial displacement of the UWB locator relative to the depth camera; Record all nodes' initial view key frames as reference frames before job start, and compare the current view changes at intervals during the job. If the node position offset is detected to exceed the set threshold (based on experience), for example, three-dimensional translation exceeds 5 centimeters or rotation angle exceeds 3 degrees, the current node re-calibration process is automatically triggered to re-calculate the relative pose transformation matrix, ensuring that all nodes are still in accurate alignment state; For each group of nodes, the current frame depth map is obtained from the depth camera, and the corresponding frame echo intensity map is obtained from the millimeter wave radar, and the current three-dimensional positioning coordinates of the nodes are obtained from the UWB; Perform gradient calculation on each frame of depth map to obtain edge intensity map and extract the area with obvious on-site edge changes: ; In the formula, is the depth gradient amplitude at the pixel point, and are the depth gradients in x / y directions, is the depth image of the current frame, is an indicator function, which is 1 when the condition in the parentheses is true, otherwise it is 0, is the occlusion edge determination threshold, which is automatically set according to the gradient distribution, for example, taking the mean plus one standard deviation, is the occlusion edge mask map, 1 represents the occlusion edge; At the same time, the millimeter wave radar image is subjected to strong threshold determination to identify the area with occlusion and detection blind area: ; In the formula, is the radar blind area mask map, 1 represents the blind area, is the radar blind area determination threshold, which is set by experience, is the echo intensity of the millimeter wave radar image at pixel (x, y); The local occlusion perception map generated by all nodes is mapped to the main coordinate system through coordinate transformation and voxel merging to obtain a global occlusion perception map; On the basis of obtaining the global occlusion map, a three-dimensional semantic map is further constructed, three-dimensional point clouds are recovered from the depth map, pixel-to-three-dimensional space mapping is completed through the camera internal parameter, the image pixel position corresponding to each point is recorded, a semantic segmentation operation is performed on the RGB image, each pixel is assigned a corresponding semantic class label, all spatial points are encoded according to the corresponding pixel label to form a sparse point cloud with semantic labels, the sparse point cloud with semantic labels is registered in the same coordinate system through GNSS time synchronization and pose transformation to form a global three-dimensional sparse map with semantic labels; On the basis of the constructed semantic map, device class semantic points are extracted and a device topology map is constructed, specifically including screening out points belonging to the device class from the semantic point cloud and performing spatial clustering, the center position of each device is extracted as a node in the map, then visibility judgment is performed between any two device nodes, if the line of sight between the two points is reachable in the occlusion perception map and the Euclidean distance is within a limited range, it is considered that the two exist a connection relationship, an adjacency matrix of the device topology is established, and then a device spatial connection graph is constructed, all device nodes and their connection relationships form a complete topology structure, the device topology map is fused with the deployment node position to provide a structure constraint and node forwarding reference map for subsequent communication path optimization.

[0029] The application realizes accurate identification and labeling of dead angle occlusion, narrow channel and blind area in a limited space by fusing multi-modal sensors such as laser ranging, depth camera and millimeter wave radar, constructing risk weight map and occlusion perception map; meanwhile, GNSS time synchronization and double plane board calibration are used to ensure the spatial and temporal alignment of data between devices; further, semantic segmentation is used to generate three-dimensional point clouds with semantic labels and construct a device topology map, realizing structured identification and connection mapping of device nodes, effectively improving the rationality of perception node deployment, the accuracy of environment mapping and the controllability of communication path planning, and providing safe, efficient and structure-traceable perception and communication support for limited space operation scenarios.

[0030] S2, dynamically generating a communication Mesh topology based on the line-of-sight occlusion information, and triggering path reconstruction when occlusion occurs between nodes, scheduling edge nodes for image acquisition and transmission on the basis of communication guarantee; Specifically, a communication Mesh topology is dynamically generated based on the line-of-sight occlusion information, and path reconstruction is triggered when occlusion occurs between nodes, that is, each pair of connected nodes in the device topology map is traversed, the three-dimensional coordinates of node i and node j are obtained from the three-dimensional semantic map and and a straight line interpolation method is used to generate a straight line from to the spatial path sampling point sequence of ; wherein, is the spatial position of the starting node, is the spatial position of the target node, g is the number of segmented sampling points, is the position of the kth sampling point; Retrieving the occlusion state of each sampling point in the path in space from the global occlusion perception map, if the sampling point is in the occlusion region in the global occlusion perception map, calling the corresponding reflection intensity value recorded by the millimeter wave radar and adding it to the occlusion intensity set O, calling the UWB sensor array to measure the spatial distance between node i and node j in real time to obtain the ranging value, and calculating the link equivalent attenuation value of the node pair based on the ranging value and the occlusion intensity set: ; wherein, represents the physical distance equivalent to the path from node i to node j in actual communication, reflecting the degree of communication attenuation, the greater the value, the more unusable the path, is the UWB ranging value, is the reflection signal intensity of the lth occlusion detected by the millimeter wave radar on the communication path from node i to node j, reflecting the thickness and material of the occlusion, is the occlusion equivalent path conversion coefficient, representing the equivalent path length increase caused by every 1 dB occlusion, which is obtained by prior experimental calibration, and n is the number of occlusions detected on the communication path from node i to node j; Defining a set V m composed of all nodes in the communication Mesh network, which contains multiple source sensor nodes deployed in a limited space, including but not limited to communication nodes integrating UWB modules, millimeter wave radars, and visual acquisition devices, for any two nodes in the node set, if the link equivalent attenuation value between the nodes does not exceed the preset maximum communication attenuation threshold I (experimentally optimized), it is considered that the communication link can maintain communication under the current environment and should be retained in the communication topology structure, otherwise it is excluded, and all node pairs that meet the condition are collected as the edge set H m Based on V m and H m , a dynamic communication Mesh topology graph is constructed, which reflects the connection relationship between all nodes with actual communicable ability and serves as the basic structure for subsequent routing path selection and network scheduling; To adapt to the change of visibility caused by frequent movement of personnel or equipment in limited space, a mechanism of periodic detection of communication events and dynamic update of topology map is designed, which specifically includes setting a fixed time interval as the refresh period, recalculating the equivalent attenuation value of the link of all node pairs in each period, and comparing it with the result of the last period. If the link state of a node pair changes, for example, from valid to invalid or from invalid to valid, it is considered as a communication event trigger and the communication topology map is immediately updated, adding or deleting the corresponding edge connection. At the same time, the timestamp of this topology change and the changed link number are recorded and written into the system log for reference by the subsequent task scheduling and data synchronization module. This mechanism ensures that the communication network can reflect the changes in the field environment in real time, improving the reliability of data transmission and the stability of task scheduling. After the update of the communication topology map is completed, it is immediately determined whether there is a break in the connection path between the target task node pair or a failure of the forwarding node. If there is any communication link failure or shortest path change between the nodes, the dynamic path reconstruction mechanism is triggered, and the AODV (Ad hoc On-Demand Distance Vector) protocol is used as the optimal path reconstruction strategy. The specific process is as follows: When a node i needs to send task data to a target node j, if the original path has failed or is unreachable, the system will broadcast RREQ packets from i to all adjacent nodes in the current communication available graph, which contain the source node address, target node address, hop count counter, and the latest timestamp. After receiving the RREQ, the intermediate node determines whether it has a feasible path to j. If not, it continues to relay and forward to its neighbors. If a reachable path is found, the RREP (Route Reply) message is returned to the source node along the original path, and the path is established. According to the cumulative hop count and historical average link delay, multiple available paths are evaluated, and the path that meets the minimum delay criterion is selected as the current task communication link. After determining the new path, the forwarding node route cache is updated, and the updated routing table information is broadcast to all nodes on the path to ensure the stability of the data link. If a node detects a link break (e.g., the equivalent attenuation value of the link exceeds the threshold) during data transmission, it actively sends a RERR (Route Error) message to the source node and re-executes the path reconstruction. The latest communicable Mesh topology is uploaded to the cloud platform database and synchronized to the APP mobile terminal.

[0031] The application realizes high coupling of communication link state and space occlusion awareness, ensures that the Mesh structure is adaptively adjusted according to environmental changes, improves the link attenuation evaluation precision through occlusion quantitative modeling, supports reliable link determination under non-line-of-sight communication, realizes path cascade recovery through the AODV protocol, ensures uninterrupted communication between key task nodes, and enhances the intelligent adaptive ability and operation reliability of the system to the limited space communication environment through the linkage of the whole-process closed loop and the mobile terminal.

[0032] Further, on the basis of communication guarantee, the edge nodes are dispatched for image acquisition and transmission. Image acquisition instructions are issued according to preset frame rate and resolution parameters, the acquisition nodes call depth cameras to complete continuous image acquisition, and image sequences are generated through a lightweight compression algorithm (such as H.265 / HEVC (High Efficiency Video Coding)) coding, with timestamps and pose meta information. It is judged whether to perform real-time image transmission or enter the cache waiting queue according to the available bandwidth and occlusion state of the communication link. All successfully transmitted image frames are uniformly recorded into an image index table as the input basis for subsequent occlusion restoration and human body pose trajectory analysis. This process realizes the spatio-temporal consistency and controllable scheduling of image acquisition and transmission, and provides continuous and high-quality data support for behavior tuple extraction and abnormal action recognition.

[0033] S3, restore the human body pose trajectory under occlusion through the optical flow method and dynamic semantic mask, extract the behavior tuple and label the abnormal action, spatially align the behavior information and the BIM model, generate a triple and upload it as a standard risk event structure; Specifically, the human body pose trajectory under occlusion is restored through the optical flow method and dynamic semantic mask, and the behavior tuple and abnormal action are extracted and labeled. Each received image is preprocessed, specifically including: adjusting the original image to a uniform resolution (such as 640x640) required by YOLOv8, using a bilinear interpolation method to complete image scaling, then performing 0-1 normalization processing on the RGB image to enhance the model's adaptability to images under different lighting conditions, inputting the standardized image data into the YOLOv8 model for target detection, the YOLOv8 model outputs K candidate target boxes in each frame, each target contains the upper left corner coordinates, width and height size, belonging category and confidence score, filtering the detection results, only retaining the target boxes with a category of “human body” and a confidence higher than a set threshold (based on experience) and generating a human body candidate box set to eliminate false positives and background interference. Traverse the human body candidate frame set, perform image cropping operation on the region corresponding to each candidate frame in the adjacent image frames (i.e. the current frame and the previous frame), and extract the local image block of each human body in the adjacent frame. This step can effectively remove the interference of non-target regions in the image, provide clear input for subsequent local optical flow estimation, and improve the optical flow perception effect of the boundary region. In the cropping process, each candidate frame is expanded in the up, down, left and right directions by a fixed number of pixels (such as 16 pixels) as a boundary redundancy area to ensure that the action edge and context clues are included and the key area is not truncated. The expanded coordinates need to be limited within the original image size range to prevent cropping out of bounds. Perform size normalization on all expanded image blocks to adjust the expanded image blocks to the standard input size (such as 256x256 pixels) required by the RAFT model to ensure model input consistency and inference efficiency. Form an image pair input set, each pair corresponding to the local area change image of a human target in the previous and subsequent frames. The output of this step not only provides high-quality input sources for optical flow field calculation, but also significantly reduces the influence of background interference on key point propagation accuracy, ensuring more stable and accurate subsequent pose trajectory estimation. Perform fine-grained dense optical flow estimation on each pair of human image blocks through a double-branch RAFT network. The main branch is responsible for basic motion calculation, and the auxiliary branch introduces edge attention and context residual optimization. Finally, the local human region optical flow field is output, which includes: Initialize the hidden state and correlation volume of the double-branch RAFT network for the GRU recursion process inside RAFT. Input the image block into the RAFT backbone network and extract image features through the Feature Encoder Initialize the zero optical flow map Call the correlation matching mechanism inside the RAFT backbone to perform GRU recursion to generate a basic optical flow estimation map: ; In the formula, is the two-dimensional optical flow vector of the th candidate human frame at time t corresponding to the pixel point (x, y) (the final output result), is the displacement (unit: pixel) between the front and rear frames in the x direction (horizontal) at the pixel point (x, y), is the displacement (unit: pixel) between the front and rear frames in the y direction (vertical) at the pixel point (x, y); Input the current frame image into the Sobel convolution operator to extract the horizontal and vertical gradient maps, and generate an edge attention map through the channel attention mechanism (such as channel weighting in CBAM): ; In the formula,​ is an edge attention map at the pixel point (x, y), CBAM() is a convolutional attention module for fusing channel attention and spatial attention, is a gradient map of the image in the x direction, is a gradient map of the image in the y direction; The edge attention map is used to identify an occlusion or edge change region, to provide guidance for optical flow residual compensation, to input an image block into a lightweight FPN structure, and to obtain three-layer pyramid feature maps, each of which captures context information under different receptive fields; The edge attention map and the context feature map set are input into a lightweight residual network (such as 3-layer ResBlock) to generate a guided residual optical flow map: ; In the formula, is an auxiliary branch guided residual correction optical flow, is a residual displacement amount at the pixel point estimated by the context feature map and the residual network, which is the same dimension as the main branch output, is an edge attention map used to enhance the correction strength of the occlusion and edge region; The main output and the residual correction are summed pixel by pixel to generate a dense optical flow map of the current human body region: ; In the formula, is an original optical flow map estimated by the main branch, is a dense optical flow map; After completing the dense optical flow field estimation, the key point position extrapolation and occlusion judgment are realized relying on the depth map constraint, and then the occluded key points are completed and fused, and the complete key point trajectory set of the current frame is output, which specifically includes: using the MiDaS depth estimation network to generate the depth map of the current frame, and combining the dense optical flow map calculated by the RAFT, the key point position in the previous frame in each human body candidate box is forward propagated to obtain the preliminary predicted key point coordinates of the current frame, and the key points after propagation are filtered through a depth consistency judgment mechanism, if the depth change of the key point in the front and back frames is lower than a set threshold (experimentally optimized), it is considered that the key point is not occluded, and is marked as “trusted key point”, otherwise it is marked as “occluded key point”; The image block of the human body region where all the occluded key points are located is extracted, the initial coordinates of the occluded points and the human body region image block are combined and input into a pose completion network (such as PoseResNet or HRNet), and the completion network performs fine regression on the spatial position of each key point based on context information to output the accurate coordinates of the occluded key points; The trusted key points and the completed key points are merged to form a complete key point trajectory set of each candidate human body in the current frame. Each candidate human keypoint trajectory is time-aligned with the trajectories within the previous J frames, a continuous time window trajectory sequence is constructed, and the velocity vector of each keypoint is calculated: ; wherein, is the velocity vector of the jth keypoint of the ith person at time t, and is the horizontal and vertical coordinate position of the keypoint in the current frame (time t) image, and is the horizontal and vertical coordinate position of the keypoint in the previous frame (time t−1) image, is the time interval between the current frame and the previous frame; For the three points constituting the biological motion chain, the joint included angle of the current frame is calculated: ; wherein, is the included angle value of the qth key angle of the oth person at time t, is the vector from the center point (such as the elbow) to the first point (such as the shoulder), is the vector from the center point to the second point (such as the wrist); Based on the continuous time change characteristics of the velocity and the angle, a behavior feature vector is constructed, the behavior feature vectors of the last J frames are constructed into a sliding window sequence and input into a lightweight time sequence model for action recognition, including using MobileNetV2 to extract a low-dimensional embedding vector, and then using an LSTM structure to capture the trend and pattern of the behavior evolution over time, so as to output the behavior category label of each individual in the current frame, including common actions and uncommon actions; The common actions include but are not limited to standing and observing, normal walking, squatting and checking, and ladder climbing, etc. The uncommon actions include but are not limited to violent falling / sideways falling, stumbling / slipping, and sudden running, etc. A risk template set is defined, and each behavior is compared with the risk template set. If the action category belongs to the risk template set, it is judged as a potential anomaly; Meanwhile, the velocity mutation value, the angle change rate, and the displacement offset are weighted and fused to calculate the risk score of the behavior. If the risk score exceeds a preset risk threshold Z (based on experience), the behavior is marked as a high-risk behavior event; All detected abnormal behavior events are encapsulated as structured records, which include individual identification, behavior type, position coordinates, risk score, and risk level.

[0034] CBAM (Convolutional Block Attention Module) is a channel and spatial attention mechanism that enhances the model's focus on target areas and improves edge structure detection capabilities.

[0035] RAFT (Recurrent All-Pairs Field Transforms) is a dense optical flow estimation method based on GRU iteration and all-to-all correlation, with high precision and strong robustness.

[0036] MiDaS network is a neural network model for depth estimation, which generates a depth map for each frame of image.

[0037] PoseResNet / HRNet is a human key point pose estimation network that restores or completes human structure in images.

[0038] Through image scaling and normalization operations, the YOLOv8 model is ensured to run stably under various lighting and image resolution conditions, and background interference is preliminarily eliminated. The candidate human region is extracted as a cropped image block, and the double-branch RAFT network is used to estimate the basic optical flow and residual compensation optical flow. The main trend of motion is provided by the backbone network, and the edge attention map guides the occlusion repair, so that the key action details in the occluded area can still be restored. The depth map generated by MiDaS is used to judge the occlusion relationship, and the occluded points are accurately regressed and completed by PoseResNet, ensuring that stable and complete human motion trajectories can still be obtained under occlusion conditions. The key point trajectory is constructed as a velocity vector and angle feature, combined with the LSTM network to model the time evolution process of the action, and the risk score is calculated to comprehensively identify unusual actions, realizing early perception of high-risk events. The behavior recognition result is packaged as a structured abnormal event tuple, including individual ID, behavior category, spatial position, and risk score, providing standard input for subsequent BIM mapping and system response.

[0039] This method can still recover complete pose information under occlusion conditions and accurately determine abnormal behavior. Its advantages are: improving the perception accuracy of key boundary areas through attention mechanism; improving the reliability of occluded point completion using depth information; integrating time series modeling and spatial behavior features to build a more accurate action recognition mechanism. Finally, the proposed behavior structured tuple provides a basic semantic structure for subsequent spatial linkage response, safety warning, and engineering system control.

[0040] Further, the behavior information is spatially aligned with the BIM model to generate triples and upload as a standard risk event structure. According to the detection results of each abnormal event in the image frame, the corresponding behavior center point coordinates are extracted, and the three-dimensional position coordinates of the behavior center point in the camera space are restored in combination with the depth map estimate value obtained in the previous stage. The camera coordinate system is converted to the BIM model world coordinate system through the external parameter matrix, and each event position is projected to the actual component area; The spatial boundary information of all components in the BIM model is traversed, and a spatial inclusion query is performed on each mapping point to automatically determine the component number corresponding to the event. If a behavior event position falls into multiple component boundary regions, the component with the smallest enclosing volume is selected to ensure spatial matching accuracy. The key fields of the abnormal event are extracted, including the unique identifier of the personnel, the behavior type (such as illegal entry, high-altitude work without wearing a safety rope, etc.), the component ID (such as "support F5"), the semantic information of the area (such as "tower crane operation area"), and the event timestamp. The extracted key fields are organized into event triples in the form of: (personnel ID, behavior type, component ID, timestamp). To ensure efficient integration with the BIM system, the triples are semantically bound to the BIM components, and the corresponding components are highlighted in the digital twin visualization platform to assist management personnel in instantly grasping the risk distribution and construction status. The triples are also uploaded to the project early warning database in real time to form a standard risk event structured record.

[0041] Through the cooperation of the depth map and the external parameter matrix, two-dimensional image events are accurately projected into BIM three-dimensional space to realize "pixel-component" level spatial mapping, greatly improving the risk event positioning accuracy. The triple structure not only includes behavior recognition results, but also integrates component ID and area semantics, making the event semantically correspond one-to-one with the BIM system in structure, improving the context integrity of the event. By standardizing the triple record and uploading it to the project early warning database, a unified data interface can be formed to support multi-module collaborative response of the construction management system, such as alarm pushing, task scheduling, and progress evaluation. After the triples are bound to the BIM, abnormal components can be displayed in real time with highlights on the digital twin platform, and combined with the timestamp, the construction phase playback and risk tracing are supported, enhancing the project safety closed-loop management capability. Through the minimum enclosing volume optimization strategy, the matching ambiguity of the overlapping area of multiple components is avoided, ensuring the stability and consistency of the event allocation with decision logic, and reducing false positives and false negatives.

[0042] S4, the edge node performs local response according to the abnormal type, and realizes video access, alarm pushing and task scheduling functions on the App side; Specifically, the edge node performs a local response according to the abnormal type, and implements video access, alarm pushing and task scheduling functions on the App end. The edge node executes local rule matching according to the behavior type field in the event triple. If high-risk events such as “crossing the boundary” and “staying in the restricted area for too long” are identified, the edge node immediately sends a control signal to the local alarm controller, starts the on-site sound and light alarm device, and links the camera to perform area tracking. The mapped components and spatial positions are used as indexes to drive the controller to lock the camera view angle and focus on the corresponding construction area. The event structure is encapsulated into a unified format and pushed to the site management personnel mobile App. After receiving the push, the App end automatically triggers the video access module, the alarm pushing module and the task scheduling module. The video access module automatically retrieves the real-time video stream or historical video clip of the corresponding time period and corresponding area to show the process of the abnormality and assist in manual review. The alarm pushing module generates a graphic alarm notification containing the risk level, behavior category, personnel information, component positioning and timestamp, and pushes it to the corresponding responsible person, supervising engineer and project manager. The task scheduling module converts the event into a to-be-processed task based on the event level and the work area to which the component in the BIM belongs, automatically enters the engineering task flow platform, marks the responsible person and processing time limit, forms a closed-loop work order, and stores the safety detection data into the database.

[0043] The attention mechanism is used to improve the perception accuracy of key boundary areas. The depth information is used to improve the reliability of occlusion point completion. The temporal modeling and spatial behavior characteristics are fused to construct a more accurate action recognition mechanism. Finally, the proposed behavior structured tuple provides a basic semantic structure for subsequent spatial linkage response, safety warning and engineering system joint control, and has wide engineering practicality and deployability.

[0044] Further, storing safety detection data into a database means that the triple information of each time, together with the key frame image, spatial position and processing result, is uniformly generated into a structured record, which is locally cached on the edge node and uploaded to the center platform database through a secure channel. Finally, the event and response are structured into the database and archived regularly.

[0045] The embodiment also provides a limited space communication and safety detection system, comprising: A semantic mapping module is used to complete space boundary recognition, risk area annotation and global occlusion perception map generation, and construct a three-dimensional map and topological structure with device semantic information. A communication topology construction module is used to construct and dynamically update the communication Mesh topology based on the occlusion state and equivalent attenuation value of the links between nodes, to ensure the data transmission reachability between nodes. The risk identification module is configured to estimate human key point trajectories under occlusion by means of optical flow estimation and semantic mask completion, identify high-risk actions, and output structured behavior events. The spatial mapping module is configured to locate abnormal behaviors in three dimensions to BIM components and generate a triple structure record, and synchronize to the early warning database in real time. The task scheduling module is configured to perform local linkage response, video focus tracking, mobile app push and task scheduling according to the type of abnormal events.

[0046] A limited space communication and safety detection device suitable for the limited space communication and safety detection method and a limited space communication and safety detection system, comprising a body main body 1, a winding wheel 2 is rotatably installed in the body main body 1, a storage battery 3 and a communication host 4 are fixedly installed in the body main body 1, a storage bin 6 is transversely movably installed on the body main body 1, the storage bin 6 is a drawer type, a terminal switch 8 and a plurality of gas detection and communication terminals 7 are movably installed in the storage bin 6, a handheld tablet 5 is detachably installed on the top of the body main body 1, a communication cable connected with the terminal switch 8 is wound on the winding wheel 2, and a threading hole 9 is formed in the inner wall of one side of the body main body 1.

[0047] Further, in use, the worker wears the gas detection and communication terminal 7 on the body, then enters the limited space for work, carries the terminal switch 8 into the limited space, and ensures that the gas detection and communication terminal 7 on the body and the terminal switch 8 are in a suitable communication distance, at this time, the communication host 4 and the terminal switch 8 are connected and transmit signals through the communication cable, and the gas detection and communication terminal 7 and the terminal switch 8 and the communication host 4 and the handheld tablet 5 are connected through wireless signals. In work, the gas detection and communication terminal 7 can monitor the air condition of the environment where the worker is in real time, and through the camera provided on the gas detection and communication terminal 7, the commander on the side of the body main body 1 can monitor the working environment of the worker in the limited space in real time, and the commander on the side of the body main body 1 can view the working condition of the worker in real time through the handheld tablet 5 and keep real-time communication with the worker.

[0048] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit it, although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the present application, which should be covered in the scope of the claims of the present application.

Claims

1. A confined space communication and security detection method, characterized by: include, Construct an occlusion perception map and combine ranging with visual SLAM to complete the establishment of the initial 3D semantic map and device topology map; Dynamically generate a communication Mesh topology based on line-of-sight occlusion information, trigger path reconstruction when occlusion occurs between nodes, and schedule edge nodes for image acquisition and transmission based on communication guarantees; The optical flow method and dynamic semantic mask are used to restore the human posture trajectory under occlusion, extract the behavior tuple and mark the abnormal action, spatially align the behavior information with the BIM model, generate triples and upload them as standard risk event structures; The edge node performs a local response based on the exception type and implements video access, alarm push, and task scheduling functions on the App side.

2. The confined space communication and security detection method according to claim 1, wherein: The construction of the occlusion perception map and the combination of ranging and visual SLAM to complete the establishment of the initial three-dimensional semantic map and device topology map refers to using a portable laser rangefinder and a depth camera to collect spatial boundary points, construct an initial plane point set, perform edge detection on the two-dimensional projection map using OpenCV, and manually mark high-risk areas to obtain a risk space map; A deployment density weight map is generated based on the blind spots, narrow passages, and vertical wellheads marked in the risk mask. A sliding window is traversed with a fixed step size on the two-dimensional ground structure map to construct a 3x3 grid area centered on the sliding window center. The sum of the risk weights of all cells in the neighborhood is calculated. If the sum of the risk weights exceeds a preset threshold, the sliding window center is deployed as a candidate deployment point and added to the candidate point set. A set of fusion perception nodes consisting of UWB locators, millimeter-wave radars, and depth cameras are deployed at each selected candidate point. Each group of three-mode nodes is fixed on the same calibration station, and a dual-plane plate structure is used to perform coordinate alignment calibration operations. The depth map of the current frame of each group of nodes is obtained from the depth camera, the corresponding frame echo intensity map is obtained from the millimeter wave radar, and the current three-dimensional positioning coordinates of the node are obtained from the UWB. Gradient calculation is performed on each frame of the depth map to obtain an edge intensity map and extract areas with obvious edge changes on the scene. At the same time, a strong threshold judgment is performed on the millimeter-wave radar image to identify areas with occlusion and detection blind spots. The depth edge map and the millimeter-wave blind spot map are superimposed at the pixel level to generate a local occlusion perception map. The local occlusion perception maps generated by all nodes are uniformly mapped to the main coordinate system through coordinate transformation and voxel merging is performed to obtain a global occlusion perception map. Based on the global occlusion map, a three-dimensional semantic map is further constructed. Based on the constructed semantic map, device-type semantic points are extracted and a device topology map is constructed.

3. The confined space communication and security detection method according to claim 2, wherein: The communication Mesh topology is dynamically generated based on the line-of-sight occlusion information, and path reconstruction is triggered when occlusion occurs between nodes. This means traversing each pair of connected nodes in the device topology map and obtaining the three-dimensional coordinates of node i and node j from the three-dimensional semantic map. and And use the linear interpolation method to generate arrive The spatial path sampling point sequence; The occlusion status of each sampling point in the path is retrieved from the global occlusion perception map. If the sampling point is in the occlusion area in the global occlusion perception map, the millimeter wave radar is called to record the corresponding reflection intensity value and add it to the occlusion intensity set O. The UWB sensor array is called to measure the spatial distance between node i and node j in real time to obtain the ranging value. The link equivalent attenuation value of the node pair is calculated based on the ranging value and the occlusion intensity set. Define the set V consisting of all nodes in the communication Mesh network m For any two nodes in the node set, if the link equivalent attenuation value between the nodes does not exceed the preset maximum communication attenuation threshold I, then the communication link is considered to be able to maintain communication under the current environment and is retained in the communication topology structure. Otherwise, it is removed. All node pairs that meet this condition are collected as the edge set H. m Based on V m and H m Build a dynamic communication Mesh topology diagram; After the communication topology map is updated, it is immediately determined whether there is any interruption in the connection path between the target task node pair or any forwarding node failure. If there is any communication link failure between any nodes or the shortest path changes, the dynamic path reconstruction mechanism is triggered and the AODV protocol is used as the optimal path reconstruction strategy.

4. The confined space communication and security detection method according to claim 3, wherein: The method of restoring the human posture trajectory under occlusion by optical flow method and dynamic semantic mask, extracting behavior tuples and calibrating abnormal actions refers to preprocessing each received frame of image, inputting the standardized image data into the YOLOv8 model for target detection, and outputting K candidate target frames in each frame. Each target includes the upper left corner coordinates, width and height dimensions, category and confidence score. The detection results are filtered to retain only the target frames with the category of "human body" and the confidence score higher than the set threshold, and generate a set of human candidate frames; Traverse the set of candidate frames of the human body, perform image cropping operations on the areas corresponding to each candidate frame in the adjacent image frames, extract the local image blocks of each human body in the adjacent frames, and construct An input set of image pairs, each pair corresponds to a local area change image of a human target in the previous and next frames; A dual-branch RAFT network is used to perform refined dense optical flow estimation on each pair of human image blocks. The main branch is responsible for basic motion calculation, while the auxiliary branch introduces edge attention and context residual optimization, ultimately outputting a dense optical flow map of the human body area. The MiDaS depth estimation network is used to generate a depth map of the current frame. Combined with the dense optical flow map calculated by RAFT, the key point positions of the previous frame in each human candidate frame are forward propagated to obtain preliminary predicted key point coordinates for the current frame. The propagated key points are screened through a depth consistency judgment mechanism. If the depth change of a key point in the previous and subsequent frames is lower than a set threshold, it is considered to be unoccluded and marked as a "trusted key point". Otherwise, it is marked as an "occluded key point". Extract the image blocks of the human body region where all occluded key points are located, combine the initial coordinates of the occluded points with the image blocks of the human body region and input them into the pose completion network. The completion network performs a refined regression on the spatial position of each key point based on the context information and outputs the precise coordinates of the occluded key points. Merge the credible key points with the completed key points to form a complete key point trajectory set for each candidate human in the current frame. Time-align the key point trajectory of each candidate human with the trajectory in the previous J frames to construct a continuous time window trajectory sequence and calculate the velocity vector of each key point. For the three points that constitute the biological motion chain, calculate the joint angle of the current frame; Based on the continuous time variation characteristics of speed and angle, a behavior feature vector is constructed. The behavior feature vectors of nearly J frames are combined into a sliding window sequence and input into a lightweight time series model for action recognition. The behavior category label of each individual in the current frame is output, including common and uncommon actions. Define a risk template set and compare each behavior with the risk template set. If the action category belongs to the risk template set, it is judged as a potential anomaly; At the same time, the speed mutation value, angle change rate, and displacement offset are weighted and fused to calculate the risk score of the behavior. If the risk score exceeds the preset risk threshold Z, the behavior is marked as a high-risk behavior event, and all detected abnormal behavior events are encapsulated as structured records.

5. The confined space communication and security detection method according to claim 4, wherein: The spatial alignment of the behavior information with the BIM model, generating a triplet and uploading it as a standard risk event structure means extracting the corresponding behavior center point coordinates based on the detection results of each abnormal event in the image frame, and restoring the three-dimensional position coordinates of the behavior center point in the camera space in combination with the depth map estimation value obtained in the previous stage, completing the conversion from the camera coordinate system to the BIM model world coordinate system through the external parameter matrix, and projecting each event position to the actual component area; Traverse the spatial boundary information of all components in the BIM model, perform spatial inclusion query on each mapping point, automatically determine the component number corresponding to the event, extract the key fields of the abnormal event, organize them into event triples, and upload them to the project early warning database in real time to form a standard structured record of risk events.

6. The confined space communication and security detection method according to claim 5, wherein: The edge node performs a local response based on the type of anomaly and implements video access, alarm push and task scheduling functions on the App side, which means performing local rule matching based on the behavior type field in the event triplet, and sending a control signal to the local alarm controller, starting the on-site sound and light alarm equipment and linking the camera to perform area tracking, using the mapped components and spatial positions as indexes, driving the controller to lock the camera's perspective, focusing on the corresponding construction area, encapsulating the event structure into a unified format and pushing it to the construction site manager's mobile App. After receiving the push, the App automatically triggers the video access module, alarm push module and task scheduling module.

7. The confined space communication and security detection method according to claim 6, wherein: Storing the security detection data in the database means uniformly generating structured records of the triplet information of each time, together with the key frame image, spatial position and processing results, caching them locally at the edge node, and uploading them to the central platform database through a secure channel.

8. The confined space communication and security detection method according to claim 7, wherein: The scheduling of edge nodes for image acquisition and transmission based on communication guarantee refers to issuing image acquisition instructions according to preset frame rate and resolution parameters, the acquisition node calling the depth camera to complete continuous image acquisition, and combining the lightweight compression algorithm to encode and generate an image sequence, and then transmitting the data with timestamp and posture metadata.

9. A confined space communication and safety detection system, based on the confined space communication and safety detection method according to any one of claims 1 to 8, characterized in that: include, The semantic mapping module is used to complete spatial boundary identification, risk area annotation, and global occlusion perception map generation, and to build a three-dimensional map and topological structure with device semantic information; The communication topology construction module is used to build and dynamically update the communication Mesh topology based on the inter-node occlusion status and link equivalent attenuation value to ensure the reachability of data transmission between nodes; The risk identification module uses optical flow estimation and semantic masking to complete the trajectory of key points of the human body under occlusion, identify high-risk actions, and output structured behavioral events; The spatial mapping module is used to locate abnormal behaviors in three dimensions to BIM components and generate triple structured records, which are synchronized to the early warning database in real time; The task scheduling module is used to perform local linkage response, video focus tracking, mobile app push and task scheduling based on the type of abnormal event.

10. A confined space communication and safety detection device, applicable to the confined space communication and safety detection method according to any one of claims 1 to 8 and the confined space communication and safety detection system according to claim 9, characterized in that: The invention comprises a main body (1), a winding wheel (2) is rotatably installed in the main body (1), a storage battery (3) and a communication host (4) are fixedly installed in the main body (1), a storage compartment (6) is movably installed through the main body (1), a terminal switch (8) and a plurality of gas detection and communication terminals (7) are movably installed in the storage compartment (6), a handheld tablet (5) is detachably installed on the top of the main body (1), a communication cable connected to the terminal switch (8) is wound on the winding wheel (2), and a threading hole (9) for the communication cable to pass through is opened on the inner wall of one side of the main body (1).

Citation Information

Cited By

  • Component installation real-time positioning system and method based on laser reference

    CN121330061A

  • A substation low-altitude target monitoring method based on 5G-A sensing integration technology

    CN122365024A