Deep learning-based wounded rescue life detection and positioning method
Through the deep learning method of integrating multimodal sensor data, the data decoupling and navigation blinding of wounded detection and positioning in complex environments is solved, efficient positioning and rescue path optimization of wounded personnel is achieved, and rescue efficiency and safety are improved.
Patent Information
- Application Number
- CN202510504540.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-05
AI Technical Summary
The existing methods of wounded detection and positioning have problems in complex environments with data decoupling, semantic non-coordination, and navigation path blinding, making it difficult to achieve multimodal fusion, precise positioning and timely life judgment, which affects rescue efficiency and safety.
Using a deep learning-based method, multi-objective RGB-D images, lidar point clouds, millimeter wave vital sign radar data and infrared thermal image data are integrated, and the metric-semantic dual map is constructed by improving the YOLO network, and obstacle avoidance optimization navigation paths are generated.
It realizes high confidence positioning and hierarchical treatment of wounded people in complex environments, improves rescue efficiency and safety, and maintains stable detection performance under complex visual conditions such as smoke and occlusion, and generates an optimized navigation path that avoids risk areas.
Smart Images

Figure CN120428244A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of rescue technology, and in particular to a life-saving detection and positioning method for wounded persons based on deep learning. Background Art
[0002] With the gradual intelligence of urban disaster response and emergency rescue systems, how to quickly discover, accurately locate and assess the vital status of trapped injured people in complex environments has become one of the key research directions of intelligent rescue technology. At the scenes of fires, earthquakes and mining disasters, due to the complex environment, low visibility, dense crowds and severe obstruction of obstacles, traditional vision-based detection technology and single vital sign perception methods face great challenges in actual combat and cannot meet the high-demand rescue scenarios of "multimodal fusion, accurate positioning and timely life identification".
[0003] Currently, mainstream methods for detecting and locating casualties mostly rely on a single sensor. For example, target recognition and location information extraction are achieved based on RGB images or infrared thermal imaging. However, these sensing methods have significant blind spots in areas with severe occlusion, smoke, or structural collapse, which can easily lead to missed detections and misjudgments. In addition, vital sign detection usually uses independent millimeter-wave radars or infrared instruments, which lack spatial linkage with semantic recognition modules, making it difficult to build an effective life risk priority model. In the absence of a unified coordinate mapping mechanism, vital sign data and spatial posture information cannot be integrated, which seriously restricts the intelligent process of casualty classification and rescue path optimization.
[0004] On the other hand, existing target detection algorithms often use classic YOLO and Faster R-CNN structures in rescue scenarios, and perform well in indoor or routine image tasks. However, in multi-source heterogeneous data environments, especially when facing visual degradation and high-dynamic scenes, their recognition accuracy and robustness decrease significantly. At the same time, due to the lack of a deep coordination mechanism for millimeter wave and infrared non-visible light sensor data, existing systems find it difficult to obtain the three-dimensional position information and vital sign fusion score of each injured person in space in real time, resulting in unreasonable scheduling paths and unclear response priorities, which in turn affects the overall rescue efficiency.
[0005] In summary, there are multiple problems among current multimodal fusion detection, real-time spatial positioning and life status recognition, such as data decoupling, semantic non-cooperation, and navigation path blindness. Especially in restricted environments, traditional methods have obvious deficiencies in accuracy, response speed and practicality, and new integrated and multimodal collaborative methods are urgently needed to improve them. Summary of the Invention
[0006] One purpose of the present invention is to propose a life-saving detection and positioning method for wounded people based on deep learning, which effectively improves rescue efficiency and safety.
[0007] A method for detecting and locating a wounded person for life-saving according to an embodiment of the present invention based on deep learning includes the following steps:
[0008] S1. Collect and time-synchronize multi-view RGB-D images, LiDAR point clouds, millimeter-wave vital signs radar data, and infrared thermal imaging data;
[0009] S2. Input the RGB-D image into the improved YOLO network, which performs early semantic fusion in the middle layer of the pyramid and outputs pixel-level semantic masks.
[0010] S3. Based on the camera and lidar extrinsic calibration results, the depth channel corresponding to the pixel-level semantic mask is projected into a 3D coordinate system. The original lidar point cloud is assigned a semantic label for the casualty to generate a semantic point cloud of the casualty.
[0011] S4. Construct a metric map for centimeter-level pose resolution and use the semantic point cloud of the casualty to construct an environmental semantic map. A sparse keyframe mechanism is used to keep the metric map and the environmental semantic map updated synchronously, forming a metric-semantic dual-map dataset.
[0012] S5. Use pixel-level semantic masks to spatially align millimeter-wave vital sign radar data and infrared thermal imaging data, calculate respiratory amplitude and body surface temperature gradient, generate vital sign scores, and write them into the metric-semantic dual map dataset;
[0013] S6. When the distance between the mobile platform and the injured person's dynamic anchor point is less than a preset threshold, re-detection is performed to obtain an updated fusion score of the injured person's location and vital signs, and the attributes of the injured person's dynamic anchor point in the metric-semantic dual map dataset are updated;
[0014] S7. Based on the updated metric-semantic dual map dataset, the 3D coordinates of each injured person’s dynamic anchor point and rescue priority ranking are calculated to generate an obstacle avoidance optimized navigation path that avoids risky semantic areas.
[0015] Optionally, the S2 includes the following steps:
[0016] S21. RGB image I RGB and depth image I D Input is an improved YOLO network built on the CSPDarkNet backbone structure. The CSPDarkNet backbone structure includes multiple CrossStagePartial residual blocks. The main path and cross-channel path in the CrossStagePartial residual block construct a multi-scale channel response structure in parallel, and output the Stage-3 backbone feature map F base ;
[0017] S22. Insert the spatial-channel coupled attention module between Stage-3 and Stage-4 of the CSPDarkNet backbone structure. The spatial-channel coupled attention module consists of the channel attention branch CA(·) and the spatial attention branch SA(·), which acts on the backbone feature map in parallel. Forming the attention-enhanced feature map:
[0018] F attn =CA(F base )+SA(F base );
[0019] CA(·) performs channel enhancement on the local texture of the human body area, SA(·) is used to highlight the spatial edge features in the irregular occlusion area, H represents the spatial height dimension of the feature map output by Stage-3 in the casualty detection image, W represents the spatial width dimension of the feature map, and C represents the number of semantic categories, including the casualty area, occlusion area, and background area.
[0020] S23. Based on depth image I D Build environment visibility weight map Where d(x,y) represents the depth of field value of the pixel point, d max is the maximum effective perception distance, and λ is the visual degradation control coefficient; it is used to simulate the degree of interference of smoke, occlusion, and reflection on detection ability;
[0021] S24. Enhance the attention feature map F attn and the environmental visibility weight map W env (x,y) performs pixel-by-pixel multiplication fusion to generate the environment adaptive feature map F env (x, y, c), used to enhance detection robustness under complex visual conditions;
[0022] S25. Input the environment adaptive feature map into the middle layer F of the feature pyramid m , and introduces a semantic decoding branch to generate the current frame semantic guidance mask on the environment adaptive feature map
[0023] S26. Extract the semantic guidance mask generated during the previous frame detection from the cache With the current frame semantic guidance mask Perform inter-frame fusion to construct semantically consistent enhancement masks:
[0024]
[0025] Where α is the dynamic fusion coefficient, which is dynamically adjusted according to the degree of image change or confidence.
[0026] Optionally, S3 includes the following steps:
[0027] S31. Enhance the semantic consistency mask And the depth image I of the corresponding frame D Perform pixel-by-pixel registration to map the depth of field value d(x,y) corresponding to the pixel point (x,y) in the mask to the three-dimensional space point in the camera coordinate system
[0028]
[0029] Where K is the camera intrinsic parameter matrix, which contains the focal length and principal point information; d(x,y) is the depth of field value of the corresponding pixel in the depth map, which has the same distance unit as the actual physical world;
[0030] S32. Based on the external parameter calibration matrix T between the camera and the lidar cam→lidar The three-dimensional space point Transform to the lidar coordinate system to get the projection point The extrinsic calibration matrix is a 4x4 homogeneous transformation matrix that contains rotation and translation information:
[0031]
[0032] S33. Enhance the mask by all the pixels that meet the semantic consistency Projected point cloud collection Aggregate into the semantic projection point cloud of the current frame And assign a label to each point in the projected point cloud set = injured area, forming a label point set:
[0033]
[0034] in, Assign a semantic label to the point to represent the semantics of the injured area in the pixel mask corresponding to the point;
[0035] S34. Raw LiDAR point cloud data Point Cloud with Semantic Projection Perform spatial union in the lidar coordinate system, and pass the radius threshold δ match Perform nearest neighbor matching on the original lidar point cloud data The points that meet the following conditions are assigned semantic labels
[0036]
[0037] in, For the matched and labeled point cloud set, δ match is the point cloud fusion tolerance threshold, which is used to control the spatial propagation range of the projection label, pi and p j Represent the 3D points in the semantic point cloud and the original lidar point cloud respectively;
[0038] S35. The point cloud collection that is matched and annotated As the semantic point cloud of the casualty in the current frame
[0039] Optionally, the S4 includes the following steps:
[0040] S41. Collecting the visual image corresponding to the current frame The acceleration and angular velocity data A output by the inertial measurement unit (t) , LiDAR point cloud data Construct input data triples And input into the visual-inertial-laser fusion synchronous positioning and mapping system for initial motion estimation and world pose tracking;
[0041] S42. Based on the feature registration relationship between the visual image corresponding to the current frame and the lidar point cloud data, construct the residual constraint term between image frames and the reprojection error between point cloud frames, and estimate the world pose of the current frame in the world coordinate system by combining the inertial integral. And construct the world pose graph G with continuous frame world poses as edges pose in, represents the world pose of the i-th frame, E (i,j) is the inter-frame error edge, used to construct the metric map;
[0042] S43. The semantic point cloud of the casualty in the current frame Invest in the semantic graph construction module and map each =The points in the injured area are taken as semantic key points, and raster mapping is performed in space according to their three-dimensional coordinates (x, y, z) to construct a semantic raster map Each grid cell g in the semantic grid map i,j,k Represents a 3D space cell, whose state is updated by whether there is a semantic point cloud;
[0043] S44. Set the key frame selection strategy during the metric map update process, when the inter-frame motion change exceeds the set threshold θ kf Or the number of semantic point clouds changes exceeds the threshold θ sem When the current frame is set as key frame K (t) , and its corresponding world pose LiDAR point cloud data Visual Image and semantic point cloud of the wounded Synchronously add the map management module;
[0044] S45. Establishing a metric map dataset With the semantic map dataset in Represents a set of keyframes;
[0045] S46. Keyframe collection As the index basis, the metric map dataset and the semantic map dataset are updated synchronously in each map optimization cycle, and the structural consistency is maintained in the global optimization graph, and finally the metric-semantic dual map dataset M is formed. dual .
[0046] Optionally, the S5 includes the following steps:
[0047] S51. Perform multi-source data registration on the current frame pixel-level semantic consistency enhancement mask, millimeter-wave vital signs radar data, and infrared thermal image data. Map the millimeter-wave vital signs radar data and infrared thermal image data to the image coordinate system based on the joint internal and external parameter matrix group between the RGB camera, millimeter-wave radar, and infrared thermal imager to establish a spatially aligned multimodal pixel-level observation map. (t) (x, y), where each pixel contains visual semantics, millimeter wave amplitude, and infrared temperature values that can be extracted simultaneously;
[0048] S52. From the spatially aligned multimodal pixel-level observation map O (t) Extract the pixel set with semantic label as the injured area in (x,y) Extract the corresponding millimeter wave reflection amplitude sequence for each pixel (x, y) in the pixel set And construct the frequency domain life waveform response function:
[0049]
[0050] Among them, r i (x,y) represents the millimeter wave reflection amplitude at the i-th sampling, is the average amplitude at the pixel point, which is used to measure the degree of periodic micro-vibration at the point. r is the number of reflection amplitude samples collected continuously for the same pixel point in the millimeter-wave vital signs radar data;
[0051] S53. Extract the corresponding infrared thermal image temperature gradient for each pixel point (x, y)
[0052] S54. The frequency domain life waveform response function of each pixel point Temperature gradient with infrared thermal imaging Perform normalization and union to construct pixel-level fusion vital sign score
[0053] S55. Pixel-level fusion of vital signs scores Map to the grid cells in the metric semantic dual map consistent with the world coordinate system, and build a life sign confidence grid map based on the mapping relationship Each grid cell corresponds to a spatial position (x, y, z), and its value is the average score of all matching pixels at that position:
[0054]
[0055] Among them, Ω i,j,k is the set of all pixels mapped to the grid (i, j, k), |Ω i,j,k | is the number of pixels, forming a dense scoring volume map with spatial position-vital signs as elements;
[0056] S56. Transform life signs into confidence grid Write Metrics - Semantic Dual Map Dataset M dual , and the world pose graph G pose Establish index binding relationship.
[0057] Optionally, the S6 includes the following steps:
[0058] S61. World pose based on current frame With dynamic anchor collection The three-dimensional coordinates p of each semantic point of the injured person i , calculate the Euclidean distance between the mobile platform and the injured person's dynamic anchor point
[0059]
[0060] in, Indicates the world pose of the current frame The translation vector part of ;
[0061] If there is a Euclidean distance where τ near is the preset distance threshold, at the preset distance threshold τ near Nearby, the precise relative position between the current mobile platform and the injured person is obtained through the time domain arrival difference positioning algorithm And calculate the updated estimated position of the injured person at the current moment
[0062] S62. Enable the micro-vibration radar module to collect short-range radar depth maps Compare it with the current frame RGB image and depth image Perform time synchronization and spatial registration to construct a three-channel radar fusion input image Send it to the improved YOLO network for casualty detection again;
[0063] S63. Extract the close-range radar depth map from the bounding box and mask area of the injured person output by the improved YOLO network Microseismic waveforms in the image are used to construct microvibration response functions. And with the millimeter wave life waveform response function and infrared temperature gradient A weighted fusion to generate an updated vital sign fusion score
[0064] S64. Estimate the location of the injured With updated vital signs fusion score The mean value of the corresponding grid is fused into the dynamic anchor point attribute update item of the injured person, and the dynamic anchor point set is updated. The dynamic anchor point set corresponding to the current frame
[0065] Optionally, the S7 includes the following steps:
[0066] S71. From the Metric-Semantic Dual Map Dataset M dual Get the dynamic anchor point set of the current frame Extract the three-dimensional coordinates of each casualty's dynamic anchor point Life status score is recorded as The semantic label is fixed to the injured area, and a unique ID is assigned to each injured person. i , and combine to form the wounded target set H of the current frame (t) ;
[0067] S72. Score the fusion life status of each wounded target According to the order of scores from high to low, the wounded target set H (t) Sort by priority and get a priority sequence of wounded numbers Π rescue , where the higher the score, the higher the priority. The original number will not be changed during the sorting process, only the treatment order will be adjusted;
[0068] S73. From the semantic map dataset M semantic Semantic regions with specific risk labels are extracted from the dataset, including flame regions, smoke regions, collapsed regions, and highly reflective metal regions. The set of spatial locations with potential dangers is defined as the risk semantic region set R. (t) ;
[0069] S74. World pose in current frame The translation vector in is the starting point, and the three-dimensional coordinates of the injured person are As the target point, the graph search path planning algorithm is used to generate an obstacle avoidance optimization navigation path connecting the two points, and finally the obstacle avoidance optimization navigation path P corresponding to each injured person is obtained. i ,The obstacle avoidance optimization navigation path is a continuous sequence of navigation nodes from the starting point to the target point;
[0070] S75. The unique number, three-dimensional coordinates, integrated life score and corresponding obstacle avoidance optimization navigation path P of each injured person are recorded. i Integrate into a navigation command data packet, collect the data packets corresponding to all the injured, build a complete data packet set, and send the data packet set to the rescue command terminal in real time through wired or wireless communication modules for use in rescue mission issuance, resource scheduling and path guidance.
[0071] Optionally, during the process of generating the obstacle avoidance optimization navigation path, the following obstacle avoidance rules are set:
[0072] If the area passed by the path does not belong to the risk semantic area set R (t) , the path is considered passable;
[0073] If the path needs to pass through the flame area, the cost weight of the path segment will be increased to reduce its selection priority;
[0074] If the path passes through a highly reflective metal area or a thick smoke area, the path's obstacle avoidance radius will be forcibly expanded to 1.5 times the normal travel distance;
[0075] If a path inevitably crosses a collapsed area, the path will be marked as a low-confidence path and the path replanning logic will be triggered.
[0076] The beneficial effects of the present invention are:
[0077] (1) Based on the YOLO backbone network, the present invention introduces a spatial-channel coupled attention module between Stage-3 and Stage-4, and enhances the backbone feature map by combining channel attention and spatial attention in parallel, so as to maintain stable detection performance in complex scenes such as smoke, reflection, and occlusion areas. At the same time, combined with the visibility weight map constructed by the depth map, the feature map is weightedly fused pixel by pixel to form an environment-adaptive feature map, which significantly improves the semantic extraction accuracy of the injured area under complex visual conditions.
[0078] (2) The present invention establishes a semantic projection mapping from pixel coordinates to laser point clouds by aligning the semantic mask with the depth map pixel by pixel, combining it with the external parameter calibration matrix of the laser radar, and propagates the casualty labels from the image domain to the point cloud domain through the point cloud space matching mechanism, constructing a three-dimensional point cloud map with semantic attributes in the centimeter-level space. The present invention further introduces a dynamic anchor mechanism, and updates the three-dimensional position and life status score of the casualty in real time under the joint perception of UWB, millimeter-wave radar and infrared sensor, thereby achieving high-confidence target positioning and hierarchical processing.
[0079] (3) The present invention simultaneously maintains two data sets, namely the metric map and the semantic map, and synchronously updates the map structure and the semantic point cloud information of the injured person under the drive of key frames. The millimeter-wave breathing amplitude and the infrared thermal image temperature gradient are jointly modeled into a vital sign scoring map. By sorting the injured targets in the scoring map and constructing a cost map in combination with the environmental semantic risk area, the graph search algorithm is guided to generate an obstacle avoidance optimization path, effectively improving the rescue efficiency and safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0081] Figure 1 This is a flowchart of a deep learning-based wounded life-saving detection and positioning method proposed by the present invention. DETAILED DESCRIPTION
[0082] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.
[0083] refer to Figure 1 , a deep learning-based wounded life-saving detection and positioning method, comprising the following steps:
[0084] S1. Collect and time-synchronize multi-view RGB-D images, LiDAR point clouds, millimeter-wave vital signs radar data, and infrared thermal imaging data;
[0085] S2. Input the RGB-D image into the improved YOLO network, which performs early semantic fusion in the middle layer of the pyramid and outputs pixel-level semantic masks.
[0086] S3. Based on the camera and lidar extrinsic calibration results, the depth channel corresponding to the pixel-level semantic mask is projected into a 3D coordinate system. The original lidar point cloud is assigned a semantic label for the casualty to generate a semantic point cloud of the casualty.
[0087] S4. Construct a metric map for centimeter-level pose resolution and use the semantic point cloud of the casualty to construct an environmental semantic map. A sparse keyframe mechanism is used to keep the metric map and the environmental semantic map updated synchronously, forming a metric-semantic dual-map dataset.
[0088] S5. Use pixel-level semantic masks to spatially align millimeter-wave vital sign radar data and infrared thermal imaging data, calculate respiratory amplitude and body surface temperature gradient, generate vital sign scores, and write them into the metric-semantic dual map dataset;
[0089] S6. When the distance between the mobile platform and the injured person's dynamic anchor point is less than a preset threshold, re-detection is performed to obtain an updated fusion score of the injured person's location and vital signs, and the attributes of the injured person's dynamic anchor point in the metric-semantic dual map dataset are updated;
[0090] S7. Based on the updated metric-semantic dual map dataset, the 3D coordinates of each injured person’s dynamic anchor point and rescue priority ranking are calculated to generate an obstacle avoidance optimized navigation path that avoids risky semantic areas.
[0091] In this embodiment, S2 includes the following steps:
[0092] S21. RGB image I RGB and depth image I D Input is an improved YOLO network built on the CSPDarkNet backbone structure. The CSPDarkNet backbone structure includes multiple CrossStagePartial residual blocks. The main path and cross-channel path in the CrossStagePartial residual block construct a multi-scale channel response structure in parallel, and output the Stage-3 backbone feature map F base ;
[0093] S22. Insert the spatial-channel coupled attention module between Stage-3 and Stage-4 of the CSPDarkNet backbone structure. The spatial-channel coupled attention module consists of the channel attention branch CA(·) and the spatial attention branch SA(·), which acts on the backbone feature map in parallel. Forming the attention-enhanced feature map:
[0094] F attn =CA(F base )+SA(F base );
[0095] CA(·) performs channel enhancement on the local texture of the human body area, SA(·) is used to highlight the spatial edge features in the irregular occlusion area, H represents the spatial height dimension of the feature map output by Stage-3 in the casualty detection image, W represents the spatial width dimension of the feature map, and C represents the number of semantic categories, including the casualty area, occlusion area, and background area.
[0096] S23. Based on depth image I D Build environment visibility weight map Where d(x,y) represents the depth of field value of the pixel point, d max is the maximum effective perception distance, and λ is the visual degradation control coefficient; it is used to simulate the degree of interference of smoke, occlusion, and reflection on detection ability;
[0097] S24. Enhance the attention feature map F attn and the environmental visibility weight map W env (x,y) performs pixel-by-pixel multiplication fusion to generate the environment adaptive feature map F env (x, y, c), used to enhance detection robustness under complex visual conditions;
[0098] S25. Input the environment adaptive feature map into the middle layer F of the feature pyramid m , and introduces a semantic decoding branch to generate the current frame semantic guidance mask on the environment adaptive feature map
[0099] S26. Extract the semantic guidance mask generated during the previous frame detection from the cache With the current frame semantic guidance mask Perform inter-frame fusion to construct semantically consistent enhancement masks:
[0100]
[0101] Where α is the dynamic fusion coefficient, which is dynamically adjusted according to the degree of image change or confidence.
[0102] In this embodiment, S3 includes the following steps:
[0103] S31. Enhance the semantic consistency mask And the depth image I of the corresponding frame D Perform pixel-by-pixel registration to map the depth of field value d(x,y) corresponding to the pixel point (x,y) in the mask to the three-dimensional space point in the camera coordinate system
[0104]
[0105] Where K is the camera intrinsic parameter matrix, which contains the focal length and principal point information; d(x,y) is the depth of field value of the corresponding pixel in the depth map, which has the same distance unit as the actual physical world;
[0106] S32. Based on the external parameter calibration matrix T between the camera and the lidar cam→lidar The three-dimensional space point Transform to the laser radar coordinate system to obtain the projection point The extrinsic calibration matrix is a 4x4 homogeneous transformation matrix that contains rotation and translation information:
[0107]
[0108] S33. Enhance the mask by all the pixels that meet the semantic consistency Projected point cloud collection Aggregate into the semantic projection point cloud of the current frame And assign a label to each point in the projected point cloud set = injured area, forming a label point set:
[0109]
[0110] in, Assign a semantic label to the point to represent the semantics of the injured area in the pixel mask corresponding to the point;
[0111] S34. Raw LiDAR point cloud data Point Cloud with Semantic Projection Perform spatial union in the lidar coordinate system, and pass the radius threshold δ match Perform nearest neighbor matching on the original lidar point cloud data The points that meet the following conditions are assigned semantic labels
[0112]
[0113] in, For the matched and labeled point cloud set, δ match is the point cloud fusion tolerance threshold, which is used to control the spatial propagation range of the projection label, p i and p j Represent the 3D points in the semantic point cloud and the original lidar point cloud respectively;
[0114] S35. The point cloud collection that is matched and annotated As the semantic point cloud of the casualty in the current frame
[0115] In this embodiment, S4 includes the following steps:
[0116] S41. Collecting the visual image corresponding to the current frame The acceleration and angular velocity data A output by the inertial measurement unit (t) , LiDAR point cloud data Construct input data triples And input into the visual-inertial-laser fusion synchronous positioning and mapping system for initial motion estimation and world pose tracking;
[0117] S42. Based on the feature registration relationship between the visual image corresponding to the current frame and the lidar point cloud data, construct the residual constraint term between image frames and the reprojection error between point cloud frames, and estimate the world pose of the current frame in the world coordinate system by combining the inertial integral. And construct the world pose graph G with continuous frame world poses as edges pose in, represents the world pose of the i-th frame, E (i,j) is the inter-frame error edge, used to construct the metric map;
[0118] S43. The semantic point cloud of the casualty in the current frame Invest in the semantic graph construction module and map each =The points in the injured area are taken as semantic key points, and raster mapping is performed in space according to their three-dimensional coordinates (x, y, z) to construct a semantic raster map Each grid cell g in the semantic grid map i,j,k Represents a 3D space cell, whose state is updated by whether there is a semantic point cloud;
[0119] S44. Set the key frame selection strategy during the metric map update process, when the inter-frame motion change exceeds the set threshold θ kf Or the number of semantic point clouds changes exceeds the threshold θ sem When the current frame is set as key frame K (t) , and its corresponding world pose LiDAR point cloud data Visual Image and semantic point cloud of the wounded Synchronously add the map management module;
[0120] S45. Establishing a metric map dataset With the semantic map dataset in Represents a set of keyframes;
[0121] S46. Keyframe collection As the index basis, the metric map dataset and the semantic map dataset are updated synchronously in each map optimization cycle, and the structural consistency is maintained in the global optimization graph, and finally the metric-semantic dual map dataset M is formed. dual .
[0122] In this embodiment, S5 includes the following steps:
[0123] S51. Perform multi-source data registration on the current frame pixel-level semantic consistency enhancement mask, millimeter-wave vital signs radar data, and infrared thermal image data. Map the millimeter-wave vital signs radar data and infrared thermal image data to the image coordinate system based on the joint internal and external parameter matrix group between the RGB camera, millimeter-wave radar, and infrared thermal imager to establish a spatially aligned multimodal pixel-level observation map. (t) (x, y), where each pixel contains visual semantics, millimeter wave amplitude, and infrared temperature values that can be extracted simultaneously;
[0124] S52. From the spatially aligned multimodal pixel-level observation map O (t) Extract the pixel set with semantic label as the injured area in (x,y) Extract the corresponding millimeter wave reflection amplitude sequence for each pixel (x, y) in the pixel set And construct the frequency domain life waveform response function:
[0125]
[0126] Among them, r i (x,y) represents the millimeter wave reflection amplitude at the i-th sampling, is the average amplitude at the pixel point, which is used to measure the degree of periodic micro-vibration at the point. r is the number of reflection amplitude samples collected continuously for the same pixel point in the millimeter-wave vital signs radar data;
[0127] S53. Extract the corresponding infrared thermal image temperature gradient for each pixel point (x, y)
[0128] S54. The frequency domain life waveform response function of each pixel point Temperature gradient with infrared thermal imaging Perform normalization and union to construct pixel-level fusion vital sign score
[0129] S55. Pixel-level fusion of vital signs scores Map to the grid cells in the metric semantic dual map consistent with the world coordinate system, and build a life sign confidence grid map based on the mapping relationship Each grid cell corresponds to a spatial position (x, y, z), and its value is the average score of all matching pixels at that position:
[0130]
[0131] Among them, Ω i,j,k is the set of all pixels mapped to the grid (i, j, k), |Ω i,j,k | is the number of pixels, forming a dense scoring volume map with spatial position-vital signs as elements;
[0132] S56. Transform life signs into confidence grid Write Metrics - Semantic Dual Map Dataset M dual , and the world pose graph G pose Establish index binding relationship.
[0133] In this embodiment, S6 includes the following steps:
[0134] S61. World pose based on current frame With dynamic anchor collection The three-dimensional coordinates p of each semantic point of the injured person i , calculate the Euclidean distance between the mobile platform and the injured person's dynamic anchor point
[0135]
[0136] in, Indicates the world pose of the current frame The translation vector part of ;
[0137] If there is a Euclidean distance where τ near is the preset distance threshold, at the preset distance threshold τ near Nearby, the precise relative position between the current mobile platform and the injured person is obtained through the time domain arrival difference positioning algorithm And calculate the updated estimated position of the injured person at the current moment
[0138] S62. Enable the micro-vibration radar module to collect short-range radar depth maps Compare it with the current frame RGB image and depth image Perform time synchronization and spatial registration to construct a three-channel radar fusion input image Send it to the improved YOLO network for casualty detection again;
[0139] S63. Extract the close-range radar depth map from the bounding box and mask area of the injured person output by the improved YOLO network Microseismic waveforms in the image are used to construct microvibration response functions. And with the millimeter wave life waveform response function and infrared temperature gradient A weighted fusion to generate an updated vital sign fusion score
[0140] S64. Estimate the location of the injured With updated vital signs fusion score The mean value of the corresponding grid is fused into the dynamic anchor point attribute update item of the injured person, and the dynamic anchor point set is updated. The dynamic anchor point set corresponding to the current frame
[0141] In this embodiment, S7 includes the following steps:
[0142] S71. From the Metric-Semantic Dual Map Dataset M dual Get the dynamic anchor point set of the current frame Extract the three-dimensional coordinates of each casualty's dynamic anchor point Life status score is recorded as The semantic label is fixed to the injured area, and a unique ID is assigned to each injured person. i , and combine to form the wounded target set H of the current frame (t) ;
[0143] S72. Score the fusion life status of each wounded target According to the order of scores from high to low, the wounded target set H (t) Sort by priority and get a priority sequence of wounded numbers Π rescue , where the higher the score, the higher the priority. The original number will not be changed during the sorting process, only the treatment order will be adjusted;
[0144] S73. From the semantic map dataset M semantic Semantic regions with specific risk labels are extracted from the dataset, including flame regions, smoke regions, collapsed regions, and highly reflective metal regions. The set of spatial locations with potential dangers is defined as the risk semantic region set R. (t) ;
[0145] S74. World pose in current frame The translation vector in is the starting point, and the three-dimensional coordinates of the injured person are As the target point, the graph search path planning algorithm is used to generate an obstacle avoidance optimization navigation path connecting the two points, and finally the obstacle avoidance optimization navigation path P corresponding to each injured person is obtained. i ,The obstacle avoidance optimization navigation path is a continuous sequence of navigation nodes from the starting point to the target point;
[0146] S75. The unique number, three-dimensional coordinates, integrated life score and corresponding obstacle avoidance optimization navigation path P of each injured person are recorded. i Integrate into a navigation command data packet, collect the data packets corresponding to all the injured, build a complete data packet set, and send the data packet set to the rescue command terminal in real time through wired or wireless communication modules for use in rescue mission issuance, resource scheduling and path guidance.
[0147] In this embodiment, during the process of generating the obstacle avoidance optimization navigation path, the following obstacle avoidance rules are set:
[0148] If the area passed by the path does not belong to the risk semantic area set R (t) , the path is considered passable;
[0149] If the path needs to pass through the flame area, the cost weight of the path segment will be increased to reduce its selection priority;
[0150] If the path passes through a highly reflective metal area or a thick smoke area, the path's obstacle avoidance radius will be forcibly expanded to 1.5 times the normal travel distance;
[0151] If a path inevitably crosses a collapsed area, the path will be marked as a low-confidence path and the path replanning logic will be triggered.
[0152] Example 1:
[0153] At 9:37 AM on November 2, 2024, a 5.8 magnitude shallow earthquake struck County A. Its epicenter was located in a mountainous area 3 kilometers south of Dachuan Town. Roads were severely damaged, communications were disrupted, and there were multiple house collapses, landslides, and power outages. The Ministry of Emergency Management dispatched the "SRU-07," an intelligent multi-sensor search and rescue robot platform equipped with the system proposed in this invention. Equipped with a multi-source sensing module (including an RGB-D camera, a 16-line lidar, a millimeter-wave vital signs radar, and an infrared thermal imager), it arrived at the core disaster area at 10:46 AM that day.
[0154] During the initial on-site survey, the system encountered typical complex post-disaster characteristics: diffuse dust, severe obstruction, and the presence of open flames and reflective metal panels in some areas. The accuracy of the traditional image-based YOLOv5 detection model dropped significantly in this scenario. During the comparative evaluation phase, the system used both the traditional YOLOv5 model and the improved YOLO network proposed in this paper, combining spatial-channel coupled attention with environmental visibility weighting, to detect casualties in Zone A, the epicenter's core area.
[0155] First, during the input phase, SRU-07 simultaneously collects RGB images (resolution 1280×720), depth images (effective sensing distance 0.4m to 5m), laser point clouds (scanning frequency 10Hz), infrared images (resolution 640×480), and millimeter-wave radar data (sampling rate 32 times per second). To address the smoke occlusion problem in the current frame image, the improved YOLO network introduces spatial-channel coupled attention to the Stage-3 backbone features and constructs a visibility weight map W based on the depth map. env (x,y), generates a feature map F with lighting adaptation capabilities env The semantic extraction F1-score of this feature map under complex visual conditions reached 0.872, an improvement of 21.8% compared to the 0.716 of the traditional YOLOv5.
[0156] Next, the system aligns the semantic mask with the depth map pixel by pixel, and combines the external parameter calibration matrix T cam→lidar The semantic points are projected onto the 3D laser point cloud coordinate system. At 11:02:13 on November 2, 2024, SRU-07 detected a trapped man in a brick-concrete structure in the northeast corner of Area A. The semantic point cloud contained 3482 3D points. The millimeter wave reflection signal had a significant respiratory fluctuation peak frequency band of 0.25Hz, and the surface temperature gradient in the corresponding area of the infrared image exceeded 4.8°C. The system dynamically extracted the millimeter wave reflection sequence within the semantic mask. And calculate the frequency domain life waveform response function:
[0157]
[0158] in, is the mean value of the reflection amplitude, N r = 32. The calculated average value of the corresponding pixel life score is L0(x, y) = 0.735.
[0159] In terms of map construction, the system uses the current frame pose With semantic point cloud A dual dataset of metric map and semantic map is established. When the map key frame update interval is set to 0.4m displacement or the semantic point number change exceeds 30%, a total of 13 key frames are selected to construct the A area M dual Map. A total of 7 anchors with vital signs were identified on the map. The system is based on fusion scoring The priority sorting results showed that the man had the highest score and was listed as the priority rescue target with ID=003.
[0160] In terms of path planning, the system will move the flame area and the collapsed high-reflective metal area from M semanticThe risk semantic region R(t) is extracted and set as A * The path algorithm is used to calculate the path from the current pose T w to p i The proposed path optimization module performs semantic obstacle avoidance optimization on the path. Traditional path planning (without risk semantics) takes 3.84 seconds and uses eight obstacle avoidance nodes. However, the proposed path optimization module generates a path in just 2.76 seconds with six nodes, improving the obstacle avoidance rate from 72% to 95% and the average path safety score from 0.61 to 0.83.
[0161] During the actual rescue verification phase, the rescue robot dispatched the coordinated drone to assist in delivering oxygen, bandages and positioning beacons according to the system navigation instructions. The rescue team successfully rescued the man at 11:23. His heart rate recovered to 78 beats per minute, his body surface temperature rose from 32.4℃ to 35.6℃, and his vital status score improved significantly.
[0162] To further verify the stability of the proposed method in multiple scenarios, the research team conducted comparative experiments at a post-earthquake simulation site in Ya'an, setting up three scenarios (A, B, and C) (representing low occlusion, smoke occlusion, and high reflection, respectively) for system evaluation. Table 1 below shows the comparison results between the traditional YOLO+ single-modal detection system and the proposed system:
[0163] Table 1 Comparison results between the traditional YOLO+ single-modal detection system and the proposed system
[0164]
[0165] In December 2024, the team also collaborated with the Emergency Management Institute of University B to deploy this system in a mine collapse accident simulation warehouse to conduct long-distance, complex environment life detection testing. In a 25m occlusion scenario, the system detected 11 real-life simulated casualties with a false detection rate of less than 5% and a missed detection rate of less than 7%, respectively, reductions of 21.4% and 18.9% compared to traditional systems.
[0166] In summary, the casualty detection and positioning method proposed in the present invention, based on improved YOLO semantic enhancement, cross-modal projection, and vital signs scoring mechanism, can realize multimodal data fusion and dynamic annotation and path guidance in three-dimensional space in complex environments. It effectively solves the problems of low recognition rate, large positioning error, and untimely response in traditional systems in scenes with occlusion, high reflection, and low visibility, and demonstrates good practical usability and technological advancement.
[0167] Based on the YOLO backbone network, the present invention introduces a spatial-channel coupled attention module between Stage-3 and Stage-4. The backbone feature map is enhanced by parallel combination of channel attention and spatial attention, and stable detection performance can be maintained in complex scenes such as smoke, reflections, and occlusion areas. At the same time, combined with the visibility weight map constructed by the depth map, the feature map is weightedly fused pixel by pixel to form an environment-adaptive feature map, which significantly improves the semantic extraction accuracy of the injured area under complex visual conditions.
[0168] The present invention establishes a semantic projection mapping from pixel coordinates to laser point clouds by aligning the semantic mask with the depth map pixel by pixel, combining it with the extrinsic parameter calibration matrix of the lidar. Furthermore, through the point cloud spatial matching mechanism, the casualty labels are propagated from the image domain to the point cloud domain, and a three-dimensional point cloud map with semantic attributes is constructed in centimeter-level space. Furthermore, a dynamic anchor mechanism is introduced to update the three-dimensional position and vital status score of the casualty in real time under the joint perception of UWB, millimeter-wave radar and infrared sensors, thereby achieving high-confidence target positioning and hierarchical processing.
[0169] The present invention simultaneously maintains two data sets, a metric map and a semantic map, and synchronously updates the map structure and the semantic point cloud information of the casualty under the drive of key frames. It also jointly models the millimeter-wave breathing amplitude and the infrared thermal image temperature gradient into a vital sign scoring map. By sorting the casualty targets in the scoring map and constructing a cost map in combination with the environmental semantic risk areas, the graph search algorithm is guided to generate an obstacle avoidance optimization path, effectively improving rescue efficiency and safety.
[0170] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A method for detecting and locating injured persons for life-saving rescue based on deep learning, characterized in that: The steps include: S1. Collect and time-synchronize multi-view RGB-D images, LiDAR point clouds, millimeter-wave vital signs radar data, and infrared thermal imaging data; S2. Input the RGB-D image into the improved YOLO network, which performs early semantic fusion in the middle layer of the pyramid and outputs a pixel-level semantic mask. S3. Based on the camera and lidar extrinsic calibration results, the depth channel corresponding to the pixel-level semantic mask is projected into a 3D coordinate system. The original lidar point cloud is assigned a semantic label for the casualty to generate a semantic point cloud of the casualty. S4. Construct a metric map for centimeter-level pose resolution and use the semantic point cloud of the casualty to construct an environmental semantic map. A sparse keyframe mechanism is used to keep the metric map and the environmental semantic map updated synchronously, forming a metric-semantic dual-map dataset. S5. Use pixel-level semantic masks to spatially align millimeter-wave vital sign radar data and infrared thermal imaging data, calculate respiratory amplitude and body surface temperature gradient, generate vital sign scores, and write them into the metric-semantic dual map dataset; S6. When the distance between the mobile platform and the injured person's dynamic anchor point is less than a preset threshold, re-detection is performed to obtain an updated fusion score of the injured person's location and vital signs, and the attributes of the injured person's dynamic anchor point in the metric-semantic dual map dataset are updated; S7. Based on the updated metric-semantic dual map dataset, the 3D coordinates of each injured person’s dynamic anchor point and rescue priority ranking are calculated to generate an obstacle avoidance optimized navigation path that avoids risky semantic areas.
2. The method for detecting and locating injured persons for life-saving based on deep learning according to claim 1, characterized in that: The S2 comprises the following steps: S21. RGB image I RGB and depth image I D Input is an improved YOLO network built on the CSPDarkNet backbone structure. The CSPDarkNet backbone structure includes multiple CrossStagePartial residual blocks. The main path and cross-channel path in the CrossStagePartial residual block construct a multi-scale channel response structure in parallel, and output the Stage-3 backbone feature map F base ; S22. Insert the spatial-channel coupled attention module between Stage-3 and Stage-4 of the CSPDarkNet backbone structure. The spatial-channel coupled attention module consists of the channel attention branch CA(·) and the spatial attention branch SA(·), which acts on the backbone feature map in parallel. Forming the attention-enhanced feature map F attn ; S23. Based on depth image I D Build environment visibility weight map Where d(x,y) represents the depth of field value of the pixel point, d max is the maximum effective perception distance, and λ is the visual degradation control coefficient; it is used to simulate the degree of interference of smoke, occlusion, and reflection on detection ability; S24. Enhance the attention feature map F attn and the environmental visibility weight map W env (x,y) performs pixel-by-pixel multiplication fusion to generate the environment adaptive feature map F env (x, y, c), used to enhance detection robustness under complex visual conditions; S25. Input the environment adaptive feature map into the middle layer F of the feature pyramid m , and introduces a semantic decoding branch to generate the current frame semantic guidance mask on the environment adaptive feature map S26. Extract the semantic guidance mask generated during the previous frame detection from the cache With the current frame semantic guidance mask Perform inter-frame fusion to construct semantically consistent enhancement masks: Where α is the dynamic fusion coefficient, which is dynamically adjusted according to the degree of image change or confidence.
3. The method for detecting and locating injured persons for life-saving rescue based on deep learning according to claim 1, characterized in that: The S3 includes the following steps: S31. Enhance the semantic consistency mask And the depth image I of the corresponding frame D Perform pixel-by-pixel registration to map the depth of field value d(x,y) corresponding to the pixel point (x,y) in the mask to the three-dimensional space point in the camera coordinate system S32. Based on the external parameter calibration matrix T between the camera and the lidar cam→lidar The three-dimensional space point Transform to the laser radar coordinate system to obtain the projection point The extrinsic calibration matrix is a 4x4 homogeneous transformation matrix that contains rotation and translation information; S33. Enhance the mask by all the pixels that meet the semantic consistency Projected point cloud collection Aggregate into the semantic projection point cloud of the current frame And assign a label l = injured area to each point in the projected point cloud set to form a label point set S34. Raw LiDAR point cloud data Point Cloud with Semantic Projection Perform spatial union in the lidar coordinate system, and pass the radius threshold δ match Perform nearest neighbor matching on the original lidar point cloud data The points that meet the following conditions are assigned semantic labels l, and the matched and labeled point cloud set is obtained. : in, For the matched and labeled point cloud set, δ match is the point cloud fusion tolerance threshold, which is used to control the spatial propagation range of the projection label, p i and p j Represent the 3D points in the semantic point cloud and the original lidar point cloud respectively; S35. The point cloud collection that is matched and annotated As the semantic point cloud of the casualty in the current frame 4. The method for detecting and locating injured persons for life-saving based on deep learning according to claim 3, characterized in that: The S4 comprises the following steps: S41. Collecting the visual image corresponding to the current frame The acceleration and angular velocity data A output by the inertial measurement unit (t) , LiDAR point cloud data Construct input data triplets and feed them into the visual-inertial-laser fusion simultaneous localization and mapping system for initial motion estimation and world pose tracking; S42. Based on the feature registration relationship between the visual image corresponding to the current frame and the lidar point cloud data, construct the residual constraint term between image frames and the reprojection error between point cloud frames, and estimate the world pose of the current frame in the world coordinate system by combining the inertial integral. And construct the world pose graph with continuous frame world poses as edges E (i,j) },in, represents the world pose of the i-th frame, E (i,j) is the inter-frame error edge, used to construct the metric map; S43. The semantic point cloud of the casualty in the current frame The semantic map construction module is used to construct a semantic grid map by mapping each point with a semantic label l = injured area as a semantic key point according to its three-dimensional coordinates (x, y, z) in space. Each grid cell g in the semantic grid map i,j,k Represents a 3D space cell, whose state is updated by whether there is a semantic point cloud; S44. Set the key frame selection strategy during the metric map update process, when the inter-frame motion change exceeds the set threshold θ kf Or the number of semantic point clouds changes exceeds the threshold θ sem When the current frame is set as key frame K (t) , and its corresponding world pose LiDAR point cloud data Visual Image and semantic point cloud of the wounded Synchronously add the map management module; S45. Establishing a metric map dataset With the semantic map dataset in Represents a set of keyframes; S46. Keyframe collection As the index basis, the metric map dataset and the semantic map dataset are updated synchronously in each map optimization cycle, and the structural consistency is maintained in the global optimization graph, and finally the metric-semantic dual map dataset M is formed. dual .
5. The method for detecting and locating injured persons for life-saving based on deep learning according to claim 4, characterized in that: The S5 comprises the following steps: S51. Perform multi-source data registration on the current frame pixel-level semantic consistency enhancement mask, millimeter-wave vital signs radar data, and infrared thermal image data. Map the millimeter-wave vital signs radar data and infrared thermal image data to the image coordinate system based on the joint internal and external parameter matrix group between the RGB camera, millimeter-wave radar, and infrared thermal imager to establish a spatially aligned multimodal pixel-level observation map. (t) (x, y), where each pixel contains visual semantics, millimeter wave amplitude, and infrared temperature values that can be extracted simultaneously; S52. From the spatially aligned multimodal pixel-level observation map O (t) Extract the pixel set with semantic label as the injured area in (x,y) Extract the corresponding millimeter wave reflection amplitude sequence for each pixel (x, y) in the pixel set And construct the frequency domain life waveform response function : Among them, r i (x,y) represents the millimeter wave reflection amplitude at the i-th sampling, is the average amplitude at the pixel point, which is used to measure the degree of periodic micro-vibration at the point. r is the number of reflection amplitude samples collected continuously for the same pixel point in the millimeter-wave vital signs radar data; S53. Extract the corresponding infrared thermal image temperature gradient for each pixel point (x, y) S54. The frequency domain life waveform response function of each pixel point Temperature gradient with infrared thermal imaging Perform normalization and union to construct pixel-level fusion vital sign score S55. Pixel-level fusion of vital signs scores Map to the grid cells in the metric semantic dual map consistent with the world coordinate system, and build a life sign confidence grid map based on the mapping relationship Each grid cell corresponds to a spatial position (x, y, z), and its value is the average score of all matching pixels at that position: Among them, Ω i,j,k is the set of all pixels mapped to the grid (i, j, k), |Ω i,j,k | is the number of pixels, forming a dense scoring volume map with spatial position-vital signs as elements; S56. Transform life signs into confidence grid Write Metrics - Semantic Dual Map Dataset M dual , and the world pose graph G pose Establish index binding relationship.
6. The method for detecting and locating injured persons for life-saving based on deep learning according to claim 5, characterized in that: The S6 comprises the following steps: S61. World pose based on current frame With dynamic anchor collection The three-dimensional coordinates p of each semantic point of the injured person i , calculate the Euclidean distance between the mobile platform and the injured person's dynamic anchor point If there is a Euclidean distance where τ near is the preset distance threshold, at the preset distance threshold τ near Nearby, the precise relative position between the current mobile platform and the injured person is obtained through the time domain arrival difference positioning algorithm And calculate the updated estimated position of the injured person at the current moment S62. Enable the micro-vibration radar module to collect short-range radar depth maps Compare it with the current frame RGB image and depth image Perform time synchronization and spatial registration to construct a three-channel radar fusion input image Send it to the improved YOLO network for casualty detection again; S63. Extract the close-range radar depth map from the bounding box and mask area of the injured person output by the improved YOLO network Microseismic waveforms in the image are used to construct microvibration response functions. And with the millimeter wave life waveform response function and infrared temperature gradient A weighted fusion to generate an updated vital sign fusion score S64. Estimate the location of the injured With updated vital signs fusion score The mean value of the corresponding grid is fused into the dynamic anchor point attribute update item of the injured person, and the dynamic anchor point set is updated. The dynamic anchor point set corresponding to the current frame 7. The method for detecting and locating injured persons for life-saving based on deep learning according to claim 6, characterized in that: The S7 comprises the following steps: S71. From the Metric-Semantic Dual Map Dataset M dual Get the dynamic anchor point set of the current frame Extract the three-dimensional coordinates of each casualty's dynamic anchor point Life status score is recorded as The semantic label is fixed to the injured area, and a unique ID is assigned to each injured person. i , and combine to form the wounded target set H of the current frame (t) ; S72. Score the fusion life status of each wounded target According to the order of scores from high to low, the wounded target set H (t) Sort by priority and get a priority sequence of wounded numbers Π rescue , where the higher the score, the higher the priority. The original number will not be changed in the sorting, only the treatment order will be adjusted; S73. From the semantic map dataset M semantic Semantic regions with specific risk labels are extracted from the dataset, including flame regions, smoke regions, collapsed regions, and highly reflective metal regions. The set of spatial locations with potential dangers is defined as the risk semantic region set R. (t) ; S74. World pose in current frame The translation vector in is the starting point, and the three-dimensional coordinates of the injured person are As the target point, the graph search path planning algorithm is used to generate an obstacle avoidance optimization navigation path connecting the two points, and finally the obstacle avoidance optimization navigation path P corresponding to each injured person is obtained. i ,The obstacle avoidance optimization navigation path is a continuous sequence of navigation nodes from the starting point to the target point; S75. The unique number, three-dimensional coordinates, integrated life score and corresponding obstacle avoidance optimization navigation path P of each injured person are recorded. i Integrate into a navigation command data packet, collect the data packets corresponding to all the injured, build a complete data packet set, and send the data packet set to the rescue command terminal in real time through wired or wireless communication modules for use in rescue mission issuance, resource scheduling and path guidance.
8. The method for detecting and locating injured persons for life-saving rescue based on deep learning according to claim 7, characterized in that: During the generation of the obstacle avoidance optimization navigation path, the following obstacle avoidance rules are set: If the area passed by the path does not belong to the risk semantic area set R (t) , the path is considered passable; If the path needs to pass through the flame area, the cost weight of the path segment will be increased to reduce its selection priority; If the path passes through a highly reflective metal area or a thick smoke area, the path's obstacle avoidance radius will be forcibly expanded to 1.5 times the normal travel distance; If a path inevitably crosses a collapsed area, the path will be marked as a low-confidence path and the path replanning logic will be triggered.
Citation Information
Cited By
Indoor personnel positioning method and system based on multi-modal perception
CN121252813A
Emergency rescue real-time human body detection method and equipment based on time sequence motion feature enhancement
CN121305671A
Positioning method and system for driving and anchoring all-in-one machine in coal mine driving roadway
CN121346766A
Laser driving intelligent regulation and control method and system based on deep learning
CN121454989A
AR (Augmented Reality) precise navigation method and system for mountainous area construction safety
CN121498702A