Disaster rescue system and disaster rescue method

By integrating multi-source data acquisition and fusion processing technology in the disaster rescue system, building a knowledge map and using GNN for inference, the problems of single detection means and insufficient data fusion in the existing system are solved, and accurate detection and precise rescue guidance for rescue personnel to be rescued at the disaster site are achieved.

CN120219133APending Publication Date: 2025-06-27WUHAN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510348792.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing disaster rescue system has problems such as single life detection methods, insufficient data fusion, low intelligence, insufficient real-time performance and need to be improved, making it difficult to accurately detect rescue workers waiting for disaster sites.

Method used

A disaster rescue system is designed, including a multi-source data acquisition unit, a data processing unit and a human-computer interaction unit. By collecting multimodal data (such as radar data, infrared image data, sound data, gas data, video data), the characteristics of multimodal data are fused to build a first knowledge graph, and the graph neural network (GNN) is used for inference to indicate the human body's orientation at the disaster site.

Benefits of technology

Accurate and efficient detection of rescue personnel waiting for disaster sites has been achieved, improving the accuracy of life detection and the accuracy of rescue operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219133A_ABST
    Figure CN120219133A_ABST
Patent Text Reader

Abstract

The invention discloses a disaster rescue system and a disaster rescue method, and belongs to the technical field of disaster rescue, the disaster rescue system comprises a multi-source data acquisition unit, a data processing unit and a man-machine interaction unit which are connected in sequence, and the multi-source data acquisition unit is used for acquiring multi-modal data of a disaster site; the multi-modal data comprises radar data, infrared image data, sound data, gas data and video data; the data processing unit is used for fusing the multiple features of the multi-modal data to obtain a fused feature set, constructing a first knowledge graph based on the fused feature set, and reasoning the first knowledge graph by adopting a graph neural network GNN to obtain a reasoning result, and the reasoning result is used for indicating the orientation of a human body in a disaster site; and the man-machine interaction unit is used for displaying the reasoning result. The disaster rescue system can accurately and efficiently detect the to-be-rescued personnel in a disaster site.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of disaster rescue, and particularly to a disaster rescue system and a disaster rescue method. Background Art

[0002] In the building collapse accidents caused by natural disasters or man-made accidents, the key to rescue work is to quickly and accurately detect the location and vital sign status of trapped people.

[0003] Disaster rescue systems in related technologies generally have problems such as single life detection means, insufficient data fusion, low intelligence level, insufficient real-time performance, and reliability to be improved. The limitations of single detection means make it difficult for disaster rescue systems to adapt to complex rescue environments, and the lack of an effective multi-source information fusion mechanism results in the inability to make full use of the comprehensive advantages of various sensors. Eventually, the disaster rescue systems in related technologies cannot accurately detect the people to be rescued at the disaster site. Summary of the Invention

[0004] The present disclosure provides a disaster rescue system and a disaster rescue method, which can accurately and efficiently detect the people to be rescued at the disaster site. The technical solutions at least include the following: In a first aspect, a disaster rescue system is provided. The disaster rescue system includes a multi-source data acquisition unit, a data processing unit, and a human-computer interaction unit connected in sequence. The multi-source data acquisition unit is used to acquire multi-modal data at the disaster site, and the multi-modal data includes radar data, infrared image data, sound data, gas data, and video data. The data processing unit is used to fuse multiple features of the multi-modal data to obtain a fused feature set, construct a first knowledge graph based on the fused feature set, and use a graph neural network (GNN) to reason about the first knowledge graph to obtain a reasoning result, where the reasoning result is used to indicate the orientation of the human body at the disaster site. The human-computer interaction unit is used to display the reasoning result.

[0005] Optionally, the data processing unit is further configured to fuse multiple features of the multimodal data in the following manner to obtain a fused feature set: convert the features of each modality in the multimodal data into an evidence form to obtain multiple evidences, where the evidences include probabilities or confidences; based on the improved Dempster-Shafer evidence theory, fuse the multiple evidences respectively under a spatial location recognition framework, a vital sign recognition framework, and an environmental state recognition framework to obtain a fused feature set, where the fused feature set includes the probabilities of each element in the spatial location recognition framework, the vital sign recognition framework, and the environmental state recognition framework; wherein, in the improved Dempster-Shafer evidence theory, each evidence has an adaptive weight, and the probability of any element in the spatial location recognition framework, the vital sign recognition framework, and the environmental state recognition framework is determined based on the evidence corresponding to the element and the adaptive weight of the evidence.

[0006] Optionally, the data processing unit is further configured to adjust the adaptive weight in the following manner: obtain environmental data, where the environmental data includes light intensity, environmental noise, and temperature; obtain a first threshold corresponding to a first modality, where the first modality is data of any modality in the multimodal data; adjust the weight of the evidence corresponding to the first modality based on the magnitude relationship between the environmental data and the first threshold.

[0007] Optionally, the data processing unit is further configured to adjust the adaptive weight in the following manner: obtain the signal-to-noise ratio and variance of a first modality, where the first modality is data of any modality in the multimodal data; when the signal-to-noise ratio of the first modality changes, adjust the weight of the evidence corresponding to the first modality based on the signal-to-noise ratio of the first modality, or when the variance of the first modality changes, adjust the weight of the evidence corresponding to the first modality based on the variance of the first modality.

[0008] Optionally, in the improved Dempster-Shafer evidence theory, the probability of the th element in the th recognition framework is determined by the following formula:

[0009] where represents the probability of the th element in the th recognition framework, , represents the spatial location, represents the vital sign, represents the environmental state, is the The element in a recognition framework corresponding to the th piece of evidence, is the adaptive weight of

[0010] Optionally, the th layer of the GNN is updated using the following formula:

[0011] where is the parameter of node in the th layer of the GNN, is the set of nodes adjacent to node , is the parameter of node in the th layer of the GNN, the node is one of the nodes in is the attention weight of node to node , is the parameter matrix of the th layer of the GNN, is a non-linear activation function.

[0012] Optionally, the first knowledge graph includes entities and relationships between entities. The entities of the first knowledge graph include vital signs, spatial locations, environmental states, and rescue resources; the relationships between the entities of the first knowledge graph include causal relationships, spatial relationships, temporal relationships, and inference relationships; constructing the first knowledge graph based on the fusion feature set includes: coupling the probabilities of the elements in the fusion feature set with multiple entities of the first knowledge graph to obtain the first knowledge graph.

[0013] Optionally, the data processing unit is further configured to perform three-dimensional scene modeling on the disaster site based on the radar data and the video data to obtain a first model of the disaster site; the human-computer interaction unit includes a visualization display module, and the visualization display module is configured to couple the infrared image data with the first model and then display the first model, and couple the inference result with the first model and then display the first model.

[0014] Optionally, the data processing unit is further configured to generate a rescue plan based on the inference result, where the rescue plan includes at least one rescue path for rescuing a first human body, and the first human body is a human body to be rescued at the disaster site; the visualization display module is further configured to mark the rescue path in the rescue plan on the first model.

[0015] In a second aspect, a disaster rescue method is further provided, including: collecting multimodal data of a disaster site based on a disaster rescue system; obtaining an inference result output by the disaster rescue system; and performing a rescue on a person to be rescued based on the inference result; where the disaster rescue system and the inference result are obtained based on the disaster rescue system described in the first aspect.

[0016] The beneficial effects brought by the technical solution provided by the embodiments of the present disclosure at least include: In the embodiments of the present disclosure, through the multi-source data collection unit, the collection of multimodal data of the disaster site is realized. The data processing unit processes the multimodal data, where multiple features of the multimodal data are fused to obtain a fused feature set. A first knowledge graph is constructed based on the fused feature set, and a GNN is used to perform inference on the first knowledge graph to obtain an inference result, and the inference result is used to indicate the position of the human body at the disaster site. In this way, when detecting life at the disaster site, the multimodal data of the disaster site can be effectively utilized, and the advantages of each modal data can be integrated to achieve life detection.

[0017] In addition, the knowledge graph has rich semantic information and logical information. After constructing the first knowledge graph based on the fused feature set and using a GNN to perform inference on the first knowledge graph, it is equivalent to making full use of the advantages of the multimodal data having comprehensive data and the knowledge graph having rich semantic information and logical information, so as to improve the accuracy of the finally obtained inference result, that is, life detection can be accurately performed based on this inference result, providing precise guidance for rescue operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 Shows a schematic structural diagram of a disaster rescue system provided by an exemplary embodiment of the present disclosure; Figure 2 Shows a schematic structural diagram of a disaster rescue system provided by another exemplary embodiment of the present disclosure; Figure 3 The flowchart of the disaster rescue method provided by an exemplary embodiment of the present disclosure is shown. Detailed implementation manners

[0020] Unless otherwise defined, technical terms or scientific terms used herein shall have the ordinary meanings as understood by those of ordinary skill in the art to which the present disclosure pertains. The terms "first", "second", "third" and similar terms used in the specification and claims of the patent application of the present disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Similarly, terms such as "a" or "an" do not denote a quantity limitation, but mean that there is at least one. The terms "including" or "comprising" and similar terms mean that the elements or items appearing before "including" or "comprising" cover the elements or items listed after "including" or "comprising" and their equivalents, and do not exclude other elements or items. The terms "connected" or "coupled" and similar terms are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.

[0021] To make the objectives, technical solutions and advantages of the present disclosure clearer, the following will further describe the embodiments of the present disclosure in detail with reference to the accompanying drawings.

[0022] The life detection means of the disaster rescue system in the related art are relatively single. For example, only one of the methods such as acoustic detection technology, infrared thermal imaging technology, life detection radar, and search and rescue dog-assisted detection is used to achieve life detection.

[0023] Among them, the acoustic detection technology locates by detecting the sound or vibration signals emitted by the trapped persons, but it is easily interfered by environmental noise and has a limited detection depth; the infrared thermal imaging technology detects by using the heat emitted by the human body, but it is easily affected by the environmental temperature and obstacles and has weak penetration ability; although the life detection radar has a certain penetration ability and can detect the breathing and heartbeat characteristics of the human body, its positioning accuracy is not high in a complex environment and it is easily interfered by multiple targets; although the search and rescue dog-assisted detection can find the trapped persons by smell detection, it is greatly limited by environmental conditions and cannot obtain accurate position information. Therefore, using only a single method for life detection is easily interfered by the environment at the disaster site, and ultimately results in the disaster rescue system in the related art being unable to accurately detect the persons to be rescued at the disaster site.

[0024] Figure 1 The structural schematic diagram of the disaster rescue system provided by an exemplary embodiment of the present disclosure is shown. This disaster rescue system is applicable to life detection at disaster sites such as earthquakes, collapses, and explosions. Refer to Figure 1, the disaster relief system includes a multi-source data acquisition unit 101, a data processing unit 102, and a human-computer interaction unit 103 that are connected in sequence.

[0025] The multi-source data acquisition unit 101 is used to collect multi-modal data at the disaster site. The multi-modal data includes radar data, infrared image data, sound data, gas data, and video data.

[0026] Optionally, the multi-source data acquisition unit 101 includes a micro radar sensor, an infrared thermal imaging sensor, a gas sensor, and a high-definition camera.

[0027] Among them, the micro radar sensor is used to collect radar data at the disaster site. Exemplarily, the operating frequency of the micro radar sensor is 24 GHz, and the detection distance is 0 - 30 meters.

[0028] The infrared thermal imaging sensor is used to collect infrared image data at the disaster site. Exemplarily, the resolution of the infrared thermal imaging sensor is at least 640×480, and the temperature resolution is at least 0.05℃.

[0029] The acoustic sensor array is used to collect sound data at the disaster site. Exemplarily, the acoustic sensor array is an 8-channel acoustic sensor array with a sampling frequency of 48 kHz.

[0030] The gas sensor is used to collect gas data at the disaster site. The gas data includes the concentration of life characteristic gases, such as the concentration of gases like CO2, NH3, etc.

[0031] The high-definition camera is used to collect video data at the disaster site. Exemplarily, the resolution of the high-definition camera is 1920×1080, and the frame rate is 60 fps. Since the disaster site is usually an extreme environment, in order to cope with the extreme environment, the high-definition camera can also be equipped with a wide-angle lens. The field of view angle of this wide-angle lens can be 120°, and the minimum illumination is 0.001 lux. In addition, the high-definition camera can also have an optical image stabilization function.

[0032] Optionally, the sampling frequency range of each sensor in the multi-source data acquisition unit 101 is 1 - 100 Hz. To adapt to the data acquisition requirements in different scenarios and ensure the real-time and synchronous nature of the data.

[0033] The data processing unit 102 is used to fuse multiple features of the multi-modal data to obtain a fused feature set, construct a first knowledge graph based on the fused feature set, and use a graph neural network (Graph Neural Network, GNN) to reason about the first knowledge graph to obtain a reasoning result. The reasoning result is used to indicate the orientation of the human body at the disaster site.

[0034] After the multi-source data acquisition unit 101 acquires multi-modal data, the data processing unit 102 first needs to preprocess the multi-modal data. The preprocessing includes signal denoising, normalization, image preprocessing, and time synchronization. Among them, signal denoising, normalization, and time synchronization are performed on all multi-modal data, and image preprocessing is performed on the data of the image modality. The data of the image modality includes infrared image data and video data.

[0035] The acquired multi-modal data usually has noise. Through signal denoising, the noise in the multi-modal data can be removed, and the signal quality can be improved. Optionally, wavelet transform combined with an adaptive filtering algorithm is used to implement signal denoising. First, the first signal is decomposed into sub-bands of different frequencies through wavelet transform, the effective signal is retained, the high-frequency noise is removed, and the denoised first signal is obtained. Then, the first signal is adaptively filtered, and the filter parameters are dynamically adjusted according to the characteristics of the first signal to further remove the noise of the first signal. Among them, the first signal is the data of any one modality in the multi-modal data.

[0036] There are differences in dimensions between multi-modal data. Through normalization, the dimensional differences between the data of each modality can be eliminated. Exemplarily, the Z-score normalization method can be used for normalization. There are many implementation methods of the Z-score normalization method in the related art, which are omitted here for detailed description.

[0037] Due to environmental interference at the disaster site (such as low visibility), the data of the image modality often has problems such as noise, blurriness, and geometric distortion.

[0038] Image preprocessing includes steps such as Gaussian filtering denoising, perspective transformation correction, and illumination compensation. Image preprocessing can improve the image quality and reduce the impact of environmental interference on the image data. Adaptive histogram equalization can enhance the contrast of the image, making the details in the image clearer. Especially in low-light environments, it can effectively improve the visibility of the image. Using Gaussian filtering denoising technology to smooth the image can remove noise interference and retain the effective information in the image. Perspective transformation correction technology can correct the possible geometric distortion of the image, ensure the spatial consistency of the image, and avoid image distortion caused by the viewing angle or sensor position. Illumination compensation technology can adjust the illumination conditions of the image, reduce the impact of uneven illumination on the image quality, and ensure stable image data in different illumination environments.

[0039] The multimodal data collected by the multi-source data acquisition unit 101 is obtained through multiple sensors. There may be errors in the timing of the data collected by the multiple sensors. Therefore, it is necessary to perform time synchronization processing on the multimodal data. In the embodiments of the present disclosure, timestamp alignment technology can be used for time synchronization processing to ensure the timing consistency of the data collected by each sensor.

[0040] The preprocessed multimodal data will be stored in the cache, and then the data processing unit 102 will process the multimodal data in the cache. The data caching mechanism ensures that data will not be lost during data transmission and guarantees the real-time nature of the data.

[0041] The data processing unit 102 first needs to extract the features of different-modal data, and then it can fuse the multiple features of the multimodal data to obtain a fused feature set.

[0042] For different-modal data, the extracted features are different, and the features extracted from each modality are all features related to the detection of a living body (human body). The features of radar data include Doppler frequency shift and range profile; the features of infrared image data include human hot spots and fire hot spots; the features of sound data include spectral features and Mel-Frequency Cepstral Coefficients (MFCC); the features of gas data include CO2 concentration, NH3 concentration, and the source directions of CO2 gas and NH3 gas; the features of video data include Histogram of Oriented Gradients (HOG), Optical Flow Feature (OFF), etc.

[0043] The radar sensor detects the target by emitting electromagnetic waves and receiving the reflected signals. When the target (i.e., the human body) has a small movement (such as breathing or heartbeat), the frequency of the reflected signal will change. This phenomenon is called Doppler frequency shift. Through the Doppler frequency shift of the radar data, the detection of small movements of the human body can be realized, so as to judge whether there is a human body. The range profile of the radar data reflects the distance between the target and the radar. Through the range profile of the radar data, the specific position of the human body can be determined, so as to achieve precise positioning of the trapped person. The range profile can also help distinguish multiple targets and reduce multi-target interference. There are many relevant techniques for obtaining the Doppler frequency shift and range profile, which are omitted here for detailed description.

[0044] An infrared thermal imaging sensor generates a temperature distribution image (i.e., infrared image data) by detecting the infrared radiation emitted by an object. The human body usually emits specific heat, so hot spot areas will appear in the temperature distribution image; a fire source emits a large amount of heat, so hot spot areas will also appear in the infrared image data, and these hot spot areas are significantly different from the hot spot areas related to the human body. By analyzing multiple hot spot areas in the infrared image data, the hot spot area related to the human body, i.e., the human body hot spot area, can be located.

[0045] Optionally, a deep learning method is used to determine the human body hot spot area from multiple hot spot areas. This method includes the following two steps: The first step is to obtain a set of infrared thermal imaging data.

[0046] The set of infrared thermal imaging data contains multiple infrared images, which include humans and non-humans (such as animals, mechanical equipment, etc.). Correspondingly, there are also hot spot areas of humans and non-humans in the infrared images. The hot spot areas in the multiple infrared images are labeled to label the human body hot spot areas and non-human body hot spot areas; The second step is to fine-tune the pre-trained image classification model.

[0047] The pre-trained image classification model includes but is not limited to ResNet, VGG, or EfficientNet, etc.

[0048] During fine-tuning, the labeled set of infrared thermal imaging data is divided into a training set, a validation set, and a test set. The pre-trained image classification model is trained using the training set. During the training process, the backpropagation algorithm is used to optimize the model parameters, and the cross-entropy loss function is used to measure the difference between the predicted results of the model and the true labels. The validation set is used to verify the performance of the model, and the model hyperparameters (such as learning rate, batch size, etc.) are adjusted to improve the model performance. The test set is used to evaluate the final performance of the model, and metrics such as accuracy, recall rate, and F1 score are calculated.

[0049] After fine-tuning, the infrared thermal imaging image is input into the image classification model, and the image classification model will output the probability that each hot spot is generated by a human body. A threshold (such as greater than or equal to 0.5) can be preset in advance. If the probability of a certain hot spot area in the output of the image classification model is greater than the threshold, it means that the hot spot area is a human body hot spot area.

[0050] In this way, the data processing unit 102 can automatically and accurately identify the human body hot spot area in the infrared image, reduce misjudgment, and improve the accuracy of life detection.

[0051] Trapped persons may emit calls for help or knocking sounds, and these sound signals will exhibit specific patterns in the frequency spectrum. By analyzing the frequency spectrum characteristics of the sound data, these sound signals related to the human body can be identified, thereby determining whether there are humans awaiting rescue. MFCC coefficients (Mel Frequency Cepstral Coefficients) are a feature extraction method commonly used in speech recognition. MFCC coefficients can capture the frequency spectrum envelope of the sound, which helps to more effectively identify human voices in a noisy environment. By extracting the MFCC coefficients of the sound data, the detailed characteristics of the sound can be further analyzed to improve the accuracy of sound recognition.

[0052] CO2 gas and NH3 gas are gases related to the human body. Whether there is a human can be indirectly judged by the CO2 concentration and NH3 concentration. For example, human respiration releases CO2, so an abnormal increase in the CO2 concentration may indicate the presence of a living being nearby. By analyzing the changing trend of the CO2 concentration, it can be determined whether there are trapped persons within a certain range; NH3 is usually produced by human sweat or excrement, especially in an enclosed space, and an increase in the NH3 concentration can be used as supplementary evidence for the presence of a living being.

[0053] When there are gas sensors in multiple different directions at the disaster site, the data of these gas sensors combined can reflect the distribution of CO2 gas and NH3 gas in the area. By the distribution of CO2 gas and NH3 gas at the disaster site, combined with environmental factors (such as wind speed, temperature, etc.), the source direction of CO2 gas and NH3 gas can be inferred, and this source direction can reflect the direction of the human body.

[0054] For video data, before extracting the features of the video data, it is first necessary to perform scene matching through SIFT (Scale-Invariant Feature Transform) and SURF (Speeded-Up Robust Features) so that the same scene can be identified from different perspectives. Then the features of the video data can be extracted. The features of the video data include Histogram of Oriented Gradients (HOG), Optical Flow Feature (OFF), etc.

[0055] Deep learning can also be used to process video data to extract multi-scale features, which are used for human target detection and human pose estimation. In implementation, ResNet50 can be used as the backbone network of the convolutional neural network to extract multi-scale features. The extracted multi-scale features can be used to implement human target detection through the YOLOv5 algorithm (such as detecting the human body to be rescued in the video); the multi-scale features can also be used to implement human pose estimation through the human key point detection algorithm based on HRNet. That is, the features of the video data can also include data related to human target detection and human pose estimation.

[0056] After the features of each modality in the multi-modal data are extracted, the data processing unit 102 is used to fuse the multiple features of the multi-modal data to obtain a fused feature set.

[0057] Optionally, the data processing unit 102 is also used to fuse the multiple features of the multi-modal data by the following steps a-b to obtain a fused feature set.

[0058] Step a, convert the features of each modality in the multi-modal data into evidence form to obtain multiple evidences.

[0059] The evidence includes probability or confidence.

[0060] The features of each modality in the multi-modal data come from data extracted by different types of sensors (such as radar, infrared, acoustic, gas sensors, etc.), and the forms of these features are different. However, through certain processing and conversion, these features can be unified into evidence form for fusion using the improved Dempster-Shafer evidence theory.

[0061] The features of radar data include Doppler frequency shift and range spectrum. By analyzing the Doppler frequency shift and range spectrum, the position and motion state of the human body to be rescued can be obtained, and then the Doppler frequency shift and range spectrum can be converted into the probability or confidence about the existence of the human body to be rescued, that is, the evidence corresponding to the radar data.

[0062] The features of infrared image data include the human body hot spot area. By analyzing the human body hot spot area, it can be judged whether there is a human body to be rescued, and then the human body hot spot area can be converted into the probability or confidence about the existence of the human body to be rescued, that is, the evidence corresponding to the infrared image data.

[0063] The features of sound data include spectral features and MFCC coefficients. By analyzing the spectral features and MFCC coefficients, it can be judged whether there is a call for help or a knocking sound, and then the spectral features and MFCC coefficients can be converted into the probability or confidence of detecting a sound signal, that is, the evidence corresponding to the sound data.

[0064] The characteristics of gas data include the CO2 concentration, the NH3 concentration, and the source directions of CO2 gas and NH3 gas. By analyzing the changes in the CO2 concentration and the NH3 concentration and their source directions, it is possible to indirectly determine whether there is a person in need of rescue. Furthermore, the CO2 concentration, the NH3 concentration, and the source directions of CO2 gas and NH3 gas can be converted into the probability or confidence level regarding the existence of a person in need of rescue, that is, the evidence corresponding to the gas data.

[0065] The characteristics of video data include the relevant data of human target detection and human pose estimation. By analyzing the relevant data of human target detection and human pose estimation, it is possible to determine whether there is a person in need of rescue. Furthermore, the characteristics of video data can be converted into the probability or confidence level regarding the existence of a person in need of rescue, that is, the evidence corresponding to the video data.

[0066] Each piece of evidence can be represented as a probability value or a confidence level value. Usually, the value range of the probability value or the confidence level value is between 0 and 1.

[0067] Step b: Based on the improved Dempster - Shafer evidence theory, multiple pieces of evidence are fused respectively under the spatial position recognition framework, the vital sign recognition framework, and the environmental state recognition framework to obtain a fused feature set.

[0068] The fused feature set includes the probabilities of each element in the spatial position recognition framework, the probabilities of each element in the vital sign recognition framework, and the probabilities of each element in the environmental state recognition framework.

[0069] Among them, in the improved Dempster - Shafer evidence theory, each piece of evidence has an adaptive weight, and the probability of any element in the spatial position recognition framework, the vital sign recognition framework, and the environmental state recognition framework is determined based on the evidence corresponding to the element and the adaptive weight of the evidence.

[0070] There is a recognition framework in the Dempster - Shafer evidence theory. The recognition framework is a set that contains all possible mutually exclusive and complete elementary events. The recognition framework represents all possible hypotheses or states in the problem domain. In the embodiments of the present disclosure, a problem in the problem domain of any recognition framework is an element.

[0071] In the identification framework of the traditional Dempster-Shafer evidence theory, the probability of each element is determined by the Basic Probability Assignment (BPA). Usually, empirical values are used for the probability assignment of each element by the BPA, which may lead to incorrect probability assignment for each element. In the embodiments of the present disclosure, an adaptive weight is adopted to implement the BPA, and a dynamic weight assignment mechanism is introduced to dynamically adjust the weight according to factors such as the reliability of sensors, environmental conditions, and historical data, so as to improve the Dempster-Shafer evidence theory, enhance the accuracy and robustness of multi-sensor fusion, and meet the application requirements in complex and changeable disaster rescue environments.

[0072] Optionally, in the improved Dempster-Shafer evidence theory, the probability of the -th element in the -th identification framework is determined by formula (1), and formula (1) is the BPA in this improved Dempster-Shafer evidence theory.

[0073] (1) In formula (1), represents the probability of the -th element in the -th identification framework, , represents the spatial position, represents the vital signs, represents the environmental state, is the -th evidence corresponding to the -th element in the -th identification framework, is 's adaptive weight.

[0074] Optionally, 's adaptive weight is calculated by formula (2).

[0075] (2) In formula (2), represents the reliability score of the -th evidence, and this reliability score is related to the modality corresponding to the -th evidence, represents the reliability score of the -th evidence, 's value range is from 1 to , is the The total number of evidences corresponding to the th element in a recognition framework. The essence of formula (2) is to use the proportion of the reliability score of the th evidence in the total reliability score of the evidences corresponding to the th element in the th recognition framework as the adaptive weight. The meanings of other parameters in formula (2) are the same as those in formula (1), and are not elaborated here.

[0076] Here, when takes , the th recognition framework is the spatial location recognition framework. When takes , the th recognition framework is the vital sign recognition framework. When takes , the th recognition framework is the environmental state recognition framework.

[0077] When takes , multiple elements in the th recognition framework are multiple locations. At this time, the probability of each element is the probability that the person to be rescued is in each location. For example, represents the probability that the person to be rescued is in the th location. represents the th evidence corresponding to the th location. For example, the spectral characteristics of a certain radar data indicate that there may be a person to be rescued at the th location, then this radar data is an evidence corresponding to the th location.

[0078] When takes , multiple elements in the th recognition framework are multiple life states. Life states include the state of life existence (such as alive, in danger, uncertain), physiological parameters (such as respiratory activity, cardiac activity, body temperature state), and consciousness state (such as awake, coma), etc. At this time, the probability of each element is the probability that the person to be rescued is in this life state. For example, represents the probability that the person to be rescued is in the th life state. represents the th evidence corresponding to the For example, if the spectral characteristics of a certain sound data indicate that a human body to be rescued is making a distress call, then the distress call is an evidence corresponding to the two elements of the state of life (surviving) and the state of consciousness (awake).

[0079] When taking at this time, multiple elements in the th recognition framework are multiple environmental states. The environmental state is used to represent the environmental conditions at the disaster site. The environmental state includes environmental parameters (such as temperature state, humidity state, air quality), risk factors (such as fire source, toxic gas, structural instability), and rescue conditions (such as passability, visibility, sound propagation conditions). At this time, the probability of each environmental state is the probability that the disaster site is in this environmental state. For example represents the probability that the disaster site is in the th environmental state. represents the th evidence corresponding to the th environmental state. For example, if the fire hot spot area of a certain infrared image data indicates that there is a fire source at the disaster site, then the infrared image data is an evidence corresponding to the element of risk factor (fire source).

[0080] In a possible implementation, since the external environment will also affect the data collected by the sensor, for example, it will affect the sensor accuracy, the adaptive weight can be adjusted based on the environmental data. In this case, the data processing unit 102 implements the adjustment of the adaptive weight by the following steps c - e.

[0081] Step c, obtain environmental data.

[0082] The environmental data includes light intensity, environmental noise, and temperature.

[0083] The light intensity can be obtained through a high - definition camera, the environmental noise can be obtained through an acoustic sensor array, and the temperature can be obtained through an infrared thermal imaging sensor.

[0084] Step d, obtain the first threshold corresponding to the first modality.

[0085] The first modality is any one of the modalities in the multi - modal data; Step e, adjust the weight of the evidence corresponding to the first modality based on the magnitude relationship between the environmental data and the first threshold.

[0086] Adjusting the weight corresponding to the evidence of the first modality in the embodiments of the present disclosure is essentially to adjust the reliability score of the evidence of the first modality. An initial value can be set for the reliability score. For example, the initial value is set to 100, indicating that by default all evidence is reliable in the initial state. Then, the reliability score of each piece of evidence is adjusted according to steps c - e, so as to achieve adaptive weights.

[0087] When considering the influence of the environment on the accuracy of sensors, three main situations are targeted: the influence of light intensity on infrared thermal imaging sensors and high - definition cameras; the influence of environmental noise on acoustic sensor arrays; and the influence of temperature on infrared sensors.

[0088] The lower the light intensity, the less effective information the camera can collect. At this time, the video data collected by the high - definition camera may become blurred, while the infrared image data collected by the infrared thermal imaging sensor is not affected by light. Therefore, a first light threshold can be set. When the light intensity is lower than the first light threshold, the reliability score of the evidence corresponding to the video data is reduced, and at the same time, the reliability score of the evidence corresponding to the infrared image data is increased (for example, it can be increased by 30%).

[0089] The higher the light intensity, the clearer the video data collected by the high - definition camera. Therefore, a second light threshold can be set. When the light intensity is higher than the second light threshold, the reliability score of the evidence corresponding to the video data is increased (for example, it can be increased by 25%).

[0090] Exemplarily, the first light threshold is 50 lux, and the second light threshold is 1000 lux.

[0091] The greater the environmental noise, the greater the proportion of environmental noise in the sound data collected by the acoustic sensor array. At this time, the reliability of the sound data is reduced. Therefore, a noise threshold can be set. When the environmental noise exceeds the noise threshold, the reliability score of the evidence corresponding to the sound data is reduced (for example, when the environmental noise exceeds the noise threshold, it is reduced by 40%, and for every 10 dB increase in environmental noise, the reliability score of the evidence corresponding to the acoustic sensor array is further reduced by 10%). Exemplarily, the noise threshold is 70 dB.

[0092] When the ambient temperature is close to the human body temperature, the human hot spot area in the infrared image data is not prominent. At this time, there may be errors in the human hot spot area identified by deep learning. Therefore, two temperature thresholds can be set. When the ambient temperature is less than the first temperature threshold, or when the ambient temperature is greater than the second temperature threshold, the reliability score of the evidence corresponding to the infrared image data is increased. When the ambient temperature is close to the human body temperature (32°C - 38°C), the reliability score of the evidence corresponding to the infrared image data is reduced.

[0093] Exemplarily, the first temperature threshold is 0°C, and the second temperature threshold is 35°C.

[0094] In another possible implementation, since the signal-to-noise ratio of the sensor can reflect the reliability of the data collected by the sensor, and the variance of the data collected by the sensor can reflect the stability of the data collected by the sensor, the adaptive weight can be adjusted based on the signal-to-noise ratio and the variance. In this case, the data processing unit 102 adjusts the adaptive weight by adopting the following steps f-g.

[0095] Step f, obtain the signal-to-noise ratio and variance of the first modality.

[0096] The first modality is the data of any one modality in the multi-modal data.

[0097] Step g, when the signal-to-noise ratio of the first modality changes, adjust the weight corresponding to the evidence of the first modality based on the signal-to-noise ratio of the first modality; or, when the variance of the first modality changes, adjust the weight corresponding to the evidence of the first modality based on the variance of the first modality.

[0098] Optionally, step g includes: taking the signal-to-noise ratio in the initial state of the first modality as the initial signal-to-noise ratio (the signal-to-noise ratio in the initial state can be the signal-to-noise ratio of the first modality when the disaster rescue system first enters a certain disaster site). Compared with the initial signal-to-noise ratio, every time the signal-to-noise ratio of the first modality decreases by 3 dB, the reliability score of the evidence corresponding to the first modality decreases by 15%. In addition, there is also a signal-to-noise ratio threshold for each modality. When the signal-to-noise ratio of the first modality is lower than the signal-to-noise ratio threshold of the first modality, the reliability score of the evidence corresponding to the first modality decreases by 50%.

[0099] The variance of the first modality can reflect the stability of the data of the first modality. If the fluctuation of the variance of the first modality exceeds the fluctuation threshold, it indicates that the data stability of the first modality is poor, and there may be a sensor failure. At this time, the reliability score of the evidence corresponding to the first modality decreases by 20%.

[0100] Through the above steps c-e or steps f-g, BPA can be implemented using the adaptive weight, which is equivalent to introducing a dynamic weight allocation mechanism in BPA, thereby improving the Dempster-Shafer evidence theory and enhancing the accuracy and robustness of multi-sensor fusion to meet the application requirements in the complex and changeable disaster rescue environment.

[0101] After obtaining the fusion feature set, the data processing unit 102 is further configured to construct a first knowledge graph based on the fusion feature set.

[0102] Optionally, the first knowledge graph includes entities and the relationships between entities. The entities of the first knowledge graph include vital signs, spatial location, environmental status, and rescue resources.

[0103] The vital sign entity represents the physiological state of the trapped person. The vital sign entity includes the state of life existence (such as alive, in danger, uncertain), physiological parameters (such as respiratory activity, cardiac activity, body temperature state), consciousness state (such as conscious, unconscious), etc.

[0104] The spatial location entity represents the spatial information in the rescue scene. The spatial location entity includes the location of the trapped person (three-dimensional coordinates and uncertainty range), the location of obstacles (such as walls, beams, rubble piles), spatial regions (such as danger zones, safe passages, search and rescue priority areas), etc.

[0105] The environmental state entity represents the conditions of the rescue environment. The environmental state entity includes environmental parameters (such as temperature state, humidity state, air quality), risk factors (such as fire sources, toxic gases, structural instability), rescue conditions (such as passability, visibility, sound propagation conditions), etc.

[0106] The rescue resource entity represents the equipment and personnel available for rescue. The rescue resource entity includes rescue equipment (such as life detectors, excavation equipment, medical equipment), rescue personnel (such as search and rescue team members, medical personnel), rescue strategies (such as direct contact, remote detection, demolition rescue), etc.

[0107] The relationships between the entities in the first knowledge graph include causal relationships, spatial relationships, temporal relationships, and reasoning relationships.

[0108] Causal relationships represent the causal connections between entities. For example, "physiological activities" cause "environmental changes" (such as respiratory activity causing an increase in the surrounding CO2 concentration), and "environmental state" affects "vital signs" (such as a high-temperature environment affecting the body temperature state), etc.

[0109] Spatial relationships: represent the spatial connections between entities. For example, the "trapped person" is located in the "spatial region", and the "obstacle" blocks the "rescue passage", etc.

[0110] Temporal relationships: represent the change relationships of entity states over time. For example, the temporal evolution trend of "vital signs", the change pattern of "environmental parameters", etc.

[0111] Reasoning relationships: represent the relationships that can be used for logical reasoning. For example, "multiple sources of evidence" support the "judgment of life existence", and "environmental conditions" restrict the "selection of rescue plans", etc.

[0112] Each entity and relationship can carry attributes. For example, the normal range of physiological parameters (such as the normal range of respiratory activity intensity), the accuracy of location information (such as the uncertainty radius of coordinates), the confidence level of the relationship (such as the reliability score of a certain reasoning relationship), etc.

[0113] Based on the above entities and the relationships between entities, a preset knowledge graph can be constructed. Based on this preset knowledge graph and the fused feature set, the data processing unit 102 can construct a first knowledge graph. The method includes: coupling the probabilities of each element in the fused feature set with multiple entities in the first knowledge graph to obtain the first knowledge graph.

[0114] Here, the multiple entities and the relationships between entities in the preset knowledge graph and the first knowledge graph are the same. The essence of the first knowledge graph is to couple the probabilities of each element in the fused feature set on the basis of the preset knowledge graph. Coupling the probabilities of each element in the fused feature set on the basis of the preset knowledge graph is also to store the probabilities of each element in the fused feature set corresponding to the corresponding entities in the preset knowledge graph, which is equivalent to assigning a probability and a feature parameter to the entities in the preset knowledge graph.

[0115] In addition, it is also necessary to map the multi-modal data into the first knowledge graph through semantic relationships, so that each entity in the first knowledge graph has relevant parameters of the disaster scene. After mapping the multi-modal data into the first knowledge graph through semantic relationships, the spatial location entity in the first knowledge graph contains three-dimensional coordinates and spatial relationship information; the life state entity in the first knowledge graph contains comprehensive parameters indicating whether there are people in the disaster scene; the environmental state entity in the first knowledge graph can represent the degree of harm of the environment to the human body.

[0116] The knowledge graph itself is a manifestation of semantic relationships and logical relationships. The first knowledge graph obtained after the above processing contains both semantic relationships and logical relationships related to disaster rescue, as well as on-site parameters of the disaster scene and the confidence levels of each on-site parameter (i.e., the probabilities of each element in the fused feature set). At this time, the data processing unit 102 can use a GNN to reason about the first knowledge graph to obtain a reasoning result.

[0117] Optionally, the first knowledge graph can be stored in the form of a graph database (such as Neo4j) or a relational database, which is convenient for subsequent querying and reasoning operations on the first knowledge graph. The graph database can efficiently process graph-structured data and support complex graph traversal and query operations.

[0118] The GNN can effectively process graph-structured data, capture the relationships between entities, and dynamically adjust the importance of different entities and relationships through an attention mechanism. The input of the GNN is the nodes (equivalent to entities) and edges (equivalent to the relationships between entities) in the first knowledge graph. Each node and edge can carry attributes (such as confidence levels, parameter ranges, etc.). Through a multi-layer message passing mechanism, each node can be gradually updated.

[0119] Optionally, the The $l$-th layer is updated using Equation (3): (3) where is the parameter of node in the -th layer of the GNN, is the set of nodes adjacent to node , is the parameter of node in the -th layer of the GNN, node is a node in , is the attention weight of node to node and is calculated using Equation (4), is the parameter matrix of the -th layer of the GNN,

[0120] (4) In Equation (4), is an activation function that is used to introduce non-linear features, is the parameter vector of the attention mechanism, denotes concatenating and these two vectors.

[0121] During the inference process, through the iterative update of multiple layers of GNNs, the GNN can capture complex entity relationships and attribute information in the knowledge graph. Finally, based on the node representations updated by the GNN, the data processing unit 102 can generate an inference result. The inference result is used to indicate the probability of each entity in the knowledge graph related to the human body to be rescued. Among entities of the same type, the entity with the highest probability is the entity related to the human body to be rescued, and thus the location of the human body to be rescued can be determined based on these entities.

[0122] For example, if there are multiple spatial location entities in the knowledge graph, the spatial location entity with the highest probability in the inference result is the location where the human body to be rescued is located.

[0123] The human-computer interaction unit 103 is used to display the inference result. After the data processing unit obtains the inference result, the inference result can be displayed through the human-computer interaction unit 103.

[0124] In the embodiments of the present disclosure, through the multi-source data acquisition unit, multi-modal data of the disaster scene is acquired, and the multi-modal data is processed by the data processing unit. Among them, multiple features of the multi-modal data are fused to obtain a fused feature set, a first knowledge graph is constructed based on the fused feature set, and a GNN is used to reason about the first knowledge graph to obtain a reasoning result, and the reasoning result is used to indicate the orientation of the human body in the disaster scene. In this way, when detecting life in the disaster scene, the multi-modal data of the disaster scene can be effectively utilized, and the advantages of each modal data can be integrated to achieve life detection.

[0125] In addition, the knowledge graph has rich semantic information and logical information. After constructing the first knowledge graph based on the fused feature set and using GNN to reason about the first knowledge graph, it is equivalent to making full use of the advantages of the multi-modal data having comprehensive data and the knowledge graph having rich semantic information and logical information, so as to improve the accuracy of the finally obtained reasoning result, that is, based on this reasoning result, life detection can be accurately performed, providing precise guidance for rescue operations.

[0126] Figure 2 Fig. shows the structural schematic diagram of a disaster rescue system provided by an exemplary embodiment of the present disclosure. Refer to Figure 1 , the disaster rescue system includes a multi-source data acquisition unit 101, a data processing unit 102, and a human-computer interaction unit 103 that are connected in sequence. Among them, the human-computer interaction unit 103 includes a visualization display module 201, an interaction module 202, and an alarm module 203.

[0127] The relevant content of the multi-source data acquisition unit 101 and the data processing unit 102 is the same as that in Figure 1 and is omitted here for detailed description.

[0128] Optionally, the visualization display module 201 is a display screen, and the visualization display module 201 is used to display the reasoning result. In the interface design of the display screen, the screen brightness and contrast can be automatically adjusted according to the ambient light, support automatic switching between day / night modes, and the night mode uses low-brightness red light display to protect night vision ability. The interface layout and function priorities can be dynamically adjusted according to the user's operation habits, and dedicated interface templates are provided for different rescue scenarios.

[0129] The interaction module 202 is, for example, a waterproof keyboard, physical function buttons, or a voice input system, and the interaction module 202 is used to control the disaster rescue system.

[0130] In some embodiments, the visualization display module 201 and the interaction module 202 can also be combined, for example, combined into a touch screen.

[0131] In a possible implementation, the data processing unit 102 is further used to perform three-dimensional scene modeling of the disaster site based on the radar data and the video data to obtain a first model of the disaster site. The visualization display module 201 is used to display the first model after coupling the infrared image data with the first model, and to display the first model after coupling the inference result with the first model.

[0132] Radar data and video data can reflect the conditions of the disaster site. Based on the radar data and video data, the data processing unit 102 can perform three-dimensional reconstruction of the disaster site to obtain a first model. Exemplarily, the data processing unit 102 uses a SLAM algorithm to combine point cloud data in the radar data and RGB images captured in the video data to achieve high-precision scene reconstruction.

[0133] When establishing the first model, the octree data structure can also be used to optimize the first model storage and rendering efficiency, support multi-resolution rendering, dynamically adjust the model detail level according to the viewpoint distance, implement a WebGL-based 3D engine, and support interactive operations such as model rotation, scaling, and sectioning.

[0134] After the visualization module 201 couples the infrared image data with the first model, it is equivalent to seamlessly integrating the thermal map into the first model, thereby providing spatial context information for the infrared image.

[0135] Optionally, the visualization display module 201 is also used to display the features of the infrared image data, such as the hot spots of the human body, and supports the overlay display of multiple layers of thermal maps, including temperature distribution, sound intensity, gas concentration, etc. The color range can also be dynamically adjusted according to the data distribution through an adaptive color mapping algorithm to enhance the contrast and realize a time series thermal map to display the changing trend of the infrared features of the human body to be rescued over time. In addition, the visualization display module 201 can also display the thermal value statistics in a certain area.

[0136] After the visualization module 201 couples the reasoning result with the first model, it is equivalent to marking the position and life status of the person to be rescued in the first model. Different color codes can also be used to represent different confidence levels of the entity.

[0137] The position of the person to be rescued displayed by the visualization module 201 can be an accurate three-dimensional coordinate with an accuracy better than m, and can also display the coordinate confidence interval to visualize the uncertainty of each 3D coordinate in the form of an ellipse or cube. It supports dual representation of relative and absolute coordinates, and realizes coordinate history trajectory recording to track the possible movement of trapped people.

[0138] The life status output displayed by the visualization module 201 includes the estimated respiratory rate and heart rate, provides an assessment of the vital sign intensity, displays a time-varying trend graph of the vital sign parameters, and evaluates the life status of the trapped person according to the inference result of the knowledge graph.

[0139] Optionally, the data processing unit 102 is further configured to generate a rescue plan based on the inference result. The rescue plan includes at least one rescue path for rescuing the first human body, where the first human body is the human body to be rescued at the disaster site; the visualization module 201 is further configured to mark the rescue path in the rescue plan on the first model.

[0140] The rescue path displayed by the visualization module 201 is generated by the data processing unit 102 based on the inference result. The data processing unit 102 can generate multiple optional rescue paths, mark the difficulty and risk level of each path, intelligently recommend the best rescue tools and methods, such as the selection of demolition equipment, approach methods, etc., estimate the rescue time window, help the rescue team formulate a time plan, and support the manual adjustment and optimization of the rescue plan.

[0141] Optionally, the visualization module 201 is further configured to display the waveforms of multi-modal data. For example, using the WebGL drawing library, set the waveform update rate of at least 60 frames per second to achieve synchronous display of multi-channel waveforms. The displayed waveforms include radar signals, sound waveforms, vital sign parameters, etc. When displaying waveforms, it supports functions such as waveform zooming, panning, pausing, and playback, integrates waveform analysis tools, including spectrum analysis, correlation analysis, etc., and realizes the function of automatic detection and marking of abnormal waveforms to help the staff quickly identify potential life signals.

[0142] Optionally, the visualization module 201 is further configured to display the features of video data. For example, it can overlay target detection frames on the video data in real time and mark the positions of the persons to be rescued; it can also use a skeleton model to visualize the human posture to help evaluate the status of the persons to be rescued. When displaying the features of video data, it can display the scene semantic segmentation results through color coding, and use different color codings to distinguish walls, obstacles, channels, etc. It supports multi-view fusion display, can simultaneously display the visual analysis results from different angles, and realizes the visual tracking function to continuously track and identify the trapped persons.

[0143] The alarm module 203 is used to prompt the staff. The alarm module 203 includes a multi-level alarm mechanism, which is divided into emergency alarms, important prompts, and general notifications according to the degree of urgency. When implemented, the alarm module 203 can be a multi-modal alarm method, including sound, vibration, visual flashing, and AR prompts. The alarm rules can be customized by the staff, and the staff can adjust the corresponding alarm thresholds according to actual needs to achieve multi-level alarms.

[0144] The interaction module 202 is used to control and select the content displayed on the visualization display module 201, supporting multiple interaction methods such as touch, voice, gesture, and physical buttons. When the interaction module 202 includes voice interaction, the interaction module 202 supports natural language instructions such as "zoom in on this area" and "mark this point". When the interaction module 202 includes AR glasses, in the AR glasses mode, the interaction module 202 supports gaze tracking and gesture control, and designs quick operations in the emergency mode to ensure fast operation under extreme conditions.

[0145] This disaster rescue system also supports multi-terminal data synchronization. The rescue command center and on-site staff can view the same information, implement the collaborative annotation function, provide real-time communication functions, and design a permission management system where different roles have different operation permissions.

[0146] When implementing the above functions of the disaster rescue system, it is also necessary to integrate and optimize the above functions. Optionally, use WebAssembly technology to optimize the performance of browser-side computationally intensive tasks, implement GPU-accelerated image processing and 3D rendering, ensure smooth display at 60fps, design a low-power mode to automatically reduce the resource consumption of non-critical functions when the battery is low, implement data compression transmission, optimize real-time communication in low-bandwidth environments, and support the offline working mode to keep the core functions available in case of network interruption.

[0147] The performance of the disaster rescue system can be verified according to the set performance evaluation indicators. Exemplarily, the performance evaluation indicators include reference performance indicators such as detection performance (detection distance: ≤30m, positioning accuracy: ±0.5m, recognition rate: ≥95%, false alarm rate: ≤1%), system performance (response time: ≤1s, data processing delay: ≤100ms, running time: ≥8h), and vision system performance (target detection accuracy: ≥90%, pose estimation accuracy: ≤50mm, scene reconstruction accuracy: ≤100mm, vision processing frame rate: ≥30fps).

[0148] It should be noted that when the disaster rescue system provided in the above embodiment conducts disaster rescue, only the above-mentioned division of each functional unit is used for illustration. In actual applications, the above functions can be allocated to different functional units according to needs, that is, the internal structure of the device is divided into different functional units to complete all or part of the functions described above.

[0149] The following are method embodiments of this application. For details not described in detail in the method embodiments, reference can be made to the above system embodiments.

[0150] Figure 3 The flowchart of a disaster rescue method provided by an exemplary embodiment of the present disclosure is shown. Refer to Figure 3 and the method includes: In step 301, based on the disaster rescue system, multi-modal data at the disaster site is collected.

[0151] In step 302, the inference result output by the disaster rescue system is obtained.

[0152] In step 303, based on the inference result, rescue operations are carried out on the personnel to be rescued.

[0153] Among them, the disaster rescue system and the inference result are based on Figure 1 or Figure 2 the obtained disaster rescue system.

[0154] The disaster rescue method provided by the above embodiment and the embodiment of the disaster rescue system belong to the same concept. For the specific implementation process, please refer to the system embodiment and will not be elaborated here.

[0155] The above are only optional embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A disaster relief system, characterized in that: The disaster relief system comprises a multi-source data acquisition unit, a data processing unit and a human-computer interaction unit connected in sequence. The multi-source data acquisition unit is used to collect multi-modal data at the disaster site, and the multi-modal data includes radar data, infrared image data, sound data, gas data, and video data; The data processing unit is used to fuse multiple features of the multimodal data to obtain a fused feature set, construct a first knowledge graph based on the fused feature set, and use a graph neural network GNN to infer the first knowledge graph to obtain an inference result, and the inference result is used to indicate the position of a human body in the disaster site; The human-computer interaction unit is used to display the reasoning result.

2. The disaster relief system according to claim 1, characterized in that: The data processing unit is further configured to implement the fusion of multiple features of the multimodal data to obtain a fused feature set in the following manner: Converting the features of each modality in the multimodal data into a form of evidence to obtain multiple pieces of evidence, wherein the evidence includes probability or confidence; Based on the improved Dempster-Shafer evidence theory, the multiple evidences are fused in the spatial position recognition framework, the vital sign recognition framework, and the environmental state recognition framework to obtain a fused feature set, wherein the fused feature set includes the probability of each element in the spatial position recognition framework, the vital sign recognition framework, and the environmental state recognition framework; Among them, in the improved Dempster-Shafer evidence theory, each evidence has an adaptive weight, and the probability of any element in the spatial position recognition framework, the vital sign recognition framework, and the environmental state recognition framework is determined based on the evidence corresponding to the element and the adaptive weight of the evidence.

3. The disaster relief system according to claim 2, characterized in that: The data processing unit is further configured to adjust the adaptive weight in the following manner: Acquiring environmental data, wherein the environmental data includes light intensity, environmental noise, and temperature; Obtaining a first threshold corresponding to a first modality, where the first modality is data of any modality in the multimodal data; Based on the size relationship between the environmental data and the first threshold, the weight corresponding to the evidence of the first modality is adjusted.

4. The disaster relief system according to claim 2, characterized in that: The data processing unit is further configured to adjust the adaptive weight in the following manner: Acquire a signal-to-noise ratio and a variance of a first modality, where the first modality is data of any one modality in the multimodal data; In the case where the signal-to-noise ratio of the first modality changes, adjusting the weight corresponding to the evidence of the first modality based on the signal-to-noise ratio of the first modality, or, When the variance of the first modality changes, the weight corresponding to the evidence of the first modality is adjusted based on the variance of the first modality.

5. The disaster relief system according to any one of claims 2 to 4, characterized in that: In the modified Dempster-Shafer evidence theory, The first The probability of an element is determined using the following formula: in, Indicates the The first The probability of an element, , Indicates the spatial position, Indicates vital signs, Indicates the state of the environment. For the said The first The element corresponding to Evidence, for The adaptive weight of .

6. The disaster relief system according to any one of claims 1 to 4, characterized in that: The GNN The layer is updated using the following formula: in, is the first Node in layer Parameters, For the node The set of adjacent nodes, is the first Node in layer The parameters of the node for A node in For the node For the node The attention weight, is the first The parameter matrix of the layer, is a non-linear activation function.

7. The disaster relief system according to any one of claims 2 to 4, characterized in that: The first knowledge graph includes entities and relationships between entities, and the entities of the first knowledge graph include vital signs, spatial locations, environmental conditions, and rescue resources; The relationships between entities in the first knowledge graph include causal relationships, spatial relationships, temporal relationships, and reasoning relationships; The constructing a first knowledge graph based on the fused feature set includes: The probability of each element in the fused feature set is coupled with multiple entities of the first knowledge graph to obtain the first knowledge graph.

8. The disaster relief system according to any one of claims 1 to 4, characterized in that: The data processing unit is further used to perform three-dimensional scene modeling on the disaster site based on the radar data and the video data to obtain a first model of the disaster site; The human-computer interaction unit includes a visualization display module, and the visualization display module is used to display the first model after coupling the infrared image data with the first model, and to display the first model after coupling the inference result with the first model.

9. The disaster relief system according to claim 8, characterized in that: The data processing unit is further used to generate a rescue plan based on the inference result, wherein the rescue plan includes at least one rescue path for rescuing a first human body, wherein the first human body is a human body to be rescued at the disaster site; The visual display module is also used to mark the rescue path in the rescue plan on the first model.

10. A disaster relief method, characterized in that: The method comprises: Based on the disaster relief system, collect multimodal data at the disaster site; Obtaining the inference result output by the disaster relief system; Based on the inference result, rescuing the rescuer; Wherein, the disaster relief system and the inference result are obtained based on the disaster relief system described in any one of claims 1 to 9.

Citation Information

Cited By

  • Disaster information emergency processing method and device based on large model, equipment and medium

    CN121707801A

  • Disaster information emergency processing method and device based on large model, equipment and medium

    CN121707801B