Shaft rescue multi-mode converged communication system and method
By employing multimodal data acquisition, preprocessing, feature extraction, attention mechanism weighted fusion, and entropy coding compression, the problems of signal attenuation and insufficient multimodal information fusion in vertical shaft rescue were solved, enabling efficient transmission and processing of multiple types of information and improving rescue efficiency and success rate.
Patent Information
- Application Number
- CN202511866901.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-02-10
AI Technical Summary
In vertical shaft rescue operations, severe signal attenuation, insufficient multimodal information fusion capabilities, and poor anti-interference capabilities lead to unstable communication, making it difficult to achieve comprehensive transmission and processing of multiple types of information, thus affecting rescue efficiency and success rate.
By employing multimodal data acquisition, preprocessing, feature extraction, attention mechanism weighted fusion, entropy coding compression, and encoded transmission, audio, video, and environmental parameter information are integrated to generate rescue analysis results.
It has achieved effective integration and stable transmission of multimodal information, improved rescue efficiency and success rate, and provided reliable rescue decision support.
Smart Images

Figure CN121509530A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of communication technology, and in particular relates to a multimodal fusion communication system and method for vertical shaft rescue. Background Technology
[0002] Shaft rescue, as a special emergency rescue scenario, faces numerous technical challenges. After a shaft accident, trapped personnel are usually located deep underground in a complex and changeable environment with extremely poor communication conditions. Traditional rescue communication methods have many limitations, such as weak signal penetration, limited transmission distance, and insufficient multimodal information fusion capabilities, which severely restrict rescue efficiency and success rates.
[0003] Current vertical shaft rescue communication technologies mainly suffer from the following problems: First, traditional wireless communication methods experience severe signal attenuation in the vertical shaft environment, making it difficult to achieve stable and reliable long-distance communication; second, existing communication systems are mostly single-mode transmissions, which cannot simultaneously meet the comprehensive transmission needs of multiple types of information such as audio, video, and environmental parameters; third, the internal environment of vertical shafts is complex, with numerous interference factors, and traditional communication systems lack sufficient anti-interference capabilities; finally, there is a lack of effective fusion and processing mechanisms for multi-source heterogeneous data, making it difficult to achieve collaborative information analysis and intelligent decision-making.
[0004] During shaft rescue operations, rescuers need to obtain real-time audio and video information from trapped personnel, as well as environmental parameters (such as oxygen levels and toxic gas concentrations). They also need to maintain two-way audio communication with the trapped personnel and provide necessary lighting support. The fusion, transmission, and processing of this multimodal information is crucial for improving rescue efficiency and ensuring rescue safety. However, currently, there is a lack of a shaft rescue communication method that can effectively integrate multimodal information such as audio, video, and environmental parameters, and achieve efficient fusion, transmission, and processing. Summary of the Invention
[0005] Therefore, it is necessary to provide a multimodal fusion communication system and method for shaft rescue to address the aforementioned technical problems.
[0006] Firstly, this application provides a multimodal fusion communication method for shaft rescue, including:
[0007] S1. Collect multimodal data from the vertical shaft rescue site; multimodal data includes audio signals, video signals, and environmental parameter signals;
[0008] S2. Preprocess the multimodal data to obtain preprocessed multimodal data;
[0009] S3. Extract features from the preprocessed multimodal data to obtain multimodal feature data;
[0010] S4. Perform attention-based weighted fusion on the multimodal feature data to generate fused multimodal feature data;
[0011] S5. Perform entropy coding compression on the fused multimodal feature data to obtain compressed multimodal feature data;
[0012] S6. Encode the compressed multimodal feature data to obtain encoded multimodal feature data, and transmit the encoded multimodal feature data to the ground command center. The encoded multimodal feature data is used to instruct the ground command center to analyze the encoded multimodal feature data and generate rescue analysis results. The rescue analysis results include the status assessment of trapped personnel, environmental risk assessment, and rescue strategy recommendations.
[0013] In one embodiment, the fused multimodal feature data is subjected to entropy coding compression to obtain compressed multimodal feature data, including:
[0014] S11. Model the probability distribution of the fused multimodal feature data to obtain the feature probability distribution;
[0015] S12. Calculate the information entropy based on the feature probability distribution to obtain the feature entropy value;
[0016] S13. Perform entropy encoding compression on the fused multimodal feature data based on the feature entropy value to obtain compressed multimodal feature data.
[0017] In one embodiment, the information entropy is calculated based on the feature probability distribution to obtain the feature entropy value, including:
[0018] S21. Discretize the characteristic probability distribution to obtain the discrete probability mass function;
[0019] S22. Calculate the self-information of each feature symbol based on the discrete probability mass function to obtain the set of self-information of the feature symbols;
[0020] S23. By weighted summation of the self-information sets of feature symbols, the feature entropy value is obtained using the following formula:
[0021]
[0022] in, Representing feature data Feature entropy value, This represents the total number of feature symbols after discretization. Indicates the first A characteristic symbol, Indicates the first The probability of each characteristic symbol appearing.
[0023] In one embodiment, the compressed multimodal feature data is encoded to obtain encoded multimodal feature data, and the encoded multimodal feature data is transmitted to the ground command center, including:
[0024] S31. Channel coding is performed on the compressed multimodal feature data, and the encoded data stream is obtained by adding redundant check bits;
[0025] S32. Interleave the encoded data stream by rearranging the data bits to obtain the interleaved data stream.
[0026] S33. Digitally modulate the interleaved data stream to map the digital bit sequence into a complex symbol sequence to obtain the baseband modulated signal;
[0027] S34. Upconvert and amplify the baseband modulation signal to generate a transmission radio frequency signal;
[0028] S35. Transmit the radio frequency signal to the ground command center.
[0029] In one embodiment, multimodal feature data is weighted and fused using an attention mechanism to generate fused multimodal feature data, including:
[0030] S41. Perform linear projection transformation on the feature vectors of each modality in the multimodal feature data to obtain the query vector, key vector and value vector corresponding to each modality;
[0031] S42. Calculate the attention scores between modal features based on the query vector and key vector to obtain the original attention score matrix. The calculation formula is as follows:
[0032]
[0033] in, Indicates the first The query vector of the modality and the first modality The raw attention scores between the key vectors of each modality Indicates the first A query vector of modalities, Indicates the first The key vector of each modality This represents the dimension of the key vector. This is a scaling factor used to prevent the gradient from becoming too small;
[0034] S43. Normalize the original attention score matrix to obtain the attention weight distribution among different modes;
[0035] S44. Based on the attention weight distribution, the value vectors are weighted and summed to obtain the context-aware feature vectors of each modality;
[0036] S45. Concatenate all context-aware feature vectors to generate fused multimodal feature data.
[0037] Secondly, this application also provides a multimodal fusion communication system for shaft rescue, comprising:
[0038] The data acquisition module is used to collect multimodal data from the vertical shaft rescue site; the multimodal data includes audio signals, video signals, and environmental parameter signals.
[0039] The data preprocessing module is used to preprocess multimodal data to obtain preprocessed multimodal data;
[0040] The data feature processing module is used to extract features from the preprocessed multimodal data to obtain multimodal feature data;
[0041] The feature data fusion module is used to perform attention-based weighted fusion of multimodal feature data to generate fused multimodal feature data.
[0042] The fusion data encoding module is used to perform entropy encoding compression processing on the fused multimodal feature data to obtain compressed multimodal feature data;
[0043] The data transmission module is used to encode the compressed multimodal feature data to obtain encoded multimodal feature data, and then transmit the encoded multimodal feature data to the ground command center. The encoded multimodal feature data is used to instruct the ground command center to analyze the encoded multimodal feature data and generate rescue analysis results. The rescue analysis results include the status assessment of trapped personnel, environmental risk assessment, and rescue strategy recommendations.
[0044] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.
[0045] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.
[0046] The aforementioned multimodal fusion communication system and method for shaft rescue first collects multimodal data such as audio signals, video signals, and environmental parameter signals from the shaft rescue site. Then, the collected multimodal data is preprocessed to improve data quality. Subsequently, multimodal feature data is extracted from the preprocessed multimodal data. Next, an attention mechanism is used to weightedly fuse the multimodal feature data to generate more comprehensive fused feature data. Afterward, the fused feature data undergoes entropy coding compression to reduce data redundancy. Finally, the compressed data is encoded and transmitted to the ground command center, where it analyzes the data to generate results such as trapped personnel status assessment, environmental risk assessment, and rescue strategy recommendations. This achieves effective integration and stable transmission of multimodal information, solving problems such as weak signal penetration, insufficient multimodal information fusion, and poor anti-interference capability in traditional methods. It provides reliable support for rescue decision-making and improves rescue efficiency and success rate. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0048] Figure 1 This is a flowchart illustrating a multimodal fusion communication method for shaft rescue in one embodiment;
[0049] Figure 2 This is a schematic diagram of the structure of a multimodal fusion communication system for shaft rescue in one embodiment;
[0050] Figure 3 This is a schematic diagram of a computer device in one embodiment. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0052] refer to Figure 1 The paper presents a flowchart of a multimodal fusion communication method for shaft rescue, which includes the following steps:
[0053] S1. Collect multimodal data from the vertical shaft rescue site.
[0054] Multimodal data includes audio signals, video signals, and environmental parameter signals.
[0055] Specifically, regarding audio signal acquisition, considering the enclosed space inside the shaft, strong echo interference, and the possibility of sudden noise such as falling objects, a distributed noise-resistant audio acquisition unit is adopted. This unit consists of multiple high-sensitivity omnidirectional microphones, using electret condenser structures. The microphones cover the audible frequency range and part of the infrasound frequency range to capture the faint cries for help from trapped personnel and abnormal noises from equipment operation. The microphones have built-in adaptive noise reduction chips, which initially filter out environmental noise through an adaptive filtering algorithm based on minimum mean square error. The installation location of the acquisition units needs to be determined in conjunction with the shaft structure. Typically, multiple units are evenly distributed around the bottom of the rescue pod, while mobile acquisition nodes are suspended above areas where trapped personnel may be active via rescue ropes. The sampling rate and quantization accuracy of all audio acquisition units are uniformly set to ensure the preservation of audio signal details.
[0056] The video signal acquisition utilizes a wide dynamic range, low-light, high-definition camera with a variable focal length lens that supports autofocus to adapt to varying lighting conditions at different depths within the shaft. In well-lit areas near the shaft opening, the camera can clearly capture a wide-area scene by adjusting the focal length. In low-light conditions underground, the camera automatically switches to night vision mode and activates its built-in infrared illumination module to meet real-time requirements. The camera is mounted on a rotatable gimbal at the front of the rescue pod. The gimbal supports horizontal rotation and pitch adjustment, enabling precise tracking of the trapped personnel's location via remote ground control.
[0057] Environmental parameter signal acquisition employs a multi-sensor array, integrating oxygen concentration sensors, toxic gas sensors, temperature sensors, humidity sensors, and air pressure sensors. The oxygen concentration sensor utilizes an electrochemical principle; the toxic gas sensor employs a composite structure combining catalytic combustion and electrochemical methods. The measurement range and accuracy of each gas meet relevant industry standards. The temperature sensor is a platinum resistance type; the humidity sensor is capacitive; and the air pressure sensor is piezoresistive. The sensor array connects to the main control module via a data acquisition card. The sampling frequency setting must balance real-time performance with power consumption. All acquisition units achieve data alignment through timestamp synchronization technology, keeping time synchronization errors within a small range to avoid deviations in subsequent fusion analysis due to data asynchrony.
[0058] S2. Preprocess the multimodal data to obtain preprocessed multimodal data.
[0059] Specifically, targeted preprocessing algorithms are employed to address issues such as noise, redundancy, and outliers in the collected raw multimodal data, improving data quality and laying the foundation for feature extraction. While preprocessing strategies differ for different data types, the processing flow must remain coordinated. For audio signals, preprocessing consists of three core steps: noise suppression, echo cancellation, and signal enhancement. First, based on initial noise reduction, an improved spectral subtraction method is used to further suppress environmental noise. This method first processes the audio signal into frames, applying a Hanning window to each frame to reduce spectral leakage. Then, a fast Fourier transform is performed to convert the time-domain signal to the frequency domain. The noise spectrum during speech-free periods is used as an initial noise template, and noise estimates are updated in real-time during speech periods. A spectral gain function is constructed using the energy difference between speech and noise in the frequency domain, attenuating noise-dominant frequency bands.
[0060] Secondly, for the two-way communication scenario between the rescue pod inside the shaft and the ground, an adaptive echo cancellation algorithm is introduced. This algorithm constructs an adaptive filter to simulate the echo path and updates the filter coefficients using a normalized least mean square algorithm, effectively canceling the audio signals emitted by the rescue equipment itself and the echoes formed by reflections from the shaft wall. Finally, a speech activity detection algorithm distinguishes between speech segments and non-speech segments. For speech segments, gain adjustment based on short-time energy is applied to enhance the strength of weak speech signals, while non-speech segments are muted to reduce data redundancy.
[0061] For video signals, preprocessing focuses on addressing issues such as uneven illumination, image blur, and noise. First, a Retinex-based illumination equalization algorithm is employed to decompose the image into reflection and illumination components. The illumination component is estimated using Gaussian filtering, then logarithmically transformed and stretched before being fused with the reflection component to achieve brightness equalization under different lighting conditions. Second, to address image blur caused by camera shake or air disturbance within the shaft, a blind deconvolution-based image sharpening algorithm is used. A sharp image is recovered by estimating the point spread function (PSF), which employs the maximum a posteriori probability criterion and incorporates prior edge information to improve estimation accuracy. Finally, a hybrid denoising algorithm combining Gaussian filtering and median filtering is used. Gaussian filtering suppresses Gaussian white noise, while median filtering eliminates salt-and-pepper noise, preserving image edge details.
[0062] For environmental parameter signals, preprocessing mainly includes outlier detection and correction, data smoothing, and standardization. Outlier detection uses the Grubbs criterion, classifying detected values exceeding a certain range as outliers. Outliers are corrected using linear interpolation of data from adjacent time points. If outliers occur at multiple consecutive time points, a sensor fault alarm is triggered. Data smoothing employs a moving average filtering algorithm to eliminate fluctuations caused by random noise from the sensor. Standardization uses the Z-score standardization method to convert each environmental parameter into standardized data with a mean of 0 and a standard deviation of 1. The formula is as follows:
[0063]
[0064] In the formula, This represents standardized environmental parameter data. Raw data representing environmental parameters. This represents the mean of all the original data for this environmental parameter. This represents the standard deviation of all raw data for this environmental parameter. This standardization process ensures that environmental parameters with different dimensions can participate in subsequent fusion processing. All preprocessed multimodal data are stored in a unified data format, encapsulated in a specific format, including fields such as data type, collection timestamp, data length, and specific numerical value, facilitating subsequent module calls.
[0065] S3. Perform feature extraction on the preprocessed multimodal data to obtain multimodal feature data.
[0066] Specifically, key features that characterize the essential attributes of the data are extracted from preprocessed audio, video, and environmental parameter data to achieve data dimensionality reduction while retaining effective information relevant to rescue decisions, providing feature support for subsequent fusion processing. For audio signals, feature extraction focuses on the time domain, frequency domain, and cepstral domain features of speech, specifically including short-time energy, short-time zero-crossing rate, Mel-frequency cepstral coefficients, and linearly predicted cepstral coefficients.
[0067] Short-time energy features are obtained by calculating the sum of squares of each frame of audio signal, reflecting the intensity changes of the speech signal, and can be used to determine whether a trapped person is active. Short-time zero-crossing rate is obtained by counting the number of times the time-domain waveform crosses zero level in each frame of audio signal. Combined with short-time energy, it can effectively distinguish speech from non-speech signals, especially exhibiting high robustness in low signal-to-noise ratio environments. Mel frequency cepstral coefficients are the core features of the audio signal, and the extraction process is as follows: First, the audio signal is pre-emphasized to enhance high-frequency components; then, it is framed and windowed, and a fast Fourier transform is performed on each frame to obtain the spectrum. The spectrum is then passed through a Mel triangular filter group, and the output energy of each filter is calculated; the energy value is logarithmically calculated, and then the correlation of each dimension of features is removed by discrete cosine transform. The first few discrete cosine transform coefficients are taken as Mel frequency cepstral coefficient features, and the first and second differences of these coefficients are calculated to form a complete Mel frequency cepstral coefficient feature vector.
[0068] Linear prediction cepstral coefficients are obtained through linear prediction analysis. Based on the linear prediction model of the speech signal, the linear prediction coefficients are solved, and a discrete cosine transform is applied to obtain the linear prediction cepstral coefficient feature vector. This vector is then fused with the Mel frequency cepstral coefficient feature to form an audio feature vector. For video signals, feature extraction is divided into two parts: shallow visual features and deep semantic features. Shallow visual features include color features, texture features, and shape features. Color features use histogram features in the HSV color space, dividing the three channels into several intervals to construct color histogram features. Texture features use gray-level co-occurrence matrix features, calculating gray-level co-occurrence matrices at different distances and angles, and extracting texture parameters such as contrast, energy, entropy, and correlation to form texture features. Shape features use Hu moment features, calculating the invariant moments of the image to describe the contour shape of the trapped person, unaffected by translation, rotation, or scaling.
[0069] Deep semantic features are extracted using a lightweight convolutional neural network. Preprocessed video frames are scaled to a specific size and input into a pre-trained convolutional neural network model. The final fully connected layer is removed, and the output of the global average pooling layer is used as the feature vector. Shallow and deep features are concatenated, and dimensionality reduction is achieved through principal component analysis to obtain the video feature vector. For environmental parameter signals, feature extraction includes statistical and trend features. Statistical features are calculated for standardized data within a certain time window, including mean, variance, maximum, minimum, median, and interquartile range, reflecting the overall state of environmental parameters during that period.
[0070] Trend features are calculated using linear regression to determine the slope and correlation coefficient of data within the window, reflecting the changing trends of environmental parameters. For example, a negative slope and a large absolute value for oxygen concentration indicate a rapid decrease in downhole oxygen. Simultaneously, a threshold exceedance feature is introduced for toxic gas parameters. When the concentration of a toxic gas exceeds a safe threshold, this feature is set to 1; otherwise, it is set to 0, enabling rapid triggering of risk alarms. Statistical features, trend features, and threshold exceedance features are fused to form an environmental parameter feature vector. Finally, the feature vectors from audio, video, and environmental parameters are concatenated to obtain initial multimodal feature data, preparing for subsequent attention-weighted fusion.
[0071] S4. Perform attention-based weighted fusion on the multimodal feature data to generate fused multimodal feature data.
[0072] Specifically, an attention mechanism is employed to achieve adaptive weighted fusion of multimodal features. The core principle is to dynamically allocate weights based on the importance of different modal features in the rescue scenario, highlighting key information, suppressing redundant information, and improving the representational capability of the fused features. First, the initial multimodal feature data obtained from S3 is modally split, yielding audio feature matrices, video feature matrices, and environmental parameter feature matrices. Then, each modal feature matrix is standardized, and L2 normalization is used to adjust the magnitude of each feature vector to 1, eliminating the influence of differences in the magnitude of features from different modalities. A dual attention mechanism combining multi-head self-attention and intermodal attention is adopted, and the specific implementation process is as follows.
[0073] The first step is to construct an intra-modal self-attention module, performing self-attention calculations on the feature matrices of the three modalities respectively. Taking audio features as an example, the audio feature matrix is transformed into a query matrix, a key matrix, and a value matrix through three linear transformation matrices. The self-attention weights are calculated using a scaled dot product attention mechanism, as shown in the following formula:
[0074]
[0075] In the formula, This represents the output of the self-attention calculation. The query matrix is used to calculate the relevance with the key matrix to determine attention weights. The key matrix is used in conjunction with the query matrix to calculate attention weights. The value matrix represents the object on which the attention weights are applied, and the final output is obtained by weighting the weights and the value matrix. This represents the softmax activation function, which normalizes the weights to the range of 0 to 1, ensuring that the sum of all weights is 1. The dimension of the key matrix represents the number of elements in the matrix. and The product result divided by This is to mitigate the vanishing gradient problem caused by dimensionality growth. Through self-attention computation, the correlation information within audio features is enhanced; for example, the weights of consecutive semantic segments in a speech signal are increased. The intramodal self-attention computation process for video and environmental parameter features is consistent with that for audio, except that the dimension of the linear transformation matrix is adjusted according to the feature dimension.
[0076] The second step involves constructing an intermodal attention module. This module concatenates the modal features after intramodal self-attention processing to obtain a joint feature matrix. A linear transformation converts the joint feature matrix into a unified-dimensional modal feature vector. Subsequently, the correlation between each modal feature and the joint feature is calculated as the intermodal attention weights. Specifically, the cosine similarity between each modal feature and the joint feature is first calculated. The cosine similarity values are then softmax normalized to obtain the intermodal weights, with the sum of all weights being 1.
[0077] The third step introduces a dynamic weight adjustment factor for the rescue scenario. Based on real-time rescue needs reported by the ground command center and changes in the underground environment, the weights between modes are adjusted a second time. For example, when the concentration of toxic gases underground exceeds the standard, the weight adjustment factor for environmental parameters is increased, while the weight adjustment factors for audio and video are decreased; when it is necessary to confirm the status of trapped personnel, the weight adjustment factors for video and audio are increased, while the weight adjustment factors for environmental parameters are decreased. The corrected weight calculation formula is as follows:
[0078]
[0079] In the formula, This represents the corrected audio modal weights. This represents the audio modal weights before correction. The weight adjustment factor representing the audio modality. This represents the video modal weights before correction. The weight adjustment factor representing the video modality. Represents the modal weights of the environmental parameters before correction. This represents the weight adjustment factor for the environmental parameter modality. The calculation method for the weights after video and environmental parameter modal correction is the same as that for audio modality; only the corresponding parameters need to be replaced.
[0080] The fourth step involves multiplying the corrected weights by the features processed with intramodal self-attention to obtain weighted modal features. These weighted modal features are then fused element-wise to obtain a fused multimodal feature matrix. This matrix retains the unique information of each modality while highlighting features most relevant to the current rescue scenario through attention weights. For example, when environmental risks are high, the contribution of environmental parameters to the fused feature matrix increases significantly, providing more valuable feature data for subsequent compression and analysis.
[0081] S5. Perform entropy coding compression on the fused multimodal feature data to obtain compressed multimodal feature data.
[0082] Specifically, while ensuring the integrity and validity of the fused feature data, an entropy coding algorithm is used to reduce data redundancy and data transmission volume, adapting to the limited communication bandwidth resources in shaft rescue, while ensuring that the compressed data stream can be fully decoded and recovered by the ground command center. First, the fused multimodal feature matrix undergoes data format conversion, converting floating-point format feature values to fixed-point format using rounding. This conversion error is within the acceptable range for rescue feature analysis, while significantly reducing the number of bits used for data storage.
[0083] Subsequently, the transformed fixed-point feature matrix undergoes differential encoding preprocessing. This involves calculating the difference between adjacent feature samples to reduce data correlation. Specifically, let the first... The feature vector of each sample is , No. The feature vector of each sample is Then the vector after differential encoding is For the first sample The data itself is used as the result of differential encoding. Through differential encoding, the numerical range of the feature data is significantly reduced, laying the foundation for subsequent entropy encoding to improve compression efficiency.
[0084] Arithmetic coding is used as the core compression algorithm. Compared to Huffman coding, arithmetic coding does not require building a code table and is more efficient at compressing low-probability symbols, making it more suitable for data streams with uneven probability distributions, such as feature data. The specific implementation process of arithmetic coding is as follows: First, the matrix after differential coding is normalized by converting all differences into non-negative integers, achieved by adding an offset corresponding to the absolute value of the largest negative value of the fixed-point number. Then, an adaptive probability model is constructed. Without pre-collecting the probability distribution of data, the probability of each symbol is updated in real time during the encoding process. Initially, all symbols have equal probabilities. As encoding progresses, the probability estimate is dynamically adjusted based on the frequency of occurrence of the encoded symbols.
[0085] Encoding is performed using interval partitioning. The entire encoding interval is divided into different sub-intervals based on the probability of the current symbol. The sub-interval corresponding to the symbol to be encoded becomes the new encoding interval. This process is repeated until all symbols are encoded. The final output encoding result is a decimal within the final interval, represented in binary as a compressed bitstream. To further improve compression efficiency, a predictive encoding stage is introduced before arithmetic encoding. A linear prediction model is used to predict the feature data after differential encoding. The feature value of the current sample is predicted based on the differential feature values of the previous few samples. The residual between the predicted value and the actual value is calculated, and arithmetic encoding is performed on the residual. This method can further reduce the dynamic range of the data, making the residual data more concentrated near zero, thus improving the encoding compression ratio.
[0086] After a combined compression process of "fixed-point conversion - differential coding - predictive coding - arithmetic coding", the fused feature data achieves a high compression ratio. Simultaneously, the mean square error between the decoded feature data and the original data is small, ensuring that the ground command center can conduct accurate rescue analysis based on the decoded data. The formula for calculating the mean square error is as follows:
[0087]
[0088] In the formula, Represents the mean square error. The total number of samples representing the feature data. The first representing the original feature data Each sample value The first part represents the decoded feature data. For each sample value, the mean square error is obtained by calculating the sum of squares of the differences between the original and decoded values of all samples and taking the average value. The smaller this value is, the smaller the deviation between the decoded data and the original data.
[0089] S6. Encode the compressed multimodal feature data to obtain encoded multimodal feature data, and transmit the encoded multimodal feature data to the ground command center.
[0090] The ground command center is used to analyze the coded multimodal feature data and generate rescue analysis results; the rescue analysis results include the status assessment of trapped personnel, environmental risk assessment, and rescue strategy recommendations.
[0091] Specifically, encoding ensures the reliability of data transmission and adapts to the complex channel environment of vertical shafts, while ground analysis outputs accurate rescue decision support information based on feature data. The encoding process combines channel coding and source coding. First, cyclic redundancy check (CRC) coding is performed on the compressed bitstream to generate a CRC code. A specific standard is selected, and the check code is obtained by performing modulo-2 division between the compressed bitstream and the generator polynomial. This check code is then appended to the end of the compressed bitstream for use by the ground command center to verify transmission errors after receiving data.
[0092] Subsequently, low-density parity-check codes are used for channel coding. Low-density parity-check codes have coding performance close to the Shannon limit, strong anti-interference ability, and are suitable for complex silo channel environments. The coding rate and code length can be set as needed. Encoding is achieved by constructing a sparse parity-check matrix. The row and column weights of the parity-check matrix must ensure a balance between coding complexity and performance. The bitstream encoded by low-density parity-check codes is the final encoded multimodal feature data, which has strong anti-noise and anti-interference capabilities.
[0093] The transmission process employs a dual-mode redundant transmission architecture combining wireless and wired connections. Wireless transmission utilizes a combination of ultra-low frequency (UHF) radio waves and very high frequency (VHF) relays. UHF radio waves possess strong penetrating power, enabling long-distance communication between the underground and surface environments for transmitting core characteristic data. VHF relays, on the other hand, use relay nodes installed at intervals along the inner wall of the shaft to supplement the transmission of characteristic data with high real-time requirements. The dual-mode transmission employs time-division multiplexing to avoid signal interference and includes a data retransmission mechanism. When the ground command center detects a data error through cyclic redundancy check, it immediately sends a retransmission request to the underground system. The underground transmission module completes the retransmission within a short time upon receiving the request.
[0094] Wired transmission serves as a redundancy backup, achieved through fiber optic cables carried by the rescue pod. Fiber optic transmission offers higher speeds and automatically switches to wired transmission when wireless transmission experiences severe interference, ensuring continuous data transmission. The analysis phase at the ground command center is achieved through a multi-module intelligent analysis system. This system first decodes the received coded data, sequentially performing low-density parity check decoding, cyclic redundancy check, arithmetic decoding, differential decoding, and fixed-point to floating-point conversion, ultimately recovering the fused multimodal feature data.
[0095] The trapped personnel status assessment module then analyzes audio and video features. The audio features use a support vector machine classifier to determine the trapped personnel's voice status, including whether they are conscious, weak, or unconscious. The video features use a target detection algorithm to locate the trapped personnel and combine them with a facial feature extraction module to determine their vital signs, such as whether they have facial movements or eye contact. The overall status assessment results are then output, including normal, requiring emergency medical care, or in critical condition.
[0096] The environmental risk assessment module, based on environmental parameter characteristics, employs the analytic hierarchy process (AHP) to construct a risk assessment model. Parameters such as oxygen content, toxic gas concentration, and temperature are used as assessment indicators, assigned different weights, and a comprehensive risk value is calculated. Based on the risk value, risk levels are categorized, including safe, warning, and hazardous. When the risk level is reached, an audible and visual alarm is immediately triggered. The formula for calculating the comprehensive risk value is as follows:
[0097]
[0098] In the formula, Represents the overall risk value. Represents the number of evaluation indicators. Representing the The weights of each evaluation indicator are determined using the analytic hierarchy process (AHP). Representing the The normalized scores of each evaluation indicator are used, and the sum of the weights of all indicators is 1. The comprehensive risk value is obtained by multiplying the weights of each indicator by their corresponding scores.
[0099] The rescue strategy recommendation module, based on the results of status and risk assessments, combined with a pre-set rescue rule base and case reasoning system, outputs specific rescue strategies. For example, when trapped personnel are in critical condition and the environment is safe, the recommended strategy is to quickly lower a rescue pod and equip it with emergency medical personnel; when the environment is dangerous, the recommended strategy is to first improve the underground environment through remote equipment before implementing the rescue, including methods such as supplying oxygen. All analysis results are presented to command personnel in a visual interface, and structured reports are generated and stored, providing precise decision support for rescue operations.
[0100] In the aforementioned multimodal fusion communication method for shaft rescue, multimodal data such as audio signals, video signals, and environmental parameter signals from the shaft rescue site are first collected. The collected multimodal data is then preprocessed to improve data quality. Key features are then extracted from the preprocessed multimodal data to form multimodal feature data. Next, an attention mechanism is used to weightedly fuse the multimodal feature data to strengthen the synergistic correlation of information from each modality. Afterward, the fused feature data is entropy-encoded and compressed to reduce transmission redundancy. Finally, the compressed data is encoded and transmitted to the ground command center, where it analyzes and generates rescue analysis results including assessments of the trapped personnel's condition, environmental risks, and rescue strategy recommendations. This scheme effectively solves the problems of unstable signal transmission and insufficient multimodal information fusion capabilities in traditional shaft rescue communication, achieving efficient transmission and comprehensive utilization of multiple types of information, providing reliable support for rescue decision-making, and improving rescue efficiency and success rate.
[0101] To further illustrate the solutions of this application, a specific example is provided below. A multimodal fusion communication system for shaft rescue includes front-end data acquisition equipment, data transmission methods, and back-end presentation and interaction methods.
[0102] Front-end data acquisition equipment: Multimodal data includes real-time monitoring data of the underground environment, the location of trapped personnel, and their vital signs. The front-end audio and video acquisition module uses a multi-functional probe, integrating video acquisition, voice intercom, and lighting functions. The voice intercom function supports sound acquisition and playback. In dark environments, this probe can provide supplemental lighting and offer psychological support to trapped personnel, helping to alleviate tension and fear. Simultaneously, the acquired audio and video data needs to undergo noise reduction and encoding processing to ensure the effectiveness of data transmission, laying the foundation for subsequent multimodal data fusion processing and ground-based presentation.
[0103] Data transmission method: The front-end equipment is a multimodal data acquisition device deployed in the vertical shaft, and the back-end is a ground-based human-machine interaction platform. The multimodal data collected by the front end is transmitted to the back-end ground software in real time through wired or wireless transmission via concentric composite rope, ensuring the continuity and timeliness of data transmission and providing data support for real-time monitoring, analysis and interactive operation on the ground.
[0104] Backend presentation and interaction: The ground-based software interface can display key information related to shaft rescue in real time, including the vital signs of trapped personnel (including heart rate, body temperature, and level of consciousness), the depth of trapped personnel from the ground, and real-time environmental parameters underground (including oxygen content, toxic gas concentration, temperature, humidity, and air pressure). When the underground environmental parameters exceed the warning value, an audible and visual alarm will be triggered. Simultaneously, the interface can display real-time video from underground, and ground operators can use the software's pan-tilt-zoom function to control the camera's zoom, pitch, and rotation, and it supports real-time voice communication with those underground. After the multimodal data is transmitted to the ground, the backend software is presented using a two-screen interactive platform: the left side is the audio and video interaction interface, equipped with gimbal control and microphone operation buttons. Clicking to turn on the microphone starts the intercom, and clicking to turn off the microphone switches to mute mode. The gimbal control buttons can be used to rotate, tilt, zoom, and perform other operations on the camera; the right side is the data monitoring interface, which displays real-time information such as underground oxygen content, types and concentrations of toxic gases, and vital signs of trapped personnel. The system uses a built-in algorithm model to conduct a comprehensive risk assessment of various data. When the change in any real-time environmental data in the well exceeds the critical value, an audible and visual alarm will be automatically triggered.
[0105] In an optional embodiment, the fused multimodal feature data is subjected to entropy coding compression to obtain compressed multimodal feature data, including the following steps:
[0106] S11. Model the probability distribution of the fused multimodal feature data to obtain the feature probability distribution.
[0107] Specifically, statistical learning methods are used to mine the probability distribution patterns of the fused feature data, providing a foundational model for subsequent information entropy calculation and entropy coding. Since the residual data of the fused multimodal feature data, after differential coding and predictive coding, typically exhibits specific statistical distribution characteristics rather than a completely random uniform distribution, common distribution types include Laplace distribution, Gaussian mixture distribution, and generalized Gaussian distribution. The modeling process first employs a sliding window mechanism to divide the residual feature data into blocks. The window size is adaptively adjusted according to the sample density of the feature data to ensure that the feature data within each data block has approximately consistent distribution characteristics. For each data block, the maximum likelihood estimation method is used to estimate the distribution parameters. Taking the Laplace distribution model as an example, its probability density function is:
[0108]
[0109] In the formula, Represents residual eigenvalues The probability density, Location parameter (mean) representing the distribution. The representative scale parameter (related to variance) can be calculated using maximum likelihood estimation based on the residual samples within the data block to obtain the parameter that maximizes the probability of the sample's occurrence. and To ensure modeling accuracy, a model selection criterion is introduced. By calculating the Kullback-Leibler divergence (KL divergence) between different candidate distributions (Lappland distribution, Gaussian distribution, etc.) and the data block, the distribution with the smallest KL divergence is selected as the optimal probability distribution model for that data block. The formula for calculating KL divergence is:
[0110]
[0111] In the formula, Represents the true distribution With model distribution The differences between them The empirical probability distribution of the residual data. The probability distribution of the candidate models is represented by the KL divergence value. The smaller the value, the better the model distribution fits the actual data distribution. The optimal distribution models of all data blocks are integrated to form a complete feature probability distribution, providing a basis for subsequent information entropy calculation.
[0112] S12. Calculate the information entropy based on the feature probability distribution to obtain the feature entropy value.
[0113] Specifically, information entropy is a core indicator for measuring data uncertainty. Its value directly reflects the average amount of information contained in the feature data and is also the theoretical upper limit of entropy coding compression efficiency. The calculated feature entropy value will serve as the core basis for adjusting coding parameters. The calculation of information entropy is based on the Shannon entropy definition. For the discretized residual feature data, the formula for calculating information entropy is:
[0114]
[0115] In the formula, Represents the feature entropy value. This represents the number of feature values after discretization. Representing the The probability of each feature taking a value (determined by the feature probability distribution) The base of the logarithm is usually 2, in which case the entropy value is in bits, representing the average number of encoding bits required for each feature value. In actual calculations, the residual feature data is first adaptively discretized. The number and width of the discrete intervals are determined based on the concentration of the feature probability distribution. For feature value ranges with high probability density, finer discrete intervals are set to reduce discretization error; for ranges with low probability density, the interval width is increased to reduce computational complexity. After discretization, the probability of each discrete feature value is calculated using the probability distribution model of each data block. Substituting the values into the Shannon entropy formula yields the local entropy value for each data block. Then, a weighted average is calculated based on the proportion of data volume in each block to obtain the global feature entropy value for the entire fused feature data. The formula for calculating the global feature entropy value is:
[0116]
[0117] In the formula, Represents the global feature entropy value. Represents the number of data blocks. Representing the The proportion of data in each data block to the total data volume Representing the The local entropy value of each data block can be calculated using this weighted method, which can fully consider the differences in information contribution of different data blocks, making the global entropy value more reflective of the information characteristics of the overall data.
[0118] S13. Perform entropy encoding compression on the fused multimodal feature data based on the feature entropy value to obtain compressed multimodal feature data.
[0119] Specifically, guided by feature entropy values, entropy coding parameters are dynamically adjusted to achieve a balance between coding efficiency and data integrity. The core principle is to select the optimal coding mode and parameter configuration based on feature entropy values, ensuring that the average bit rate of actual coding approaches the theoretical limit determined by the feature entropy values. First, basic coding parameters are determined based on global feature entropy values. When the value is small (indicating low data uncertainty and high redundancy), a higher encoding compression ratio is used, and compression efficiency is improved by increasing the interval division precision of arithmetic coding and reducing coding redundancy bits; when When the value is large (indicating high data density and low redundancy), the compression ratio is appropriately reduced to prioritize data transmission integrity and avoid information loss due to over-compression. Simultaneously, a block-based encoding strategy is implemented based on the local entropy values of each data block. More aggressive compression parameters are used for data blocks with low local entropy values (such as stable feature data in environmental parameters), while conservative compression parameters are used for data blocks with high local entropy values (such as audio feature data containing sudden speech changes in trapped individuals), achieving differentiated encoding. During the encoding process, the actual encoding bit rate is compared with the feature entropy value in real time, and the encoding parameters are dynamically corrected through a feedback adjustment mechanism. For example, when the actual bit rate is more than 10% higher than the feature entropy value, the probability model update frequency is automatically optimized to improve encoding efficiency; when the actual bit rate is significantly lower than the feature entropy value, data check bits are added to improve data transmission reliability. Arithmetic coding is used as the core encoding algorithm, with the feature probability distribution model as the initial probability model for arithmetic coding. During the encoding process, the model parameters are updated in real time based on the feedback of the feature entropy value, ensuring that the encoding process always adapts to the distribution characteristics of the feature data. After encoding, the generated compressed bitstream is structured and encapsulated. Metadata information such as the feature entropy value and probability distribution type of each data block is added before the bitstream, facilitating the ground command center to quickly obtain encoding parameters during decoding and improving decoding efficiency. After the above entropy encoding processing based on feature entropy values, the resulting compressed multimodal feature data not only has a high compression ratio but also ensures that the error between the decoded feature data and the original data is controlled within an acceptable range for rescue analysis, effectively adapting to the limited communication bandwidth resources in vertical shaft rescue scenarios.
[0120] In an optional embodiment, the information entropy is calculated based on the feature probability distribution to obtain the feature entropy value, including the following steps:
[0121] S21. Discretize the characteristic probability distribution to obtain the discrete probability mass function.
[0122] Specifically, since the feature probability distribution may include continuous distribution types (such as Laplace distribution and Gaussian distribution), and information entropy calculation requires discrete feature symbols, discretization is a crucial step connecting continuous distributions and entropy calculation. The core principle of discretization is to map continuous feature values to a finite number of discrete feature symbols while preserving the data distribution characteristics, and simultaneously minimizing discretization error. In practice, an adaptive interval partitioning algorithm is first used to determine the partitioning intervals based on the probability density curve of the feature probability distribution: for regions with peak probability density (i.e., regions with concentrated feature values), a smaller interval width is set to preserve the subtle differences in the data; for example, in regions where residual feature values are close to zero (highest probability density), the interval width can be set to 1% of the feature value range. For regions with flat probability density (sparse feature values), a larger interval width is used to reduce the total number of discrete symbols and lower subsequent computational complexity. After interval partitioning, each interval is defined as a discrete feature symbol. Calculate the proportion of feature samples within each interval to the total number of samples, and determine the probability of that feature being represented. The probabilities of all feature symbols constitute the discrete probability mass function, satisfying... ( From 1 to , (This represents the total number of discrete feature symbols). To verify the discretization effect, the KL divergence between the discrete distribution and the original continuous distribution needs to be calculated to ensure that the divergence value is below a preset threshold. If it exceeds the threshold, the interval partitioning parameters are readjusted until the accuracy requirements are met, ensuring the reliability of subsequent entropy calculations.
[0123] S22. Calculate the self-information of each feature symbol based on the discrete probability mass function to obtain the set of self-information of the feature symbols.
[0124] Specifically, self-information is a measure of the information contained in a single feature symbol. Its core physical meaning is: the lower the probability of a feature symbol appearing, the greater the amount of information it carries when it appears, and vice versa. According to Shannon's information theory definition, a single feature symbol... Self-information The calculation formula is:
[0125]
[0126] In the formula, Representing the The self-information of a feature symbol, in bits. The probability of this feature symbol, with a value range of (0,1). When it approaches 1 (the characteristic symbol will almost certainly appear). A value approaching 0 indicates that the symbol carries very little information; when When it approaches 0 (the characteristic symbol rarely appears). Approaching infinity indicates that the symbol provides crucial anomaly information when it appears. For example, in a rescue scenario, the probability of a feature symbol representing a "sudden increase in toxic gas concentration" is extremely low, but its self-information is extremely high, requiring close attention in subsequent processing. During calculation, all feature symbols in the discrete probability mass function are traversed, and the self-information of each symbol is calculated one by one. The results are then arranged according to the feature symbol's index to form a set of self-information values. This provides the basic data for subsequent weighted summation.
[0127] S23. By weighted summation of the self-information sets of feature symbols, the feature entropy value is obtained using the following formula:
[0128]
[0129] in, Representing feature data Feature entropy value, This represents the total number of feature symbols after discretization. Indicates the first A characteristic symbol, Indicates the first The probability of each characteristic symbol appearing.
[0130] In the above formula, Representing feature data Information entropy, measured in bits per symbol, represents the average amount of information carried by each feature symbol. This represents the total number of feature symbols after discretization; Indicates the first One characteristic symbol; Indicates the first The probability of each feature symbol appearing is also the weight of its self-information, ensuring that feature symbols with higher probabilities contribute more to the entropy value, consistent with the logic of statistical averaging. The specific calculation process consists of two steps: First, for each feature symbol, calculate its probability. With self-information The product of, i.e. The first step is to sum the contributions of each symbol to the overall information; the second step is to sum the contributions of all feature symbols to obtain the feature entropy value. Special cases need to be handled during calculation: if the probability of a certain characteristic symbol... If the symbol does not contribute to the entropy value, the calculation can be skipped; if all feature symbols are concentrated in one (n=1), then... ,at this time This indicates that the data has no uncertainty and extremely high redundancy, making it suitable for extreme compression strategies. The feature entropy value calculated through this step quantifies the information density of the feature data and provides a clear efficiency benchmark for entropy coding. Ideally, entropy coding should result in an average number of bits per symbol after encoding close to... This achieves the optimal balance between data compression and information retention.
[0131] In practical calculations, to improve efficiency, a block-based parallel computing strategy can be adopted: the feature data is split into blocks, and steps S21-S23 are performed on each data block to obtain the local entropy value of each data block. Then, based on the proportion of the number of samples in each data block (No. The global feature entropy value is calculated by weighting the number of block samples / total number of samples. The formula is:
[0132]
[0133] In the formula, The number of data blocks represents the number of blocks. This approach not only adapts to the block processing architecture of feature data, but also reflects the information importance of different data blocks through weight allocation. For example, blocks containing the voice features of trapped personnel have higher weights, and their local entropy values have a more significant impact on the global results, ensuring that the global entropy value can accurately reflect the key information characteristics in the rescue scenario.
[0134] In an optional embodiment, the compressed multimodal feature data is encoded to obtain encoded multimodal feature data, and the encoded multimodal feature data is transmitted to the ground command center, including the following steps:
[0135] S31. Channel coding is performed on the compressed multimodal feature data, and the encoded data stream is obtained by adding redundant check bits.
[0136] Specifically, the core purpose of channel coding is to add redundant information to compressed data and build error control capabilities. This allows the ground command center to correct transmission errors through redundant check bits when data is transmitted through complex channels in shafts, improving the reliability of data transmission. Channel coding employs a concatenated coding scheme of Low-Density Parity-Check (LDPC) and Cyclic Redundancy Check (CRC), forming a dual error control mechanism of "outer code + inner code." First, CRC outer code encoding is performed, using the CRC-32 standard. The generator polynomial is an internationally recognized standard polynomial. The compressed multimodal feature data is used as information bits, and modulo-2 division is performed with the generator polynomial to obtain a 32-bit CRC check code. This check code serves as the first layer of redundancy information and is appended to the end of the compressed data for preliminary error detection. Subsequently, LDPC inner code encoding is performed. LDPC codes have coding performance close to the Shannon limit and low decoding complexity, making them suitable for the real-time requirements of shaft rescue. Before encoding, the code length and code rate of the LDPC code need to be determined based on the length of the compressed data. The code length selection must match the block size of the subsequent interleaving processing, and the code rate is set to 1 / 2 to achieve a balance between encoding efficiency and anti-interference capability. The core of the LDPC code is to construct a sparse parity check matrix. The row weight of the matrix is set to a fixed value, and the column weight is determined according to the ratio of the code length to the row weight to ensure the sparsity and randomness of the matrix. The parity check matrix is used to perform a modulo-2 operation with the CRC-encoded data stream to generate LDPC check bits, which are concatenated as a second layer of redundancy after the data. Finally, an encoded data stream containing the original compressed data, CRC check code, and LDPC check code is formed. The addition of redundant check bits enables the data stream to correct random errors and some burst errors.
[0137] S32. The encoded data stream is interleaved by rearranging the data bits to obtain the interleaved data stream.
[0138] Specifically, burst errors occur in shaft channels due to rock reflections, equipment interference, etc. These errors manifest as consecutive bit errors. Simple channel coding has limited ability to correct burst errors. Interleaving, by rearranging the transmission order of data bits, disperses consecutive burst errors into discrete random errors, facilitating correction by channel coding in subsequent decoding stages. A block interleaving method is used. First, the interleaving block size is determined, with its length consistent with the LDPC code length to ensure coordination between interleaving and channel coding. The interleaver internally uses an M x N storage matrix, where the product of M and N equals the interleaving block size. The value of M is determined based on the maximum length of the burst error, typically set to 2-3 times the maximum burst error length to ensure complete dispersion. Data is written row-wise, filling each row of the storage matrix with bits of the encoded data stream until the matrix is full. Data is read column-wise, sequentially reading bits from each column of the matrix and concatenating them to form the interleaved data stream. For example, if the storage matrix is 4 rows and 5 columns, writing data row by row would be "1 2 3 4 5; 6 7 8 9 10; 11 12 13 14 15; 16 17 18 19 20", and reading data column by column would be "1 6 11 16 2 7 12 17 3 8 13 18 4 9 14 19 5 10 15 20". Using this method, if a consecutive error of "6 7 8" occurs during transmission, the error positions are dispersed to the 2nd, 6th, and 10th positions after interleaving, becoming discrete errors that can be effectively corrected using the LDPC code decoding algorithm. After interleaving, interleaving parameter identifiers, including interleaving block size, matrix row and column numbers, need to be added to the data stream header to facilitate corresponding deinterleaving processing at the ground receiver.
[0139] S33. Digitally modulate the interleaved data stream to map the digital bit sequence into a complex symbol sequence to obtain the baseband modulated signal.
[0140] Specifically, the core of digital modulation is to convert discrete digital bit sequences into continuous analog signals, adapting them to the frequency characteristics of the transmission channel for easy transmission via wireless or wired methods. Considering the characteristics of siloed channels, such as multipath fading and strong noise interference, a modulation method with strong anti-fading capability and moderate spectral efficiency needs to be selected. Quadrature Phase Shift Keying (QPSK) modulation is chosen. This modulation method maps two binary bits to a complex symbol, and its spectral efficiency is twice that of Binary Phase Shift Keying (BPSK). It also has strong anti-interference capability and better bit error rate performance than higher-order modulation methods. The modulation process is as follows: First, the interleaved data stream is grouped into pairs of two bits each, resulting in four combinations: "00", "01", "10", and "11". Then, it is mapped according to the QPSK constellation diagram. The four complex symbols in the constellation diagram correspond to the four quadrants of the complex plane; for example, "00" maps to "1+j", "01" maps to "-1+j", "10" maps to "1-j", and "11" maps to "-1-j". The mapping process ensures that the distance between constellation points corresponding to adjacent bit pairs is maximized to improve noise immunity. Finally, the mapped complex symbol sequence is output sequentially to form the baseband modulation signal. This signal's frequency range is concentrated in the low-frequency band, facilitating subsequent up-conversion processing. To adapt to different silo channel conditions, the modulation module supports dynamic switching of modulation methods. It can switch between QPSK and BPSK based on channel quality assessment results. When the channel quality is poor, it switches to the more interference-resistant BPSK; when the channel quality is good, it uses QPSK to improve transmission efficiency.
[0141] S34. Upconvert and amplify the baseband modulation signal to generate a transmission radio frequency signal.
[0142] Specifically, the baseband modulation signal has a low frequency and cannot be directly transmitted over long distances through a wireless channel. The upconversion function is to raise the frequency of the baseband signal to a radio frequency band suitable for vertical shaft transmission, while the power amplification function is to raise the signal power to a level that meets the requirements of the transmission distance, ensuring that the ground command center can receive the signal stably. The upconversion process is implemented in two mixing stages to avoid excessive noise generated by a single mixing stage: The first mixing stage mixes the baseband modulation signal with the intermediate frequency carrier signal to obtain the intermediate frequency modulation signal. The selection of the intermediate frequency needs to consider the device characteristics and anti-interference requirements, and is usually set to tens of MHz. The mixer adopts a balanced mixing structure to suppress even harmonic interference. The second mixing stage mixes the intermediate frequency modulation signal with the radio frequency carrier signal to obtain the radio frequency modulation signal. The selection of the radio frequency band needs to be combined with the silo transmission characteristics. The ultra-low frequency (ULF) band is used for long-distance penetration transmission, with a frequency range of 300Hz-3kHz. The very high frequency (VHF) band is used for relay transmission, with a frequency range of 30MHz-300MHz. Single-band or dual-band transmission can be selected according to the transmission distance and silo structure. After up-conversion, the radio frequency signal needs to undergo power amplification. A linear power amplifier is used to avoid signal quality degradation caused by nonlinear distortion. The amplified signal power needs to be determined based on the transmission distance and channel attenuation characteristics to ensure that the signal power reaching the ground receiver is higher than the receiver sensitivity, while also complying with electromagnetic compatibility standards to avoid interference with rescue equipment. An Automatic Power Control (APC) module is set up in the power amplification stage. By real-time detection of the output signal power and the signal strength fed back from the receiver, the amplifier gain is dynamically adjusted to stabilize the output power at the optimal value. When channel attenuation increases, the gain is increased; when attenuation decreases, the gain is decreased, achieving a balance between energy saving and reliable transmission.
[0143] S35. Transmit the radio frequency signal to the ground command center.
[0144] Specifically, a transmission architecture of "wireless dual-mode redundancy + wired backup" is adopted, taking into account the complex characteristics of the shaft channel to ensure the continuity and reliability of signal transmission. Wireless transmission employs a dual-mode collaboration of Ultra-Low Frequency (ULF) and Very High Frequency (VHF): the ULF transmission link utilizes the strong penetrating power of ULF radio waves to directly penetrate the shaft's rock strata and soil, achieving long-distance direct communication between the underground and the surface without the need for relay equipment. It primarily transmits core multimodal characteristic data, ensuring uninterrupted communication even in extreme conditions. The VHF transmission link is constructed by installing relay nodes at intervals along the shaft's inner wall. These relay nodes use omnidirectional antennas to receive, amplify, and forward VHF radio frequency signals transmitted from underground until reaching the surface. This link offers a higher transmission rate and is used to transmit characteristic data with high real-time requirements, such as the voice and video characteristics of trapped personnel. Both wireless links use frequency division multiplexing to avoid interference. A channel quality monitoring module is also installed to evaluate the bit error rate and signal strength of both links in real time, automatically selecting the link with better quality for data transmission, achieving seamless link switching. Wired transmission serves as a redundant backup link, implemented via armored fiber optic cables carried by the rescue pod. Fiber optic transmission offers advantages such as strong resistance to electromagnetic interference, high transmission speed, and low loss. The up-converted radio frequency signal is converted into an optical signal and transmitted to the ground via fiber optics. In the event of severe interference or interruption of the wireless link, it automatically switches to the wired link to ensure uninterrupted data transmission. A data retransmission mechanism is implemented during transmission. After receiving the signal, the ground command center uses CRC checksum and LDPC decoding to determine the data's correctness. If an error is found, a retransmission request is sent downhole. The downhole transmission module retransmits the corresponding data block within 100ms of receiving the request, ensuring data transmission integrity. After receiving the transmitted radio frequency signal, the ground command center sequentially performs power amplification, down-conversion, digital demodulation, deinterleaving, and channel decoding to recover the compressed multimodal characteristic data, which then proceeds to the analysis phase.
[0145] In an optional embodiment, the multimodal feature data is weighted and fused using an attention mechanism to generate fused multimodal feature data, including the following steps:
[0146] S41. Perform linear projection transformation on the feature vectors of each modality in the multimodal feature data to obtain the query vector, key vector and value vector corresponding to each modality.
[0147] Specifically, because the dimensions of feature vectors from different modalities differ (e.g., audio features and video features have different dimensions), directly performing attention calculations would lead to a dimensionality mismatch. The core function of linear projection transformation is to map the feature vectors of each modality to a unified high-dimensional space, while simultaneously generating three core vectors for attention calculations. In practice, three independent linear projection matrices are constructed for each modality (audio A, video V, and environmental parameters E): [Query projection matrix] Key projection matrix Value projection matrix The output dimension of all projection matrices is set to a uniform value. (This dimension is adaptively adjusted based on feature complexity and must meet the following requirements) (Compatibility with subsequent attention calculations). Using audio modality feature vectors. (dimension is) Taking (e.g., ) as an example, the projection process is as follows: , , ,in For the query vector of the audio modality, For the key vector of the audio modality, The audio modality is represented by a value vector, with all three elements having dimension d. The projection process for the video and environmental parameter modalities is the same as for the audio, ultimately resulting in three sets of vectors: the query vector set. Key vector set Value vector set This provides a dimensionally consistent input for subsequent attention score calculations.
[0148] S42. Calculate the attention scores between modal features based on the query vector and key vector to obtain the original attention score matrix. The calculation formula is as follows:
[0149]
[0150] in, Indicates the first The query vector of the modality and the first modality The raw attention scores between the key vectors of each modality Indicates the first A query vector of modalities, Indicates the first The key vector of each modality This represents the dimension of the key vector. This is a scaling factor used to prevent the gradient from becoming too small.
[0151] Specifically, attention scores are used to quantify the correlation strength between features of different modalities. Higher scores indicate a stronger synergistic effect between the features of the two modalities in rescue decision-making. For example, the audio features of a trapped person's cries for help and the motion features of people in a video typically have high attention scores. The original attention score is calculated using a scaled dot product method, as shown in the formula above. Indicates the first The query vector of the modality and the first modality The raw attention score between the key vectors of the modality is higher when the value is larger. The modality and the first The stronger the feature correlation between modalities; Indicates the first Each modality's query vector is used to actively explore associations with other modalities, such as when using environmental parameter modalities as the query. It will focus on the characteristic dimensions related to environmental risks; Indicates the first Each modality's key vector is used to provide "identification information" of its own features for matching by other modalities; The dimension representing the key vector is consistent with the projection dimension in S41. Consistent; This is a scaling factor, its function is to... When the size is large, reduce it by scaling. To ensure effective updates of attention weights, the numerical range is defined to avoid the vanishing gradient problem during softmax function calculation. During computation, a query matrix is constructed using the query vectors from all modalities. (Dimension is 3×d, where 3 represents the number of modes), construct a key matrix using the key vectors of all modes. (3×d dimension), original attention score matrix (3×3 dimension) Pass After matrix operations, divide element by element. The matrix is obtained. Corresponding to the The modality and the first The original attention score for each modality, for example This represents the strength of the feature association between the audio modality and the video modality.
[0152] S43. Normalize the original attention score matrix to obtain the attention weight distribution among different modes.
[0153] Specifically, the values in the original attention score matrix only reflect relative correlation strength and do not form a probability distribution that can be directly used for weighting. The core of the normalization process is to convert the attention score corresponding to each modality into a weight value between 0 and 1, with the sum of the weights in the same row being 1, ensuring the interpretability and additivity of the weights. The softmax function is used for normalization of the original attention score matrix. The line (corresponding to the first) The normalization formula for all association scores of each modality is:
[0154]
[0155] In the formula, Represents the normalized i-th The modality pair of the first Attention weights for each modality Let be the natural constant, and numerator be the first digit. The modality and the first The indexed result of the original score of the modality, with the denominator being the first... The normalized weights are the sum of the original scores of each modality and all modalities. Exponentialization amplifies the difference between high and low scores, strengthening the weights of key associations; the summation of the denominator ensures that the sum of the weights of the same modality to other modalities is 1, forming a complete probability distribution. For example, if the original scores of the audio modality, video, and environmental parameters are 2, 1, and 0.5 respectively, the weights after normalization will be concentrated on the video modality, reflecting the strong correlation between audio and video features. The normalized weights result in the attention weight matrix. (3×3 dimension), each element in the matrix represents the weight distribution relationship between the corresponding modes.
[0156] S44. Based on the attention weight distribution, the value vectors are weighted and summed to obtain the context-aware feature vectors of each modality.
[0157] Specifically, the value vector carries the core feature information of each modality. Through weighted summation using attention weights, the features of each modality can be integrated with effective information from other related modalities, forming new features that combine their own characteristics with contextual relevance. In the specific calculation, the attention weights of each modality are used as coefficients to linearly weight the value vectors of all modalities. The context-aware feature vector of the i-th modality... For example, the calculation formula is: (in (From 1 to 3, representing all modalities). For example, context-aware features of the audio modality. This vector retains the audio's inherent speech features while incorporating information from related features in the video, such as human actions and environmental gas concentrations, making the features more globally relevant to the rescue scenario. Through this step, each modality generates a vector with a dimension of... Context-aware feature vectors are used to form a vector set. .
[0158] S45. Concatenate all context-aware feature vectors to generate fused multimodal feature data.
[0159] Specifically, the core of the splicing operation is to integrate the context-aware features of various modalities to form a more dimensional and comprehensive fused feature set, providing rich feature support for subsequent compression and analysis. The splicing method employs feature dimension expansion, combining the three dimensions... The context-aware feature vectors are concatenated in sequence to form a dimension of The fused feature vector. If the original projection dimension For a specific value, the dimension of the concatenated and fused feature vector is the corresponding numerical value. This vector contains the core features of the three modalities: audio, video, and environmental parameters, as well as the correlation information between the modalities. For example, through weighted fusion, when the concentration of toxic gas in the environmental parameters increases, the proportion of the correlation information of the environmental parameters in the fused features will automatically increase, providing an accurate basis for risk assessment for the ground command center. For multi-sample scenarios, the fused feature vector of each sample is arranged row by row to form a fused multimodal feature matrix. (The dimension is N×3d, where N is the number of samples), and this matrix is the input data for subsequent entropy encoding compression.
[0160] In the aforementioned multimodal fusion communication method for shaft rescue, multimodal data of audio, video, and environmental parameters from the shaft rescue site are first collected. After targeted preprocessing to improve data quality, features of each modality are extracted. Then, through linear projection transformation, attention score calculation, and normalization, attention mechanism weighted fusion is achieved to obtain comprehensive features. Subsequently, probability distribution modeling, information entropy calculation, and entropy coding compression are performed to reduce data redundancy. Finally, through channel coding, interleaving, digital modulation, up-conversion amplification, and other coding and transmission methods, the data is stably transmitted to the ground command center, which generates rescue analysis results based on this data. This method effectively solves the problems of weak signal penetration, insufficient multimodal information fusion, and poor anti-interference ability of traditional methods, achieving stable and efficient transmission and collaborative analysis of multiple types of information, providing reliable support for rescue decision-making, and improving rescue efficiency and success rate.
[0161] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0162] Based on the same inventive concept, this application also provides a multimodal fusion communication system for shaft rescue to implement the aforementioned multimodal fusion communication method for shaft rescue. The solution provided by this system is similar to the implementation scheme described in the above method. Therefore, the specific limitations of one or more embodiments of the multimodal fusion communication system for shaft rescue provided below can be found in the limitations of the multimodal fusion communication method for shaft rescue described above, and will not be repeated here.
[0163] In one exemplary embodiment, such as Figure 2 As shown, a multimodal fusion communication system 200 for shaft rescue is provided, comprising:
[0164] The data acquisition module 201 is used to acquire multimodal data from the vertical shaft rescue site; the multimodal data includes audio signals, video signals, and environmental parameter signals.
[0165] The data preprocessing module 202 is used to preprocess the multimodal data to obtain preprocessed multimodal data;
[0166] The data feature processing module 203 is used to extract features from the preprocessed multimodal data to obtain multimodal feature data;
[0167] The feature data fusion module 204 is used to perform attention-based weighted fusion of multimodal feature data to generate fused multimodal feature data.
[0168] The fusion data encoding module 205 is used to perform entropy encoding compression processing on the fused multimodal feature data to obtain compressed multimodal feature data;
[0169] The data transmission module 206 is used to encode the compressed multimodal feature data to obtain encoded multimodal feature data, and transmit the encoded multimodal feature data to the ground command center. The encoded multimodal feature data is used to instruct the ground command center to analyze the encoded multimodal feature data and generate rescue analysis results. The rescue analysis results include the status assessment of trapped personnel, environmental risk assessment, and rescue strategy recommendations.
[0170] Furthermore, the integrated data encoding module 205 is also used for:
[0171] S11. Model the probability distribution of the fused multimodal feature data to obtain the feature probability distribution;
[0172] S12. Calculate the information entropy based on the feature probability distribution to obtain the feature entropy value;
[0173] S13. Perform entropy encoding compression on the fused multimodal feature data based on the feature entropy value to obtain compressed multimodal feature data.
[0174] Furthermore, the integrated data encoding module 205 is also used for:
[0175] S21. Discretize the characteristic probability distribution to obtain the discrete probability mass function;
[0176] S22. Calculate the self-information of each feature symbol based on the discrete probability mass function to obtain the set of self-information of the feature symbols;
[0177] S23. By weighted summation of the self-information sets of feature symbols, the feature entropy value is obtained using the following formula:
[0178]
[0179] in, Representing feature data Feature entropy value, This represents the total number of feature symbols after discretization. Indicates the first A characteristic symbol, Indicates the first The probability of each characteristic symbol appearing.
[0180] Furthermore, the data transmission module 206 is also used for:
[0181] S31. Channel coding is performed on the compressed multimodal feature data, and the encoded data stream is obtained by adding redundant check bits;
[0182] S32. Interleave the encoded data stream by rearranging the data bits to obtain the interleaved data stream.
[0183] S33. Digitally modulate the interleaved data stream to map the digital bit sequence into a complex symbol sequence to obtain the baseband modulated signal;
[0184] S34. Upconvert and amplify the baseband modulation signal to generate a transmission radio frequency signal;
[0185] S35. Transmit the radio frequency signal to the ground command center.
[0186] Furthermore, the feature data fusion module 204 is also used for:
[0187] S41. Perform linear projection transformation on the feature vectors of each modality in the multimodal feature data to obtain the query vector, key vector and value vector corresponding to each modality;
[0188] S42. Calculate the attention scores between modal features based on the query vector and key vector to obtain the original attention score matrix. The calculation formula is as follows:
[0189]
[0190] in, Indicates the first The query vector of the modality and the first modality The raw attention scores between the key vectors of each modality Indicates the first A query vector of modalities, Indicates the first The key vector of each modality This represents the dimension of the key vector. This is a scaling factor used to prevent the gradient from becoming too small;
[0191] S43. Normalize the original attention score matrix to obtain the attention weight distribution among different modes;
[0192] S44. Based on the attention weight distribution, the value vectors are weighted and summed to obtain the context-aware feature vectors of each modality;
[0193] S45. Concatenate all context-aware feature vectors to generate fused multimodal feature data.
[0194] In one embodiment, such as Figure 3 A computer device 300 is provided, comprising:
[0195] At least one processor 301, and at least one memory 302 communicatively connected to said processor 301; said memory stores application code executable by said processor, said application code being executed by said processor to enable said processor to perform the steps of a shaft rescue multimodal fusion communication method as described above;
[0196] The computer device may also include: sensor 303;
[0197] The processor 301, memory 302, and sensor 303 can be connected via bus 304 or other means. The figure shows an example of connection via bus 304. Figure 3 The character is represented by a single thick line, but this does not mean that there is only one bus or a type of bus.
[0198] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0199] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0200] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.
Claims
1. A multimodal fusion communication method for vertical shaft rescue, characterized in that, The method includes: S1. Collect multimodal data from the vertical shaft rescue site; the multimodal data includes audio signals, video signals, and environmental parameter signals; S2. Preprocess the multimodal data to obtain preprocessed multimodal data; S3. Perform feature extraction on the preprocessed multimodal data to obtain multimodal feature data; S4. Perform attention-based weighted fusion on the multimodal feature data to generate fused multimodal feature data; S5. Perform entropy coding compression processing on the fused multimodal feature data to obtain compressed multimodal feature data; S6. Encode the compressed multimodal feature data to obtain encoded multimodal feature data, and transmit the encoded multimodal feature data to the ground command center; wherein, the encoded multimodal feature data is used to instruct the ground command center to analyze the encoded multimodal feature data and generate rescue analysis results; the rescue analysis results include the status assessment of trapped personnel, environmental risk assessment, and rescue strategy recommendations.
2. The method according to claim 1, characterized in that, The step of performing entropy coding compression on the fused multimodal feature data to obtain compressed multimodal feature data includes: S11. Perform probability distribution modeling on the fused multimodal feature data to obtain the feature probability distribution; S12. Calculate the information entropy based on the feature probability distribution to obtain the feature entropy value; S13. Perform entropy encoding compression on the fused multimodal feature data according to the feature entropy value to obtain the compressed multimodal feature data.
3. The method according to claim 2, characterized in that, The step of calculating information entropy based on the feature probability distribution to obtain the feature entropy value includes: S21. Discretize the feature probability distribution to obtain a discrete probability mass function; S22. Calculate the self-information of each feature symbol based on the discrete probability mass function to obtain the set of self-information of the feature symbols; S23. By weighted summation of the self-information set of the feature symbols, the feature entropy value is obtained using the following formula: in, Representing feature data Feature entropy value, This represents the total number of feature symbols after discretization. Indicates the first One characteristic symbol, Indicates the first The probability of each characteristic symbol appearing.
4. The method according to claim 1, characterized in that, The process of encoding the compressed multimodal feature data to obtain encoded multimodal feature data, and transmitting the encoded multimodal feature data to the ground command center, includes: S31. Channel coding is performed on the compressed multimodal feature data, and the encoded data stream is obtained by adding redundant check bits; S32. The encoded data stream is interleaved, and the interleaved data stream is obtained by rearranging the data bit order. S33. Digitally modulate the interleaved data stream to map the digital bit sequence into a complex symbol sequence to obtain a baseband modulated signal; S34. Upconvert and amplify the baseband modulation signal to generate a transmission radio frequency signal; S35. Transmit the radio frequency signal to the ground command center.
5. The method according to claim 1, characterized in that, The step of performing attention-based weighted fusion on the multimodal feature data to generate fused multimodal feature data includes: S41. Perform a linear projection transformation on each modality feature vector in the multimodal feature data to obtain the query vector, key vector and value vector corresponding to each modality; S42. Calculate the attention scores between each modality feature based on the query vector and the key vector to obtain the original attention score matrix. The calculation formula is as follows: in, Indicates the first The query vector of the modality and the first modality The raw attention scores between the key vectors of each modality Indicates the first A query vector of modalities, Indicates the first The key vector of each modality This represents the dimension of the key vector. This is a scaling factor used to prevent the gradient from becoming too small; S43. Normalize the original attention score matrix to obtain the attention weight distribution among each modality; S44. Based on the attention weight distribution, the value vector is weighted and summed to obtain the context-aware feature vector of each modality; S45. Concatenate all context-aware feature vectors to generate the fused multimodal feature data.
6. A multimodal fusion communication system for shaft rescue, used to implement the method as described in any one of claims 1 to 5, characterized in that, The system includes: The data acquisition module is used to collect multimodal data from the vertical shaft rescue site; the multimodal data includes audio signals, video signals, and environmental parameter signals. The data preprocessing module is used to preprocess the multimodal data to obtain preprocessed multimodal data; The data feature processing module is used to extract features from the preprocessed multimodal data to obtain multimodal feature data; The feature data fusion module is used to perform attention-based weighted fusion on the multimodal feature data to generate fused multimodal feature data. The fusion data encoding module is used to perform entropy encoding compression processing on the fused multimodal feature data to obtain compressed multimodal feature data; The data transmission module is used to encode the compressed multimodal feature data to obtain encoded multimodal feature data, and transmit the encoded multimodal feature data to the ground command center; wherein, the encoded multimodal feature data is used to instruct the ground command center to analyze the encoded multimodal feature data and generate rescue analysis results; the rescue analysis results include the status assessment of trapped personnel, environmental risk assessment, and rescue strategy recommendations.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 5.