Multi-dimensional transmission data compression sensing method and system for urban emergency rescue
Through the multimodal joint compression perception framework, the redundant storage and transmission of multimodal data in urban emergency rescue is solved, efficient and anti-interference data transmission and semantic consistency are achieved, and the accuracy of command decisions is improved.
Patent Information
- Application Number
- CN202510657244.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-12
AI Technical Summary
The existing technology cannot effectively handle the joint sparseness and correlation of multimodal data in urban emergency rescue scenarios, resulting in redundant storage and transmission, and insufficient anti-interference ability in complex electromagnetic environments, affecting the accuracy of command and decision-making.
The multimodal joint compression perception framework is adopted, and the diagonal random matrix design of sparse representation and blocked diagonal random matrix, combined with cross-modal attention mechanism and error correction coding, realize efficient compression and anti-interference transmission of multimodal data.
It significantly reduces the storage and transmission overhead of text, image and audio data, improves data transmission efficiency and anti-interference ability, ensures the semantic consistency of multimodal data, and improves the accuracy of command decisions.
Smart Images

Figure CN120475181A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data compression and transmission, and in particular to a multi-dimensional transmission data compression perception method and system for urban emergency rescue. Background Art
[0002] With the rapid development of smart cities and the Internet of Things (IoT) technologies, urban fire emergency rescue scenarios are placing higher demands on the real-time collection, compression, and transmission of multimodal data. Fire scenes typically involve multidimensional data such as text alarm information (such as location descriptions and fire severity), high-resolution images / video (such as smoke spread and building structure), and ambient audio (such as cries for help and explosions). These data are highly heterogeneous, multidimensional, and complexly correlated.
[0003] In actual urban emergency rescue scenarios, firefighting terminal equipment (such as drones and wearable sensors) often has limited resources, and traditional data transmission technologies face the following core challenges. First, the amount of raw data for high-definition images and audio is huge, far exceeding the carrying capacity of narrowband emergency communication networks, which can easily lead to a data explosion. Secondly, due to the harsh transmission environment at the fire scene, there are often problems such as electromagnetic interference and signal obstruction, resulting in a high transmission bit error rate. Furthermore, the traditional isolated processing of text, images, and audio leads to semantic fragmentation and a lack of collaborative processing, making it difficult to support accurate joint decision-making.
[0004] In recent years, with the continuous development of data resources and computer technology, compressed sensing (CS) theory, which reduces data dimensionality through sparse representation and random sampling, has been widely used in data processing and efficient transmission. However, existing solutions are primarily designed for a single modality (such as images or audio) and lack a unified compression and collaborative recovery framework for multimodal data. Existing technologies are underappreciated in urban fire emergency scenarios, leaving ample room for improvement. Traditional compression methods (such as JPEG and HEVC) require full sampling before compression, resulting in high computational complexity and low compression efficiency, making them difficult to meet the real-time requirements of fire terminals. Single-modality compression schemes (such as wavelet-based image compression and MFCC audio compression) cannot effectively handle the joint sparsity and correlation of multimodal data, leading to redundant storage and transmission. Conventional error-correcting codes (such as RS codes and convolutional codes) have limited error correction performance in complex electromagnetic environments, easily losing critical information and resulting in insufficient anti-interference capabilities. Existing sparse signal recovery algorithms (such as basis pursuit) consume large computational resources and are difficult to execute quickly on mobile receiving devices. The lack of cross-modal alignment mechanism after multimodal data recovery leads to inconsistency between text description and image / audio content, affecting the accuracy of command decision-making. Summary of the Invention
[0005] In order to solve the problems of fragmentation, inefficiency and lack of reliability in multimodal data processing in existing technologies, this disclosure proposes a multi-dimensional transmission data compression sensing method for urban emergency rescue to solve the above problems.
[0006] According to one aspect of the present disclosure, a method for compressive sensing of multi-dimensional transmission data for urban emergency rescue is provided, comprising:
[0007] S10: Obtain multimodal data, where the multimodal data includes text data, image data, and audio data in an urban emergency rescue scenario;
[0008] S20. Normalizing and sparsely representing the multimodal data to obtain sparse coefficients of each modality, wherein the sparse coefficients of each modality include: text sparse basis coefficients, image wavelet coefficients, and audio Mel-frequency cepstral coefficients;
[0009] S30, constructing block diagonal random matrices for each modal sparse coefficient to perform independent compression sampling, and generating low-dimensional observation values before encoding;
[0010] S40, performing arithmetic coding and low-density parity check code error correction coding on the low-dimensional observation value before encoding to generate an anti-interference code stream;
[0011] S50, performing low-density parity check code decoding and arithmetic decoding on the anti-interference code stream to obtain decoded low-dimensional observation values, and then iterating using an orthogonal matching pursuit algorithm in combination with a block diagonal random matrix to recover the sparse coefficients of each mode;
[0012] S60, reconstructing the recovered sparse coefficients of each modality to obtain text, image and audio features;
[0013] S70. Project text, image, and audio features into a shared semantic space, calculate the joint semantic weights through a cross-modal attention mechanism, and generate semantically consistent multimodal fusion results.
[0014] Preferably, the multimodal data is normalized and sparsely represented to obtain text sparse basis coefficients, image wavelet coefficients, and audio Mel-frequency cepstral coefficients, including:
[0015] Generate text sparse basis coefficients for text data through a bidirectional encoder representation model and a k-sparse autoencoder;
[0016] By performing block normalization and discrete wavelet transform on the image data, the image wavelet coefficients are generated;
[0017] The audio Mel-frequency cepstral coefficients are generated by pre-emphasis, framing and audio feature extraction of audio data.
[0018] Preferably, the image data is subjected to block normalization and discrete wavelet transform to generate image wavelet coefficients, including:
[0019] Divide the image into multiple sub-blocks and perform normalization;
[0020] Perform discrete wavelet transform on each sub-block, and replace the high-frequency sub-band coefficients by threshold processing or neighboring block reference to obtain the image wavelet coefficients.
[0021] Preferably, the audio data is pre-emphasized, framed, and audio features are extracted to generate audio Mel-frequency cepstral coefficients, including:
[0022] Perform pre-emphasis, frame division, and Hamming window processing on the audio signal to extract the Mel-frequency cepstral coefficients and their dynamic characteristics;
[0023] Construct a block diagonal random matrix and independently compress and sample the sub-vectors of the Mel-frequency cepstral coefficients of each frame of audio;
[0024] In view of the sparsity of high-order audio Mel-frequency cepstral coefficients, the dimensionality is further reduced by combining part of the Fourier measurement matrix, and the sparse coefficients are iteratively solved through the orthogonal matching pursuit algorithm to obtain the audio Mel-frequency cepstral coefficients.
[0025] Preferably, performing arithmetic coding and low-density parity check code error correction coding on the low-dimensional observation value before encoding to generate an anti-interference code stream includes:
[0026] Perform arithmetic coding on the low-dimensional observations before encoding to generate a compact binary bit stream;
[0027] Construct the system generation matrix and the corresponding check matrix;
[0028] Inputting the bit stream into a system generation matrix and generating a codeword by matrix multiplication;
[0029] Error detection is performed using a check matrix, and error correction is iteratively performed through a belief propagation algorithm to generate an interference-resistant code stream.
[0030] Preferably, reconstructing the recovered sparse coefficients of each modality includes:
[0031] The sparse basis coefficients of the text are converted into a word vector sequence, and the word vector sequence is repaired with contextual semantics through the self-attention layer of the bidirectional encoder representation model to obtain the completed text features;
[0032] After inverse transforming the image wavelet coefficients, high-resolution image features are obtained through a super-resolution model based on a convolutional neural network;
[0033] The audio Mel-frequency cepstral coefficients are denoised by spectral subtraction and then the audio features are generated by inverse short-time Fourier transform.
[0034] Preferably, the text, image and audio features are projected into a shared semantic space, represented as:
[0035] T′=W t T,
[0036] V′=W v V,
[0037] A′=W a A,
[0038] Where T, V, and A represent text, image, and audio feature vectors respectively, and W t 、W v 、W a is a learnable matrix.
[0039] According to one aspect of the present disclosure, a multi-dimensional transmission data compression sensing system for urban emergency rescue is provided, comprising:
[0040] A multimodal data acquisition module, which obtains multimodal data, wherein the multimodal data is text data, image data, and audio data in an urban emergency rescue scenario;
[0041] Each modality sparse coefficient acquisition module performs normalization processing and sparse representation on the multimodal data to obtain each modality sparse coefficient, wherein the each modality sparse coefficient includes: text sparse basis coefficient, image wavelet coefficient and audio Mel frequency cepstral coefficient;
[0042] A low-dimensional observation value generation module before encoding is used to construct a block diagonal random matrix for each modal sparse coefficient, perform independent compression sampling, and generate a low-dimensional observation value before encoding;
[0043] The anti-interference code stream generation module performs arithmetic coding and low-density parity check error correction coding on the low-dimensional observation values before encoding to generate an anti-interference code stream;
[0044] Each modal sparse coefficient recovery module performs low-density parity check code decoding and arithmetic decoding on the anti-interference code stream, obtains the decoded low-dimensional observation value, and then uses the orthogonal matching pursuit algorithm in combination with the block diagonal random matrix to iterate and recover the sparse coefficients of each modal;
[0045] Each modal sparse coefficient reconstruction module reconstructs the recovered sparse coefficients of each modality to obtain text, image and audio features;
[0046] The multimodal fusion result generation module projects text, image and audio features into a shared semantic space, calculates the joint semantic weights through a cross-modal attention mechanism, and generates semantically consistent multimodal fusion results.
[0047] According to one aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the above-mentioned multi-dimensional transmission data compression sensing method for urban emergency rescue.
[0048] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the multi-dimensional transmission data compression sensing method for urban emergency rescue is implemented.
[0049] Compared with the prior art, the beneficial effects of the present disclosure are:
[0050] 1) This disclosure significantly reduces the storage and transmission overhead of text, image, and audio data through a multimodal joint compressed sensing framework, solves the problem of high redundancy in traditional single-modality compression, and improves data transmission efficiency in fire emergency scenarios.
[0051] 2) This paper employs a block-diagonal random matrix design with structured decomposition to reduce the storage complexity of the global matrix, improving the feasibility of deployment on resource-constrained terminal devices. Combining a k-sparse autoencoder with a high-frequency wavelet coefficient substitution strategy further enhances the sparse representation, improving reconstruction accuracy while maintaining the compression ratio.
[0052] 3) This disclosure addresses complex electromagnetic environments. The combined LDPC-arithmetic coding mechanism effectively suppresses transmission errors through layered redundancy elimination and error correction protection, significantly improving anti-interference capabilities and communication reliability. Furthermore, the cross-modal attention mechanism and dynamic time warping (DTW) technology achieve precise alignment of multimodal semantics, resolving the semantic gaps caused by data isolation in traditional solutions. This allows for the coordination and consistency of text descriptions, image details, and audio timing, significantly improving the accuracy of command decisions.
[0053] 4) This paper adopts the CNN super-resolution model and BERT semantic completion technology to quickly restore high-resolution images and complete semantics, overcoming the difficulties of detail loss and semantic ambiguity in traditional reconstruction algorithms. It combines the orthogonal matching pursuit algorithm to optimize iteration efficiency, enabling mobile terminals to complete data recovery in real time.
[0054] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure.
[0055] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.
[0057] Figure 1 A flow chart of a multi-dimensional transmission data compression sensing method for urban emergency rescue is shown;
[0058] Figure 2 The overall technical process flow chart of the multi-dimensional transmission data compression sensing method for urban emergency rescue in the disclosed example is shown;
[0059] Figure 3 The block diagram of the structure of the multi-dimensional transmission data compression sensing system for urban emergency rescue in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0060] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.
[0061] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0062] The term "and / or" herein simply describes an association relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent the existence of three situations: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" herein refers to any combination of at least two of any one or more of a plurality of items. For example, "at least one of A, B, and C" can represent any one or more elements selected from the set consisting of A, B, and C.
[0063] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.
[0064] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the embodiments of the present disclosure.
[0065] Example 1
[0066] Based on the above ideas, the embodiments of the present disclosure propose a method for compressive sensing of multi-dimensional transmission data for urban emergency rescue. Figure 1 A flow chart showing a method for compressive sensing of multi-dimensional transmission data for urban emergency rescue. The method comprises:
[0067] S10: Obtain multimodal data, where the multimodal data includes text data, image data, and audio data in an urban emergency rescue scenario;
[0068] S20. Normalizing and sparsely representing the multimodal data to obtain sparse coefficients of each modality, wherein the sparse coefficients of each modality include: text sparse basis coefficients, image wavelet coefficients, and audio Mel-frequency cepstral coefficients;
[0069] S30, constructing block diagonal random matrices for each modal sparse coefficient to perform independent compression sampling, and generating low-dimensional observation values before encoding;
[0070] S40, performing arithmetic coding and low-density parity check code error correction coding on the low-dimensional observation value before encoding to generate an anti-interference code stream;
[0071] S50, performing low-density parity check code decoding and arithmetic decoding on the anti-interference code stream to obtain decoded low-dimensional observation values, and then iterating using an orthogonal matching pursuit algorithm in combination with a block diagonal random matrix to recover the sparse coefficients of each mode;
[0072] S60, reconstructing the recovered sparse coefficients of each modality to obtain text, image and audio features;
[0073] S70. Project text, image, and audio features into a shared semantic space, calculate the joint semantic weights through a cross-modal attention mechanism, and generate semantically consistent multimodal fusion results.
[0074] The embodiment of the present disclosure provides a multi-dimensional transmission data compression sensing method for urban emergency rescue. The overall technical route flow chart is as follows: Figure 2 As shown, it includes the following contents.
[0075] Data collection and preprocessing: Collect text, image, and audio data from urban firefighting scenarios to construct a multimodal dataset.
[0076] Sparse representation: Text data is processed through the BERT model to generate word vectors, and sparse basis coefficients are extracted through the k-sparse autoencoder. Image data is divided into blocks and the sparsity is enhanced through discrete wavelet transform (DWT) and high-frequency coefficient substitution. Audio data is pre-emphasized, framed and windowed to extract Mel-frequency cepstral coefficients (MFCC).
[0077] Compressed sampling: A block diagonal random matrix (satisfying the RIP condition) is used to independently measure the sparse coefficients of each mode to generate low-dimensional observations.
[0078] Coding and transmission: The observation values are arithmetically coded to eliminate redundancy, and then LDPC error correction coding is used to generate an interference-resistant code stream for transmission.
[0079] Reconstruction at the receiving end: LDPC decoding and arithmetic decoding are performed on the code stream to restore the low-dimensional observation value; the orthogonal matching pursuit (OMP) algorithm is used to iteratively reconstruct the sparse coefficients of each mode.
[0080] Multimodal restoration and alignment: Text sparse coefficients are semantically completed through the BERT model; image wavelet coefficients are inversely transformed and the CNN super-resolution model is used to restore details; audio MFCC coefficients are denoised through spectral subtraction and inversely transformed into time domain signals; the cross-modal attention mechanism and dynamic time warping (DTW) are used to align multimodal features and output semantically consistent rescue decision information.
[0081] The specific steps include:
[0082] S10: Obtain multimodal data, where the multimodal data includes text data, image data, and audio data in an urban emergency rescue scenario.
[0083] S20. Perform normalization and sparse representation on the multimodal data to obtain sparse coefficients of each modality, where the sparse coefficients of each modality include: text sparse basis coefficients, image wavelet coefficients, and audio Mel-frequency cepstral coefficients.
[0084] In this embodiment, the multimodal data is normalized and sparsely represented to obtain text sparse basis coefficients, image wavelet coefficients, and audio Mel-frequency cepstral coefficients, including:
[0085] S201. For text data, the BERT model is used to generate semantically compact word vectors, and feature sparsification is performed through the k-sparse autoencoder to retain key activation values and filter out redundant information.
[0086] S202 , after performing block normalization on the image data, the frequency domain features are extracted using discrete wavelet transform, and the inter-block redundancy compression performance is enhanced through a high-frequency sub-band coefficient substitution strategy.
[0087] Image wavelet coefficients are generated from image data through block normalization and discrete wavelet transform, including: dividing the image into multiple sub-blocks and performing normalization processing; performing discrete wavelet transform on each sub-block, replacing the high-frequency sub-band coefficients through threshold processing or neighboring block reference to obtain image wavelet coefficients.
[0088] S203 : For the audio data, extract Mel-Frequency Cepstral Coefficients (MFCC) to represent frequency domain features through pre-emphasis, frame division and windowing processing.
[0089] Audio Mel-frequency cepstral coefficients are generated by pre-emphasis, framing and audio feature extraction of audio data, including: pre-emphasis, framing and Hamming window processing of audio signals to extract Mel-frequency cepstral coefficients and their dynamic characteristics; constructing a block diagonal random matrix to independently compress and sample the sub-vectors of the audio Mel-frequency cepstral coefficients of each frame; in view of the sparsity of high-order audio Mel-frequency cepstral coefficients, further dimensionality reduction is carried out in combination with partial Fourier measurement matrices, and the sparse coefficients are iteratively solved through the orthogonal matching pursuit algorithm to obtain the audio Mel-frequency cepstral coefficients.
[0090] S30, constructing block diagonal random matrices for each modal sparse coefficient to perform independent compression sampling, and generating low-dimensional observation values before encoding.
[0091] In this embodiment, block diagonal random matrices are constructed for the sparse features of different modalities for independent compressive sampling. Each sub-matrix element obeys a Gaussian or sub-Gaussian distribution and satisfies the constrained isometry condition (RIP). Low-dimensional observation values are generated through linear projection, which significantly reduces storage and computational complexity.
[0092] S40, performing arithmetic coding and low-density parity check code error correction coding on the low-dimensional observation value before encoding to generate an anti-interference code stream.
[0093] In this embodiment, arithmetic coding is used on the compressed observation values to eliminate statistical redundancy, and then combined with LDPC error correction coding to enhance anti-interference capabilities, generating a robust code stream for transmission. The method involves performing arithmetic coding and low-density parity check error correction coding on the low-dimensional observation values before encoding to generate an anti-interference code stream, including: performing arithmetic coding on the low-dimensional observation values before encoding to generate a compact binary bit stream; constructing a system generator matrix and a corresponding check matrix; inputting the bit stream into the system generator matrix and generating codewords through matrix multiplication; performing error detection using the check matrix and iteratively correcting errors through a belief propagation algorithm to generate an anti-interference code stream.
[0094] For the text data acquired and stored in the firefighting environment, a compressed sensing method is used to achieve efficient and reliable transmission, including the following steps:
[0095] Step 1: Use the BERT model to preprocess the text data and convert the text into BERT word vectors, which are lower-dimensional and more compact than traditional text processing methods.
[0096] Step 2: Based on the generated BERT word vector as the original signal x, a sparse autoencoder is used for sparse representation to obtain the sparse basis coefficient s in the sparse domain.
[0097] Specifically, after the text data is converted into low-dimensional and compact word vectors through the BERT model, a k-sparse autoencoder is used to perform feature sparsification. The autoencoder generates hidden layer activation values through forward propagation, which are expressed as:
[0098] z=W T x+b,
[0099] Where z is the activation vector of the hidden layer of the autoencoder, W is the weight matrix, b is the bias term used to adjust the baseline offset of the activation value, and x is the compact word vector generated by the BERT model.
[0100] Only the top k activation values with the largest absolute values are retained (the rest are forced to zero) to achieve hard threshold sparse representation of text features, specifically:
[0101]
[0102] Where Γ = supp k (z), supp k (z) function returns the index set of the first k activation values with the largest absolute value in z, (Γ) c is the complement of Γ, i.e., all other indices other than the first k maximum activation values.
[0103] This process optimizes weights through back-propagation to ensure that the sparse features can retain key semantic information while meeting the requirements of compressed sensing for signal sparsity, significantly reducing the computational complexity of subsequent processing.
[0104] Step 3: Use the diagonal block matrix as the observation matrix and linearly project the high-dimensional sparse signal into the low-dimensional observation value y output.
[0105] Specifically, a block diagonal random matrix is designed for compressive sampling based on the sparse basis coefficients of the text. Using the block diagonal random matrix as the measurement matrix, to ensure that the signal can be reliably recovered from the observation value y = Φx, the measurement matrix Φ must satisfy the constrained isometry condition (RIP). This condition is equivalent to the observation matrix and the sparse matrix being uncorrelated. This equivalence condition also derives that independent and identically distributed Gaussian random measurement matrices can become universal compressed sensing measurement matrices. Therefore, the element distribution of each diagonal submatrix needs to be close to a random Gaussian matrix or sub-Gaussian matrix to ensure good sparse signal recovery performance. Using the measurement matrix Φ to linearly transform the signal x, the observation value y can be obtained.
[0106] Step 4: Based on the observed values, arithmetic coding is used to reduce redundancy, and then LDPC error correction code is used to increase anti-interference capability.
[0107] Specifically, arithmetic coding is performed on the low-dimensional observations after compressive sampling. By statistically quantizing the probability distribution of the symbols, the floating-point observations are mapped into a sequence of discrete symbols. The encoding process iteratively shrinks the cumulative probability interval of the symbols, ultimately outputting a compact bitstream that approximates the source entropy. This step effectively eliminates statistical redundancy in the observations, resulting in a significantly higher compression rate than traditional fixed-length coding. It is particularly suitable for compressing repetitive semantic features in text data.
[0108] After arithmetic coding, LDPC error correction code is used to protect the compressed bit stream. The information bits are expanded into code words through the system generation matrix, which is expressed as:
[0109] c=s×G,
[0110] Where G is the generator matrix, s is the information bit, and c is the codeword.
[0111] Ensure that the check matrix and generator matrix meet the following constraints:
[0112] GH T =0,
[0113] Where H is the check matrix.
[0114] The receiver uses a sparse check matrix and a belief propagation algorithm for efficient error correction. By iteratively updating the probability information of variable and check nodes, the original code stream, which has been disturbed, is restored. The strong error correction capability of LDPC codes ensures the reliable transmission of text data in complex firefighting channel environments, preventing the loss of critical semantic information.
[0115] Step 5: The codeword obtained after arithmetic coding and LDPC coding can be connected to the channel for transmission.
[0116] For the image data acquired and stored in the firefighting environment, a compressed sensing method is used to achieve efficient and reliable transmission, including the following steps:
[0117] Step 1: Image preprocessing and block normalization.
[0118] Specifically, the input image is divided into B×B sub-blocks. Let the input image be a two-dimensional matrix, expressed as:
[0119]
[0120] Where H and W are the height and width of the image respectively, and the block formula is:
[0121]
[0122] Where, I ij is the i-th row and j-th column sub-block of the original image, with a size of B×B.
[0123] Each sub-block is normalized by subtracting the sub-block mean and dividing by the standard deviation to eliminate illumination differences and local contrast fluctuations and improve the stability of sparse representation. The formula is:
[0124]
[0125] μ ij =mean(I ij ),
[0126] σ ij =std(I ij ),
[0127] The normalized image blocks retain spatial structural information while reducing computational noise interference in subsequent wavelet transforms, providing standardized input for sparsification and compression.
[0128] Step 2: Use wavelet transform and wavelet coefficient substitution to perform sparse representation on the image.
[0129] Specifically, as a two-dimensional signal, the image has strong correlation in the spatial domain. Wavelet transform is a multi-scale analysis tool that can have good localization characteristics in both time and frequency domains. The discrete wavelet transform (DWT) is used to perform sparse transformation on the normalized block. The formula is:
[0130] S ij =ΨI′ ij ,
[0131] Where, is an orthogonal wavelet basis (such as Haar, Daubechies), is a sparse coefficient vector.
[0132] The image block can be further decomposed into four sub-bands under DWT: low-frequency sub-band (LL) and high-frequency sub-band (HL, LH, HH). The low-frequency sub-band LL retains the main structure of the image, and the high-frequency sub-bands HL, LH, HH contain edge and texture information.
[0133] In order to further enhance the sparsity and optimize the subsequent compression effect, some wavelet coefficients in the high-frequency subband can be Replaced by the reference coefficient or template coefficient of the neighboring block, expressed as:
[0134]
[0135] Where, Represents the high-frequency coefficients of the current block, It represents the corresponding coefficient of the reference block (such as the adjacent block), and α∈[0, 1] is the weight coefficient.
[0136] This wavelet coefficient replacement strategy improves the inter-block redundancy compression performance and improves the overall sparsity.
[0137] Step 3: Construct a block diagonal random measurement matrix and perform independent measurement and compression sampling;
[0138] Specifically, for the image wavelet coefficients, a block diagonal random measurement matrix is designed, and each sub-block corresponds to an independent sub-matrix Φ k , its structure is decomposed into:
[0139] Φ k =P k H k D k ,
[0140] Where H k is the Hadamard submatrix, D k is a random diagonal matrix, P k is the row permutation matrix. This design reduces the storage complexity of the global random matrix through local structured operations while satisfying the constrained isometry property (RIP) to ensure high-probability recovery of sparse signals. Block-independent measurement further adapts the image block processing flow and reduces hardware resource consumption.
[0141] The compressed sampling formula is:
[0142]
[0143] Where, φ ij is a submatrix in the block diagonal random measurement matrix, corresponding to the independent measurement matrix of the i-th row and j-th column image block, is the sparse coefficient vector after wavelet transformation. The purpose is to use the block diagonal random matrix φ ij For sparse signals Perform linear projection to compress high-dimensional sparse signals into low-dimensional observation vectors.
[0144] Step 4: Perform arithmetic coding and LDPC error correction on the measurement vector.
[0145] Specifically, for the compressed low-dimensional observation value y ij Vector quantization is performed to discretize continuous values into symbol sequences, and entropy coding is performed through arithmetic coding.
[0146] Assume that the symbol set is Σ={s1,s2,...,s n}, the corresponding probability is P = {p1,p2,...,p n}, define the symbol interval as:
[0147]
[0148] The iterative update interval of the encoding process is:
[0149]
[0150] The final encoding result is the interval [L n ,H n ). This encoding method is close to the source entropy, and the average code length is approximately:
[0151]
[0152] To enhance robustness, low-density parity-check (LDPC) codes are further employed for error correction coding of the bitstream. Redundant check bits are added to the bitstream to improve resilience against channel interference. The receiver uses an iterative belief propagation algorithm to correct errors, ensuring the integrity of image data in harsh transmission environments and preventing the loss of critical details (such as flame areas and escape signs) due to bit errors.
[0153] For the audio data acquired and stored in the firefighting environment, a compressed sensing method is used to achieve efficient and reliable transmission, including the following steps:
[0154] Step 1: Apply a pre-emphasis filter to the original audio signal. This filter enhances the high-frequency components through a first-order high-pass filter to compensate for high-frequency attenuation during signal transmission. The filter coefficient is usually set to an empirical value (e.g., α = 0.97) to balance high-frequency enhancement with signal fidelity.
[0155] Step 2: Divide the pre-emphasized audio signal into frames and add windows.
[0156] Specifically, the pre-emphasized audio signal is divided into short time frames (e.g., 25 milliseconds / frame), with an overlap region (e.g., 10 milliseconds) between each frame to capture the short-term stationary characteristics of the signal. A Hamming window function is applied to each frame signal to reduce spectral leakage. The formula is:
[0157]
[0158] Where L is the frame length, n is the sampling point, and represents the sampling point position within the window function.
[0159] Each frame is multiplied by the Hamming window function to obtain:
[0160] s i (n) = x i (n)w(n),
[0161] Where x i (n) is the sampling value of the original i-th frame signal at sampling point n.
[0162] Step 3: Perform a short-time Fourier transform (STFT) on each frame of the windowed signal to convert the time-domain signal into a frequency-domain spectrum. Calculate the power spectrum, which is the square of the spectrum amplitude. This spectrum reflects the energy distribution of each frequency component and provides a basis for subsequent feature extraction.
[0163] Specifically, an N-point FFT is performed on each frame signal after frame division and windowing to calculate the spectrum, also known as short-time Fourier transform (STFT), where N is usually 256 or 512, and the transformation formula is:
[0164]
[0165] The power spectrum is the spectrum line energy of the speech signal obtained by square modulo the spectrum of the speech signal. The formula is:
[0166] P=|S i (k)| 2 ,
[0167] Step 4: Apply a Mel-scaled triangular filter bank in the frequency domain. This performs a weighted integration of the power spectrum using multiple overlapping triangular filters (typically 22-40) to simulate the human hearing. The natural logarithm of the filter output is taken to enhance the discrimination of low-energy frequency bands, resulting in a logarithmic Mel-scale spectrum.
[0168] Specifically, a filter bank with M triangular filters is defined (the number of filters is similar to the number of critical bands), where M is usually 22 to 40, and 26 is a common standard. Each filter in the filter bank is triangular, with a center frequency of f(m). The response at the center frequency is 1 and decreases linearly toward 0 until it reaches the center frequency of two adjacent filters, where the response is 0. The interval between each f(m) widens as the value of m increases. The core intermediate result Y(m) output by the Mel filter bank is calculated as follows:
[0169]
[0170] Where, PB m (k) represents the frequency response of the mth Mel filter at the discrete frequency point k, that is, the weight, f m-1 and f m+1 Represents the index of the starting and ending frequency points of the mth Mel filter, B m (k) is the frequency response of the Mel filter bank, expressed as:
[0171]
[0172] Where f(m) is the center frequency of the mth Mel filter. Its value is distributed nonlinearly according to the Mel scale, with smaller intervals in low-frequency bands and larger intervals in high-frequency bands. f(m-1) and f(m+1) are the center frequencies of the left and right adjacent filters of the mth Mel filter, respectively, and are used to define the frequency coverage of the current filter.
[0173] Step 5: Perform discrete cosine transform (DCT) on the logarithmic Mel spectrum to remove the correlation between frequency bands.
[0174] Specifically, the above Y(m) is an even function, and the filter bank coefficients are decorrelated using the inverse discrete cosine transform (IDCT), and the formula is:
[0175]
[0176] Where T is the MFCC coefficient order.
[0177] Step 6: Divide the Mel-frequency cepstral coefficient (MFCC) a into sub-vectors, construct a block diagonal random matrix, and use an independent Gaussian or sub-Gaussian random matrix for linear projection on each sub-block.
[0178] Specifically, the MFCC coefficients are divided into sub-vectors, and a block diagonal random matrix is constructed. Each sub-block is linearly projected using an independent Gaussian or sub-Gaussian random matrix, which can significantly reduce the storage and computational complexity while meeting the RIP conditions of compressed sensing.
[0179] Step 7: Fourier compressed sensing projection.
[0180] Specifically, because the MFCC coefficients processed with a block-diagonal random matrix exhibit "energy concentration," high-order coefficients approach zero, meeting the sparsity requirement of compressed sensing. Combining a partial Fourier measurement matrix, the block-compressed MFCC coefficients are reprojected, leveraging frequency-domain sparsity to further reduce the data dimension and generate low-dimensional observations suitable for resource-constrained transmission scenarios.
[0181] Step 8: Quantize the compressed observations and apply LDPC error correction coding.
[0182] S50, performing low-density parity check code decoding and arithmetic decoding on the anti-interference code stream to obtain decoded low-dimensional observation values, and then iterating using an orthogonal matching pursuit algorithm in combination with a block diagonal random matrix to restore the sparse coefficients of each mode.
[0183] In this embodiment, after the receiving end performs LDPC error correction decoding and arithmetic decoding on the code stream, the orthogonal matching pursuit algorithm is used to iteratively restore the sparse features of each modality, and fast reconstruction is achieved in combination with a block diagonal matrix.
[0184] Specifically, the compressed data needs to be decoded and reconstructed at the transmission and receiving end to recover the sparse signal from a small amount of linear projection data.
[0185] After receiving the information transmitted through the channel, the receiver performs LDPC error correction decoding on the transmitted code stream, using the parity check matrix H and the belief propagation algorithm to detect and correct bit errors in transmission, and recover the low-dimensional observation value after compression coding. Subsequently, through the inverse process of arithmetic decoding, the discrete symbol sequence is restored to floating-point observation values, completing the initial recovery of data under channel noise interference and ensuring the integrity of the subsequent reconstructed input.
[0186] Orthogonal Matching Pursuit (OMP) is a greedy iterative algorithm based on a sparse signal model. It gradually selects the basis vectors (atoms) most correlated with the residual and uses orthogonal projection to minimize the residual energy, ultimately recovering the original sparse signal from the low-dimensional observation. Initialize the residual to the observation vector r0 = b, and calculate the correlation coefficient between the current value and each column of the observation matrix. The calculation formula is:
[0187] λ k =argmax|<a j , r k-1 >|,
[0188] Where a j is the jth column in the observation matrix A, r k-1 is the residual vector after the k-1th iteration, which represents the error between the current observation value and the recovered signal.
[0189] Iteratively select the observation matrix atom most relevant to the residual and add it to the support set Λ0=φ, and add the index to the support set. Update the sparse coefficients by the least squares method. The update formula is:
[0190]
[0191] r k =b-Ax k ,
[0192] Repeat the above process until the maximum number of iterations is reached or the residual value is less than the threshold, and then the original sparse signal is obtained.
[0193] For different multimodal data, the orthogonal matching pursuit algorithm will recover the sparse basis coefficients s and wavelet sparse coefficients S in the sparse domain respectively. i,j and MFCC sparse coefficients.
[0194] S60: Reconstruct the restored sparse coefficients of each modality to obtain text, image and audio features.
[0195] In this embodiment, the restored sparse coefficients of each modality are reconstructed, including:
[0196] S601: In the recovery phase, the BERT model is used to perform semantic completion on the sparse basis coefficients of the text to repair noise loss. The sparse basis coefficients of the text are converted into a sequence of word vectors. The self-attention layer of the bidirectional encoder representation model is used to perform contextual semantic repair on the word vector sequence to obtain the completed text features.
[0197] Specifically, for text information, the reconstructed sparse basis coefficients of the text are fed into a denoising autoencoder (DAE). After pre-cleaning the word embeddings, they are fed into the BERT model, which uses its deep Transformer network to complete missing or noise-corrupted semantics. A self-attention mechanism captures contextual associations, repairs semantic fragments lost due to compression or transmission (such as key instructions and location descriptions), and outputs complete and coherent text information about the fire scene.
[0198] S602: For the image, the wavelet coefficients are inverse transformed to generate low-resolution image blocks, which are then input into the CNN super-resolution model to restore details. After the image wavelet coefficients are inverse transformed, the high-resolution image features are obtained through the super-resolution model based on the convolutional neural network.
[0199] Specifically, the inverse discrete wavelet transform (IDWT) is performed on the restored image wavelet sparse coefficients, and the input is the orthogonal wavelet basis Output is the normalized low-resolution image block I′ i,j =ψ T (S i,j ), denormalizing each normalized low-resolution image patch to obtain a low-resolution image patch. This is then fed into a convolutional neural network super-resolution model. This model uses residual learning and upsampling layers to gradually restore high-frequency details (such as flame edges and building structures), improving image resolution and resolving blurring caused by traditional inverse transformations, providing a clear visual basis for rescue decisions.
[0200] S603: For audio, spectral subtraction is used to denoise the audio and the time domain signal is reconstructed by combining it with the inverse short-time Fourier transform (ISTFT). After the audio Mel-frequency cepstral coefficients are denoised by spectral subtraction, the audio features are generated by the inverse short-time Fourier transform.
[0201] Specifically, for audio data, spectral subtraction is used to suppress ambient noise on the reconstructed MFCC sparse coefficients. The statistical noise spectrum is calculated as follows:
[0202]
[0203] Where Y(f, t) is the noisy spectrum amplitude of the t-th frame signal at frequency point f.
[0204] Perform spectrum amplitude subtraction on each frequency point of the spectral sparse coefficient to proportionally reduce the noise component from the signal spectrum and retain the pure frequency domain characteristics of the voice and alarm sound. The spectrum subtraction formula is:
[0205]
[0206] Where Y(k,n) is the amplitude of the noisy spectrum at the kth frequency point in the nth frame, N(f) is the average amplitude of the noise at the frequency point f, α is the subtraction factor, and β is the spectrum lower limit threshold parameter.
[0207] The corrected spectrum is then converted into a time domain signal through the inverse short-time Fourier transform (ISTFT), and the Hanning window is superimposed to eliminate inter-frame discontinuities and output high-fidelity audio data.
[0208] S70. Project text, image, and audio features into a shared semantic space, calculate the joint semantic weights through a cross-modal attention mechanism, and generate semantically consistent multimodal fusion results.
[0209] In this embodiment, text, image, and audio features are projected into a shared semantic space, expressed as:
[0210] T′=W t T,
[0211] V′=W v V,
[0212] A′=W a A,
[0213] Where T, V, and A represent text, image, and audio feature vectors respectively, and W t 、W a 、W a is a learnable matrix.
[0214] The cross-modal attention mechanism dynamically calculates the inter-modal association weights: using text as the query, image and audio as the key-value, a joint semantic representation is generated. The weight calculation formula is:
[0215]
[0216] Where, dk is the dimension of the Key vector, which is used to scale the dot product result to prevent gradient explosion. Q is the projection of the text modality feature in the shared semantic space, which is the same as T′ above. K is the projection of the image modality feature in the shared semantic space, which is the same as V′ above. V is the projection of the audio modality feature in the shared semantic space, which is the same as A′ above.
[0217] The joint feature after weighted sum alignment is expressed as:
[0218] F align =Attention(Q, K, V)+T′
[0219] Where T′ is the projection of text features in the shared semantic space, which aims to preserve the original semantic information of the text and prevent the attention mechanism from being overly biased towards images or audio.
[0220] At the same time, dynamic time warping (DTW) is used to align time series data (such as voice and video actions) to minimize the time axis offset of multimodal signals. The formula is:
[0221] DTW(V, A)=min∑ i,j∈路径 ||V i -A j || 2 ,
[0222] Where V i is the feature of the i-th frame of the video, A j is the feature of the audio frame j.
[0223] The aligned output ensures the spatiotemporal consistency of “fire alarm voice-smoke diffusion image-command text” and ultimately outputs semantically coordinated multimodal rescue decision information.
[0224] The disclosed embodiments provide a multi-dimensional transmission data compression perception method for urban emergency rescue. It innovates the entire link from compression, transmission to recovery, and comprehensively solves the problems of fragmentation, inefficiency and insufficient reliability in multimodal data processing in existing technologies. It provides an efficient, robust, semantically coordinated end-to-end solution for urban fire emergency rescue, and has broad application prospects.
[0225] Example 2
[0226] As another aspect of the embodiment of the present disclosure, a multi-dimensional transmission data compression sensing system 100 for urban emergency rescue is also provided. Figure 3 As shown, including:
[0227] Multimodal data acquisition module 1, which obtains multimodal data, wherein the multimodal data is text data, image data and audio data in an urban emergency rescue scenario;
[0228] Each modality sparse coefficient acquisition module 2 performs normalization processing and sparse representation on the multimodal data to obtain each modality sparse coefficient, wherein the each modality sparse coefficient includes: text sparse basis coefficient, image wavelet coefficient and audio Mel frequency cepstral coefficient;
[0229] A low-dimensional observation value generation module 3 before encoding is used to construct a block diagonal random matrix for each modal sparse coefficient, perform independent compression sampling, and generate a low-dimensional observation value before encoding;
[0230] The anti-interference code stream generation module 4 performs arithmetic coding and low-density parity check code error correction coding on the low-dimensional observation value before encoding to generate an anti-interference code stream;
[0231] Each modal sparse coefficient recovery module 5 performs low-density parity check code decoding and arithmetic decoding on the anti-interference code stream, obtains the decoded low-dimensional observation value, and then uses the orthogonal matching pursuit algorithm in combination with the block diagonal random matrix to iterate and recover the sparse coefficients of each modal;
[0232] Each modality sparse coefficient reconstruction module 6 reconstructs the restored sparse coefficients of each modality to obtain text, image and audio features;
[0233] The multimodal fusion result generation module 7 projects the text, image and audio features into a shared semantic space, calculates the joint semantic weights through a cross-modal attention mechanism, and generates a semantically consistent multimodal fusion result.
[0234] In the absence of any contradiction, the above modules in the system of the embodiment of the present disclosure can implement any implementation of the above method.
[0235] Based on the description of the above embodiments, it can be seen that the embodiments of the present disclosure can achieve the following technical effects:
[0236] 1) This disclosure significantly reduces the storage and transmission overhead of text, image, and audio data through a multimodal joint compressed sensing framework, solves the problem of high redundancy in traditional single-modality compression, and improves data transmission efficiency in fire emergency scenarios.
[0237] 2) This paper employs a block-diagonal random matrix design with structured decomposition to reduce the storage complexity of the global matrix, improving the feasibility of deployment on resource-constrained terminal devices. Combining a k-sparse autoencoder with a high-frequency wavelet coefficient substitution strategy further enhances the sparse representation, improving reconstruction accuracy while maintaining the compression ratio.
[0238] 3) This disclosure addresses complex electromagnetic environments. The combined LDPC-arithmetic coding mechanism effectively suppresses transmission errors through layered redundancy elimination and error correction protection, significantly improving anti-interference capabilities and communication reliability. Furthermore, the cross-modal attention mechanism and dynamic time warping (DTW) technology achieve precise alignment of multimodal semantics, resolving the semantic gaps caused by data isolation in traditional solutions. This allows for the coordination and consistency of text descriptions, image details, and audio timing, significantly improving the accuracy of command decisions.
[0239] 4) This paper adopts the CNN super-resolution model and BERT semantic completion technology to quickly restore high-resolution images and complete semantics, overcoming the difficulties of detail loss and semantic ambiguity in traditional reconstruction algorithms. It combines the orthogonal matching pursuit algorithm to optimize iteration efficiency, enabling mobile terminals to complete data recovery in real time.
[0240] The present disclosure also provides an electronic device comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to implement the aforementioned multi-dimensional transmission data compression sensing method for urban emergency rescue. The electronic device can be provided as a terminal, server, or other device.
[0241] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon. When executed by a processor, the computer program instructions implement the aforementioned method for compressing multi-dimensional transmission data for urban emergency rescue. The computer-readable storage medium may be a non-volatile computer-readable storage medium.
[0242] Those skilled in the art will understand that in the specific implementation of the above-mentioned multi-dimensional transmission data compression sensing method and system for urban emergency rescue, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0243] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0244] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technical improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A multi-dimensional transmission data compression sensing method for urban emergency rescue, characterized by: The steps include: S10: Obtain multimodal data, where the multimodal data includes text data, image data, and audio data in an urban emergency rescue scenario; S20. Normalizing and sparsely representing the multimodal data to obtain sparse coefficients of each modality, wherein the sparse coefficients of each modality include: text sparse basis coefficients, image wavelet coefficients, and audio Mel-frequency cepstral coefficients; S30, constructing block diagonal random matrices for each modal sparse coefficient to perform independent compression sampling, and generating low-dimensional observation values before encoding; S40, performing arithmetic coding and low-density parity check code error correction coding on the low-dimensional observation value before encoding to generate an anti-interference code stream; S50, performing low-density parity check code decoding and arithmetic decoding on the anti-interference code stream to obtain decoded low-dimensional observation values, and then iterating using an orthogonal matching pursuit algorithm in combination with a block diagonal random matrix to recover the sparse coefficients of each mode; S60, reconstructing the recovered sparse coefficients of each modality to obtain text, image and audio features; S70. Project text, image, and audio features into a shared semantic space, calculate the joint semantic weights through a cross-modal attention mechanism, and generate semantically consistent multimodal fusion results.
2. The method according to claim 1, characterized in that The multimodal data is normalized and sparsely represented to obtain text sparse basis coefficients, image wavelet coefficients, and audio Mel-frequency cepstral coefficients, including: Generate text sparse basis coefficients for text data through a bidirectional encoder representation model and a k-sparse autoencoder; By performing block normalization and discrete wavelet transform on the image data, the image wavelet coefficients are generated; The audio Mel-frequency cepstral coefficients are generated by pre-emphasis, framing and audio feature extraction of audio data.
3. The method according to claim 2, characterized in that The image data is normalized by blocks and discrete wavelet transform to generate image wavelet coefficients, including: Divide the image into multiple sub-blocks and perform normalization; Perform discrete wavelet transform on each sub-block, and replace the high-frequency sub-band coefficients by threshold processing or neighboring block reference to obtain the image wavelet coefficients.
4. The method according to claim 2, characterized in that By pre-emphasis, framing and audio feature extraction of audio data, audio Mel-frequency cepstral coefficients are generated, including: Perform pre-emphasis, frame division, and Hamming window processing on the audio signal to extract the Mel-frequency cepstral coefficients and their dynamic characteristics; Construct a block diagonal random matrix and independently compress and sample the sub-vectors of the Mel-frequency cepstral coefficients of each frame of audio; In view of the sparsity of high-order audio Mel-frequency cepstral coefficients, the dimensionality is further reduced by combining part of the Fourier measurement matrix, and the sparse coefficients are iteratively solved through the orthogonal matching pursuit algorithm to obtain the audio Mel-frequency cepstral coefficients.
5. The method according to claim 1, wherein Perform arithmetic coding and low-density parity check error correction coding on the low-dimensional observation value before encoding to generate an anti-interference code stream, including: Perform arithmetic coding on the low-dimensional observations before encoding to generate a compact binary bit stream; Construct the system generation matrix and the corresponding check matrix; Inputting the bit stream into a system generation matrix and generating a codeword by matrix multiplication; Error detection is performed using a check matrix, and error correction is iteratively performed through a belief propagation algorithm to generate an interference-resistant code stream.
6. The method according to claim 1, characterized in that Reconstruct the recovered sparse coefficients of each modality, including: The sparse basis coefficients of the text are converted into a word vector sequence, and the word vector sequence is repaired with contextual semantics through the self-attention layer of the bidirectional encoder representation model to obtain the completed text features; After inverse transforming the image wavelet coefficients, high-resolution image features are obtained through a super-resolution model based on a convolutional neural network; The audio Mel-frequency cepstral coefficients are denoised by spectral subtraction and then the audio features are generated by inverse short-time Fourier transform.
7. The method according to claim 1, characterized in that Project text, image, and audio features into a shared semantic space, represented as: T′=W t T, V′=W v V, A′=W a A, Where T, V, and A represent text, image, and audio feature vectors respectively, and W t 、W v 、W a is a learnable matrix.
8. A multi-dimensional transmission data compression sensing system for urban emergency rescue, characterized by: include: A multimodal data acquisition module, which obtains multimodal data, wherein the multimodal data is text data, image data, and audio data in an urban emergency rescue scenario; Each modality sparse coefficient acquisition module performs normalization processing and sparse representation on the multimodal data to obtain each modality sparse coefficient, wherein the each modality sparse coefficient includes: text sparse basis coefficient, image wavelet coefficient and audio Mel frequency cepstral coefficient; A low-dimensional observation value generation module before encoding is used to construct a block diagonal random matrix for each modal sparse coefficient, perform independent compression sampling, and generate a low-dimensional observation value before encoding; The anti-interference code stream generation module performs arithmetic coding and low-density parity check error correction coding on the low-dimensional observation values before encoding to generate an anti-interference code stream; Each modal sparse coefficient recovery module performs low-density parity check code decoding and arithmetic decoding on the anti-interference code stream, obtains the decoded low-dimensional observation value, and then uses the orthogonal matching pursuit algorithm in combination with the block diagonal random matrix to iterate and recover the sparse coefficients of each modal; Each modal sparse coefficient reconstruction module reconstructs the recovered sparse coefficients of each modality to obtain text, image and audio features; The multimodal fusion result generation module projects text, image and audio features into a shared semantic space, calculates the joint semantic weights through a cross-modal attention mechanism, and generates semantically consistent multimodal fusion results.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the multi-dimensional transmission data compression sensing method for urban emergency rescue according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the multi-dimensional transmission data compression sensing method for urban emergency rescue according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Multi-source cooperative sensing information low-delay transmission method based on intelligent traffic networking
CN120785925A
Multi-modal emotion fusion robot voice style conversion method and device
CN121191489A
Multi-source detection-oriented high-definition video compression method and system
CN121691728A
A high-definition video compression method and system for multi-source detection
CN121691728B