High resolution remote sensing image compression method and system based on block modulation imaging

By employing a block modulation imaging method and a decoding network, the computational burden and resolution limitations of remote sensing image compression on satellite platforms are resolved, achieving efficient and fast remote sensing image compression and high-quality decoding, suitable for real-time missions on satellite platforms.

CN119383355BActive Publication Date: 2025-11-25UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411378824.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2025-11-25
Estimated Expiration
2044-09-30

AI Technical Summary

Technical Problem

Existing remote sensing image compression technologies place an excessive computational burden on satellite platforms, especially in high-resolution imaging. Traditional methods require a large amount of computing resources and are not suitable for the real-time mission requirements of satellite platforms. Single-pixel imaging methods are slow and the DMD size limits the resolution.

Method used

A block-based modulation imaging method is adopted, which generates a binary mask matrix to configure optical mask devices for optical processing. By combining block division and decoding networks, optical domain compression and electrical domain decoding are achieved, reducing the computational burden. Pre-configured optical mask devices are used to replace DMD, and three-dimensional convolution and cross-attention mechanisms are introduced to improve decoding performance.

Benefits of technology

It achieves efficient remote sensing image compression on satellite platforms, reduces computing resource requirements, improves imaging speed, adapts to high-resolution image requirements, reduces energy consumption, and improves decoding quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119383355B_ABST
    Figure CN119383355B_ABST
Patent Text Reader

Abstract

The application discloses a kind of high-resolution remote sensing image compression method and system based on block modulation imaging, generate a binary mask matrix, according to this configuration optical mask device in acquisition end, when acquiring high-resolution remote sensing image, optical lens is used to collect optical signal matrix, after being handled by optical mask device, modulation is carried out by relay lens, then it is converted into electrical signal matrix by photoelectric sensor, then it is overlaid after being divided into blocks to complete compression, the measurement matrix obtained is sent to receiving end;Receiving end extracts the block mask matrix corresponding to each block from binary mask matrix according to the block mode of acquisition end, then extracts each block signal matrix from measurement matrix and constitutes three-dimensional signal matrix, input decoding network that is constructed and trained in advance, and the optical signal matrix of high-resolution remote sensing image is obtained by decoding. The application proposes block modulation imaging, improves decoding restoration quality while improving compression speed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of remote sensing image compression, and more specifically relates to a high-resolution remote sensing image compression method and system based on block modulation imaging. BACKGROUND

[0002] Remote sensing images play a key role in numerous fields, and their high-resolution data provide important support for decision-making, resource optimization, and disaster risk reduction. In recent years, the progress of remote sensing technology has significantly improved the spatial resolution of images and increased the amount of data collected, causing great pressure on the storage capacity, data transmission bandwidth, and energy resources on satellites. Therefore, the processing, storage, and communication capabilities on satellites are severely affected, limiting real-time processing capabilities and data transmission from satellites to the ground / satellite-to-satellite.

[0003] Remote sensing image compression algorithms are key technologies for solving the challenges of massive remote sensing image storage and transmission. Satellite payload image compression algorithms are mainly divided into lossless compression and lossy compression. Lossless compression reduces file size by utilizing the redundancy in image data, while perfectly reconstructing the original image without losing any information. However, lossless compression usually has a low compression ratio (Cr), which cannot meet the transmission bandwidth and storage limitations of satellites. On the contrary, lossy compression can achieve significantly higher compression ratios by allowing a certain degree of controllable information loss. However, this advantage comes at the cost of increased computational demand, which can put pressure on the limited processing capabilities on satellites. The increased computational burden brought by complex compression algorithms can lead to a decline in the performance of satellite processors, which in turn affects the execution of other critical tasks. In response to the challenge of limited computing resources on specific platforms, extensive research has been conducted in the academic field. Many strategies have been proposed to reduce the encoding complexity of compression algorithms. However, these methods have not brought significant changes, and the burden on satellite processors during high-resolution imaging is still heavy.

[0004] Based on the principle of compressed sensing, single pixel imaging (SPI) proposes a new scheme, which uses the inherent sparsity of the signal to reconstruct the signal from a much smaller number of samples than required by the Nyquist-Shannon sampling theorem, providing a better choice for remote sensing image compression. Compared with traditional compression methods that perform electrical domain calculations after the sensor captures the image, SPI performs optical domain calculations before the sensor imaging through optical elements, directly capturing the compressed representation of the scene. Compared with traditional compression methods such as JPEG and JPEG2000, SPI has a simpler encoding process and is very suitable for resource-constrained platforms. However, the application of SPI in satellite scenes faces two major challenges. The first challenge is that the single pixel camera (SPC) can only capture one measurement value at a time. This requires a long acquisition time for time series measurements, which requires static imaging and is extremely detrimental to moving satellites. Limited imaging speed will also severely affect the execution of real-time tasks. The second challenge is that the SPC relies on a digital micromirror device (Digtial Micromirror Devices, DMD) for dynamic coding. Unfortunately, the limited size of the DMD cannot meet the high resolution requirements of remote sensing images. It is worth mentioning that the current SPI decoding algorithm also has very high computational complexity, which is not suitable for high-resolution images. SUMMARY

[0005] The purpose of the present application is to overcome the shortcomings of the prior art and provide a high-resolution remote sensing image compression method and system based on block modulation imaging, which introduces compressed sensing theory into remote sensing image compression to propose block modulation imaging and the corresponding decoding method, while improving the compression speed, the decoding restoration quality is also improved.

[0006] In order to achieve the above-mentioned purpose of the application, the high-resolution remote sensing image compression method based on block modulation imaging comprises the following steps:

[0007] S1: generating a binary mask matrix Where H and W represent the height and width of the optical signal matrix in the high-resolution remote sensing image acquisition process, respectively, and the binary mask matrix M is distributed to the acquisition end and the receiving end;

[0008] S2: The acquisition end configures the optical mask device according to the binary mask matrix M. When the element value in the binary mask matrix M is 1, the optical signal of the corresponding channel passes through, and when the element value is 0, the optical signal of the corresponding channel does not pass through;

[0009] S3: The acquisition end uses an optical lens to collect the optical signal matrix of the high-resolution remote sensing image of the target area After optical mask processing by the optical mask device, modulation is performed by the relay lens, and then converted into an electrical signal matrix by the photoelectric sensor

[0010] S4: The acquisition end divides the electrical signal matrix Z of the high-resolution remote sensing image into blocks, obtaining N block electrical signal matrices z of the same size. i , i = 1, 2, ..., N, superimpose the N block electrical signal matrices and normalize them to a preset value range to obtain the measurement value matrix Y and send it to the receiving end;

[0011] S5: After receiving the measurement value matrix Y, the receiving end extracts the block mask matrix m corresponding to each block from the binary mask matrix M according to the block division method of the acquisition end. i Then, the signal matrix of each block is extracted from the measurement matrix Y.

[0012]

[0013] N block signal matrices Constructing a three-dimensional signal matrix h and w represent the height and width of the block signal matrix, respectively. These are input to a pre-built and trained decoding network to decode the optical signal matrix of the high-resolution remote sensing image.

[0014] This invention also proposes a high-resolution remote sensing image compression system based on block modulation imaging, comprising a compression module at the acquisition end and a decoding module at the receiving end, wherein:

[0015] The acquisition-end compression module includes an optical lens, an optical mask device, a relay lens, a photoelectric sensor, and a compression transmission module, wherein:

[0016] Optical lenses are used to acquire high-resolution remote sensing images of target areas using optical signal matrices.

[0017] Optical mask devices use a pre-generated binary mask matrix The configuration is as follows: when an element in the binary mask matrix M has a value of 1, the optical signal of the corresponding channel passes through; when an element has a value of 0, the optical signal of the corresponding channel does not pass through. The optical mask device performs optical masking processing on the optical signal matrix X.

[0018] The relay lens modulates the optical signal matrix after optical masking.

[0019] The photoelectric sensor converts the modulated optical signal matrix into an electrical signal matrix.

[0020] The compression and transmission module is used to divide the electrical signal matrix Z of the high-resolution remote sensing image into N blocks, resulting in N identically sized block electrical signal matrices z. i, i = 1, 2, …, N, superimpose the N block electrical signal matrices and normalize to a preset value range to obtain a measurement value matrix Y and send to the receiving end;

[0021] The receiving end comprises a receiving module and a decoding network, wherein:

[0022] The receiving module is configured to extract a block mask matrix m corresponding to each block from the binary mask matrix M according to the block manner of the acquisition end i extract each block signal matrix from the received measurement value matrix Y

[0023]

[0024] superimpose the N block signal matrices to form a three-dimensional signal matrix S h and w represent the height and width of the block signal matrix respectively;

[0025] The decoding network is configured to decode the three-dimensional signal matrix S to obtain an optical signal matrix

[0026] The present application is based on a high-resolution remote sensing image compression method and system based on block modulation imaging, a binary mask matrix is generated, the acquisition end configures the optical mask device according to the binary mask matrix, when acquiring a high-resolution remote sensing image, an optical signal matrix is acquired by using an optical lens, after optical mask processing by the optical mask device, modulation is performed by a relay lens, and then converted into an electrical signal matrix by a photoelectric sensor, the electrical signal matrix is divided into blocks and superimposed to complete compression, and the obtained measurement value matrix is sent to the receiving end; the receiving end extracts a block mask matrix corresponding to each block from the binary mask matrix according to the block manner of the acquisition end, then extracts each block signal matrix from the measurement value matrix and forms a three-dimensional signal matrix, inputs a decoding network that is pre-constructed and trained, and decodes to obtain an optical signal matrix of a high-resolution remote sensing image.

[0027] The present application has the following beneficial effects:

[0028] 1) The present application realizes Hardman product in the optical domain by setting the optical mask device, and the calculation is completed instantaneously when the light passes through the optical mask device. The whole process is isolated from other systems of the satellite, so it does not need to allocate any processor resources; after conversion into an electrical signal, the block and summation of the modulated image are the main source of computational complexity in the encoding process; the compression sensing theory makes this simple encoding method possible, compared with the traditional JPEG and JPEG2000 compression algorithm, the compression sensing transfers the computational burden from encoding to decoding, saving a large amount of computational resources of the acquisition end;

[0029] 2) Compared with the limitation of SPI paradigm, the present application only needs one exposure to obtain the compressed representation of the image, not only accelerating the image acquisition, making it very suitable for real-time tasks, but also alleviating the problem of image blur caused by satellite motion;

[0030] 3) The present application can eliminate the need for DMD, and realize modulation by using a preset light mask device, solving the limitation of commercial DMD in processing high-resolution images, and canceling the deployment of DMD can significantly save energy consumption;

[0031] 4) In terms of decoding network, the present application has smaller calculation amount than SPI, is more suitable for high-resolution images, and can introduce gated three-dimensional convolution and cross-attention mechanism to further improve the decoding performance. BRIEF DESCRIPTION OF DRAWINGS

[0032] Figure 1 is the system schematic diagram of the high-resolution remote sensing image compression method based on block modulation imaging in the present application;

[0033] Figure 2 is the specific implementation method flowchart of the high-resolution remote sensing image compression method based on block modulation imaging in the present application;

[0034] Figure 3 is the structure diagram of the decoding network in the present embodiment;

[0035] Figure 4 is the structure diagram of the gated three-dimensional convolution module in the present embodiment;

[0036] Figure 5 is the structure diagram of the 3D U-net network in the present embodiment;

[0037] Figure 6 is the structure diagram of the bidirectional cross-attention module in the present embodiment;

[0038] Figure 7 is the structure diagram of the high-resolution remote sensing image compression system based on block modulation imaging in the present application.

[0039] Figure 8 is the visualization comparison diagram of the reconstructed images of the present application and SAUNet decoding network at different compression ratios on the DOTA-v1.0 dataset in the present embodiment;

[0040] Figure 9 is the visualization comparison diagram of different tasks of the present application at different compression ratios in the present embodiment;

[0041] Figure 10 is the comprehensive performance curve diagram of the present application at different compression ratios in the present embodiment;

[0042] Figure 11is the physical layout diagram of the digital acquisition system in this embodiment;

[0043] Figure 12 is a comparison diagram of the original image and the reconstructed image with a compression ratio of 16 in this embodiment. DETAILED DESCRIPTION

[0044] The specific embodiments of the present application are described below in conjunction with the accompanying drawings, so that those skilled in the art can better understand the present application. It should be particularly noted that in the following description, when the detailed description of known functions and designs may obscure the main content of the present application, these descriptions will be omitted here.

[0045] In order to better illustrate the technical solutions of the present application, first, the principle on which the present application is based is briefly described.

[0046] Figure 1 is a system schematic diagram of the high-resolution remote sensing image compression method based on block modulation imaging in the present application. In order to facilitate the display, Figure 1 All signal matrices in the present application are taken as gray-scale images. As shown in Figure 1 The present application proposes a new remote sensing image compression paradigm, called block modulation imaging (BMI), which only needs to modulate the image signal with an optical mask, then block and sum the modulated image. The whole process has negligible impact on computing resources. The decoding scheme is based on the theory of compressed sensing, which decodes through the design of a suitable deep neural network, and realizes high-quality image recovery. The present application not only inherits the advantage of ultra-low encoding complexity of compressed sensing (CS) compared with traditional methods, but also is superior to single-pixel imaging (SPI) in practicality. Compared with SPI, the present application only needs one exposure and does not need static imaging. In addition, the present application no longer needs a digital micromirror device (DMD), uses a pre-configured optical mask device, and can adapt to high-resolution images.

[0047] Based on the above process, the present application proposes a high-resolution remote sensing image compression method based on block modulation imaging. Figure 2 is a flowchart of the specific embodiment of the high-resolution remote sensing image compression method based on block modulation imaging in the present application. As shown in Figure 2 The specific steps of the high-resolution remote sensing image compression method based on block modulation imaging in the present application include:

[0048] S201: Generate a mask matrix:

[0049] Generate a binary mask matrix Wherein H, W respectively represent the height and width of the optical signal matrix in the high-resolution remote sensing image acquisition process, the binary mask matrix M is distributed to the acquisition end and the receiving end. In practical application, the binary mask matrix M can be set according to experience, or can be randomly generated, for example, the probability of each element in the binary mask matrix M taking 1 can be set according to needs, and then the binary mask matrix is generated by random value, and the probability is usually set to 50%.

[0050] S202: Configure optical mask device:

[0051] The acquisition end configures the optical mask device according to the binary mask matrix M, and the optical signal of the corresponding channel passes when the element value in the binary mask matrix M is 1, and the optical signal of the corresponding channel does not pass when the element value is 0.

[0052] S203: Acquire high-resolution remote sensing image:

[0053] The acquisition end acquires the optical signal matrix of the high-resolution remote sensing image of the target area by using the optical lens Then, after optical mask processing by the optical mask device, modulation is performed by the relay lens, and then the electrical signal matrix is converted by the photoelectric sensor

[0054] Let the modulated optical signal matrix be Then the process of mask optical processing and modulation can be represented by the following formula:

[0055] G=X⊙M

[0056] Wherein, ⊙ represents Hardman product.

[0057] The 0-1 mask can be regarded as a sparse representation, which meets the principle of compressed sensing. This mathematical framework allows the signal to be reconstructed from a small number of measurement values.

[0058] S204: Compress and send high-resolution remote sensing image:

[0059] The acquisition end divides the electrical signal matrix Z of the high-resolution remote sensing image into blocks to obtain N block electrical signal matrices z i , i=1, 2, …, N, superimpose the N block electrical signal matrices and normalize to a preset value range to obtain the measurement value matrix Y and send it to the receiving end. The measurement value matrix Y can be represented by the following formula:

[0060]

[0061] Wherein, Nor() represents the normalization operation.

[0062] S205: Receive and decode high-resolution remote sensing image:

[0063] After receiving the measurement matrix Y, the receiving end extracts the block mask matrix m corresponding to each block from the binary mask matrix M according to the block mode of the acquisition end i Then, each block signal matrix X is extracted from the measurement matrix Y

[0064]

[0065] The N block signal matrices X are combined to form a three-dimensional signal matrix X The three-dimensional signal matrix X is input into the decoding network to obtain the optical signal matrix X h and w represent the height and width of the block signal matrix, respectively.

[0066] Block Modulated Imaging (BMI) and Video Snapshot Compressive Imaging (SCI) have the same decoding principle. Compared with the decoding of video SCI, the decoding of BMI can be considered as a more challenging task because the correlation between blocks in an image cannot be guaranteed.

[0067] Since the electrical signal matrix is divided into blocks, the optical signal matrix X can also be regarded as N blocks, denoted as X i Let Φ=[d1,d2,…,d N ] T where the superscript T represents transposition, x i =vec(X i ), d i =diag(vec(M i )), and vec() represents vectorization. Let y=vec(Y), then the following expression can be obtained:

[0068]

[0069] The remote sensing image signal is reconstructed from y is to solve a pathological optimization problem:

[0070]

[0071] where || ||2 represents the two-norm, λ represents the prior regularization coefficient, and R() represents the prior regularization term.

[0072] The depth unfolding algorithm can be used to solve the above optimization problem. The basic idea of the depth unfolding algorithm is to unfold the iterative steps of traditional optimization algorithms (such as ADMM and GAP) into stages in a deep neural network. Each stage contains a linear projection and a small neural network. Start, solve The kth iteration of the following formula:

[0073]

[0074] Where v (k) is the introduced feature matrix for reconstructing the estimated signal matrix D (k) represents the neural network operation of the kth iteration step. η (k) represents the regularization term.

[0075] In the field of image restoration, the intuitive method may be to design a neural network using two-dimensional convolution. However, experimental results show that when processing high compression rate, the deep architecture constructed by many two-dimensional CNNs is difficult to reconstruct high resolution images, while three-dimensional CNNs show higher efficiency. Three-dimensional convolution can be interpreted as a method of modeling non-local relationships between image blocks. However, due to throughput considerations, only a limited number of three-dimensional convolutions are usually used. In this embodiment, considering the remote sensing application scenario, it can be assumed that the receiving end, i.e. the decoding end, has sufficient computing resources, so the complexity of the decoding network is strategically increased to produce better performance.

[0076] Based on the above analysis, a decoding network based on three-dimensional convolution is proposed in this embodiment, called BMNet. Figure 3 is the structure diagram of the decoding network in this embodiment. As Figure 3 shown, the decoding network in this embodiment includes K-level linear projection modules and 3D U-net networks, 3D-2D conversion modules, and 2D U-net networks. Next, each module will be described in detail.

[0077] The kth linear projection module and the 3D U-net network are used for kth iteration decoding, where:

[0078] The linear projection module is used to perform linear projection on the signal matrix The feature matrix v (k) obtained by linear projection is sent to the 3D U-net network, where The formula of linear projection is as follows:

[0079]

[0080] Where the superscript T represents transposition, Φ = [d1, d2, …, d N ] T , d i = diag(vec(M i )), y = vec(Y), vec() represents vectorization, η (k) represents the regularization term of the kth level.

[0081] 3D U-net network to the feature matrix v (k) performing reconstruction operation to obtain signal matrix performing output, wherein the signal matrix obtained by the first K-1 level 3D U-net network output to the next linear projection module, k' = 1, 2, …, K-1, the signal matrix obtained by the K level 3D U-net network output to the 3D-2D conversion module.

[0082] The 3D-2D conversion module is used to convert the signal matrix into N signal matrices, and then the N signal matrices are spliced into a two-dimensional signal matrix according to the blocking mode of the acquisition end and input into the 2D-Unet network.

[0083] The 2D U-net network is used to perform reconstruction operation on the input two-dimensional signal matrix to obtain the optical signal matrix of the restored high-resolution remote sensing image It is found that using only 3D convolution for image decoding is not enough because it will cause block artifacts. Therefore, the embodiment adds a 2D U-net network at the end of the decoding network to perform result refinement, so as to learn the residual between the output of the final stage of the decoding network BMNet and the true value.

[0084] The U-net network is a commonly used neural network, which adopts an encoder-decoder structure, wherein the encoder is used to perform several times of down-sampling by using a convolution module, and the decoder is used to perform several times of up-sampling by using a convolution module, so as to realize reconstruction of the input. In order to further enhance the representation ability of the decoding network, the embodiment introduces a gating mechanism for the three-dimensional convolution module in the 3D U-net network, and proposes a gated three-dimensional convolution module. Figure 4 is a structural diagram of the gated three-dimensional convolution module in the embodiment. As Figure 4 shown, the gated three-dimensional convolution module in the embodiment includes a first three-dimensional convolution module, a LeakyReLU activation module, a second three-dimensional convolution module, a third three-dimensional convolution module, a sigmoid module, a Hadamard product module and a matrix superposition module, wherein:

[0085] The first three-dimensional convolution module is used to perform three-dimensional convolution operation on the input feature B, and the obtained feature is sent to the LeakyreLU activation module.

[0086] The LeakyReLU activation module is used to activate the received feature by using the LeakyReLU activation function, and the obtained feature is sent to the second three-dimensional convolution module.

[0087] The second three-dimensional convolution module is configured to perform a three-dimensional convolution operation on the received feature, and send the obtained feature to the Hadamard product module.

[0088] The third three-dimensional convolution module is configured to perform a three-dimensional convolution operation on the input feature B, and send the obtained feature to the sigmoid module.

[0089] The sigmoid module is configured to process the received feature by using a sigmoid function, and send the obtained feature to the Hadamard product module.

[0090] The Hadamard product module is configured to calculate the Hadamard product of the two received features, and send the obtained feature F to the matrix superposition module.

[0091] The matrix superposition module is configured to superimpose the input feature B and the feature F, and output the superimposed matrix.

[0092] According to the above description, since the signal matrix used for recovery extracted by the receiving end through the block mask matrix contains invalid elements, it may not be optimal to apply a convolution kernel with equal weights to all elements. Therefore, the embodiment sets a gated three-dimensional convolution, and an additional branch is introduced to calculate the weight, so that important features can be dynamically selected, and the accuracy of the decoded remote sensing image is improved.

[0093] In the deep unfolding algorithm, it is very important to promote information exchange between different stages. Optimizing these information exchange mechanisms can alleviate the inherent information loss in the unfolding process. To address this challenge, the embodiment proposes a bidirectional cross-attention (TWCA) module to achieve sufficient inter-stage information exchange by fully utilizing the latent vectors generated by the encoder in the 3D U-net network. Figure 5 is a 3D U-net network structure diagram in the embodiment. As Figure 5 shown, the 3D U-net network in the embodiment includes an encoder, a bidirectional cross-attention module and a decoder, wherein:

[0094] The encoder is configured to encode the feature matrix v (k) to generate a latent variable u (k) , and then send it to the bidirectional cross-attention module.

[0095] The bidirectional cross-attention module is configured to use a bidirectional cross-attention mechanism to exchange information between the hidden variable h (k -1) and the latent variable u (k) generated by the previous stage in a bidirectional manner, generate an optimized latent variable u' (k) and send it to the decoder to generate a hidden variable h (k)and sent to the next stage 3D U-net network bidirectional cross attention module, wherein h (0) = u (0) , that is, the hidden variable in the first stage 3D U-net is the latent variable generated by the encoder. This exchange makes the latent variable u (k) can integrate the information reflecting all the accumulated knowledge of the previous stage from the hidden variable h (k-1) generated by the upper stage.

[0096] The decoder is used to decode the optimized latent variable u' (k) to generate a signal matrix output.

[0097] Figure 6 is the structure diagram of the bidirectional cross attention module in this embodiment. As Figure 6 shown, the bidirectional cross attention module in this embodiment includes a first three-dimensional convolution module, a second three-dimensional convolution module, a first matrix multiplication module, a first softmax module, a second matrix multiplication module, a second softmax module, a third matrix multiplication module, a fourth matrix multiplication module, a third three-dimensional convolution module, a fourth three-dimensional convolution module, a first Hadamard product module and a second Hadamard product module, wherein:

[0098] The first three-dimensional convolution module is used for three-dimensional convolution operation on the hidden variable h (k-1) generated by the upper stage to generate a query matrix q h , a value matrix v h , and a key matrix k h , which can be represented as follows:

[0099] q h = W1 k,h h (k-1)

[0100] v h = W2 k,h h (k-1)

[0101] k h = W3 k,h h (k-1)

[0102] wherein W1 k,h , W2 k,h , W3 k,h respectively represent the weight matrix corresponding to the query, value and key corresponding to the hidden variable h (k-1) .

[0103] Then the key matrix k h is sent to the first matrix multiplication module, and the query matrix q hto a second matrix multiplication module, and the value matrix v h to a third matrix multiplication module.

[0104] The second three-dimensional convolution module is configured to perform a three-dimensional convolution operation on the latent variable u (k) to generate a query matrix q u , a value matrix v u , and a key matrix k u , which can be represented as follows:

[0105] q u = W1 k,u u (k)

[0106] v u = W2 k,u u (k)

[0107] k u = W3 k,u u (k)

[0108] wherein W1 k,u , W2 k,u , and W3 k,u represent weight matrices corresponding to the query, value, and key corresponding to the hidden variable h (k-1) , respectively.

[0109] The query matrix q u is sent to a first matrix multiplication module, the key matrix k u is sent to a second matrix multiplication module, and the value matrix v u is sent to a fourth matrix multiplication module.

[0110] The first matrix multiplication module is configured to calculate the matrix product of the key matrix k h and the query matrix q u , and send the obtained feature matrix to a first softmax module.

[0111] The first softmax module is configured to process the received feature matrix by using a softmax function, and send the obtained feature matrix f1 to a third matrix multiplication module.

[0112] The third matrix multiplication module is configured to calculate the matrix product of the value matrix v h and the feature matrix f1, and send the obtained feature matrix to a third three-dimensional convolution module.

[0113] The third three-dimensional convolution module is configured to perform a three-dimensional convolution operation on the received feature matrix, and send the obtained feature matrix g1 to a second Hadamard product module.

[0114] The second Hadamard product module is configured to calculate a Hadamard product of the feature matrix g1 and the latent variable u (k) to obtain an optimized latent variable u' (k) and output.

[0115] The second matrix product module is configured to calculate a matrix product of the key matrix k u and the query matrix q h and send the obtained feature matrix to the second softmax module.

[0116] The second softmax module is configured to process the received feature matrix by using a softmax function, and send the obtained feature matrix f2 to the fourth matrix product module.

[0117] The fourth matrix product module is configured to calculate a matrix product of the value matrix v u and the feature matrix f2, and send the obtained feature matrix to the fourth three-dimensional convolution module.

[0118] The fourth three-dimensional convolution module is configured to perform a three-dimensional convolution operation on the received feature matrix, and send the obtained feature matrix g2 to the first Hadamard product module.

[0119] The first Hadamard product module is configured to calculate a Hadamard product of the feature matrix g2 and the hidden variable h (k-1) to obtain the hidden variable h (k) and output.

[0120] The training sample of the decoding network can be a high-resolution remote sensing image actually measured, and a conventional training method can be used. The loss function can be L1 loss, L2 loss, SSIM (Structure Similarity Index Measure) loss, MS-SSIM loss, etc.

[0121] In order to better apply the high-resolution remote sensing image compression method based on block modulation imaging of the present application, the present application further provides a high-resolution remote sensing image compression system based on block modulation imaging. Figure 7 is a structural diagram of the high-resolution remote sensing image compression system based on block modulation imaging of the present application. As shown in Figure 7 the high-resolution remote sensing image compression system based on block modulation imaging of the present application comprises an acquisition end compression module 1 and a receiving end decoding module 2, wherein:

[0122] The acquisition end compression module 1 comprises an optical lens 11, an optical mask device 12, a relay lens 13, a photoelectric sensor 14 and a compression sending module 15, wherein:

[0123] The optical lens 11 is configured to acquire an optical signal matrix of a high-resolution remote sensing image of a target region

[0124] The optical mask device 12 is configured according to a pre-generated binary mask matrix M, wherein when the element value of the binary mask matrix M is 1, the optical signal of the corresponding channel passes, and when the element value of the binary mask matrix M is 0, the optical signal of the corresponding channel does not pass; the optical mask device 12 performs optical mask processing on the optical signal matrix X. The relay lens 13 modulates the optical signal matrix after the optical mask processing.

[0125] The photoelectric sensor 14 converts the modulated optical signal matrix into an electrical signal matrix

[0126] The compression and transmission module 15 is configured to divide the electrical signal matrix Z of the high-resolution remote sensing image into N block electrical signal matrices z i , i = 1, 2, …, N, superimpose the N block electrical signal matrices and normalize them to a preset value range to obtain a measurement value matrix Y and send it to the receiving end.

[0127] The receiving end decoding module 2 includes a receiving module 21 and a decoding network 22, wherein:

[0128] The receiving module 21 is configured to extract a block mask matrix m i corresponding to each block from the binary mask matrix M according to the block mode of the acquisition end, and extract each block signal matrix

[0129]

[0130] The N block signal matrices

[0131] are combined to form a three-dimensional signal matrix S .h and w represent the height and width of the block signal matrix, respectively. The decoding network 22 is configured to decode the three-dimensional signal matrix S to obtain an optical signal matrix X

[0132] . The specific structure of the decoding network 22 is shown in . Figure 3 In order to better illustrate the technical solutions of the present application, specific examples are used to experimentally verify the present application.

[0133]

[0134] DOTA-v1.0 dataset is chosen in this embodiment to train the decoding model BMNet in the present application. DOTA-v1.0 dataset is a benchmark large dataset designed for object detection in aerial images. It contains 2,806 high-resolution images covering 15 object classes, including airplane, ship, storage tank, stadium, and bridge. In this embodiment, 1,411 training set images are randomly cropped to 512x512 non-overlapping image patches. From the resulting 41,672 image patches, 39,588 images are randomly selected as the training set, and the remaining 2,084 images are used as the test set. Peak signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM) are chosen as the evaluation metrics. To be consistent with SPI, compression is only applied to the luminance channel (Y) in YCbCr space. In addition to evaluating the reconstruction quality, the annotations of DOTA-v1.0 can also be used to evaluate the mean average precision (mAP) of the downstream task of object detection under different compression rates.

[0135] To investigate the impact of compression on more fine-grained downstream tasks, ISPRS Vaihingen dataset is chosen in this embodiment to test the BMNet trained on DOTA-v1.0 without fine-tuning. ISPRS Vaihingen dataset is a benchmark for evaluating semantic segmentation algorithms in remote sensing images. It contains 33 high-resolution aerial orthoimages and corresponding digital surface models covering an urban area in Vaihingen, Germany. We only use the TOP images and do not use DSM and NDSM. We use IDs 34 and 37 as the test set, cropping the images into 31 non-overlapping image patches of size 512x512. The images of the remaining IDs are cropped into 3120 non-overlapping image patches of size 512x512 for training the semantic segmentation model DC-Swin. Mean intersection over union (mIoU) is chosen as the evaluation metric.

[0136] To facilitate comparison with more SPI methods, CBSD68 is chosen in this embodiment as an additional test set. It consists of 68 natural color images. The experiment is conducted on the luminance channel of YCbCr space. PSNR is chosen as the evaluation metric in this embodiment.

[0137] First, the BMNet model in this invention is evaluated on CBSD68, a benchmark dataset for the SPI method. For fair comparison, the model in this invention also uses a jointly trained measurement matrix. Subsequently, the performance of the BMNet model and the state-of-the-art SPI decoding network SAUNet is evaluated on the DOTA-v1.0 dataset. Due to the high computational complexity of mainstream SPI decoding algorithms, especially when applied to high-resolution images, replicating them on the DOTA-v1.0 dataset would be extremely expensive. For the DOTA-v1.0 experiments, the images in this embodiment are resized to a resolution of 256×256. The model is first trained with a compression ratio of 10, and then fine-tuned with compression ratios of 4 and 25 using the pre-trained model as a base.

[0138] Figure 8 This is a visual comparison of the images reconstructed on the DOTA-v1.0 dataset by the present invention and the SAUNet decoding network at different compression ratios in this embodiment. Table 1 is a comparison table of the average PSNR values ​​of the images reconstructed on the DOTA-v1.0 dataset by the present invention and the SAUNet decoding network at different compression ratios in this embodiment.

[0139]

[0140] Table 1

[0141] like Figure 8 As shown in Table 1, the performance of BMNet in this invention is comparable to or better than that of SAUNet (SOTA SPI method), which indicates that BMI has the potential to replace SPI as a solution for on-orbit remote sensing image compression, with both lower coding complexity and superior decoding performance.

[0142] For lossy compression methods, analyzing the impact of compression ratio on subsequent downstream tasks is crucial. Most compressed sensing methods primarily focus on minimizing reconstruction error. However, reconstruction quality (measured by metrics such as mean squared error (MSE)) may not always directly correlate with the performance requirements of a particular application. In this section, a baseline will be established for object detection and semantic segmentation tasks as a benchmark for evaluating the effectiveness of future compressed sensing-based remote sensing image compression methods.

[0143] To obtain decoding networks with different compression ratios, the embodiment first establishes a baseline BMNet trained on DOTA-v1.0 with a compression ratio of 16, and then uses this pre-trained model to fine-tune additional models with compression ratios of 4, 9, 25, 36, 49, 64, 81, and 100. For the object detection task, the pre-trained YOLO-v5s model is first evaluated on the uncompressed test set, achieving an mAP of 86.5%. Then the present method is applied to compress and reconstruct the test set at different compression ratios. Then, the reconstructed test set is input into the YOLO model for detection. For the semantic segmentation task, the DC-Swin model is first trained, and achieves an mIoU of 78.6% on the uncompressed test set. Then, the test set compressed and reconstructed by the BMI framework is used again to evaluate the performance of the trained DC-Swin model at different compression ratios.

[0144] Figure 9 is a visual comparison chart of the present application in different tasks at different compression ratios in the embodiment. Figure 10 is a comprehensive performance curve chart of the present application at different compression ratios in the embodiment. Figure 10 The performance indicators include PSNR and SSIM of image reconstruction (DOTA-v1.0 dataset), mAP of object detection (DOTA-v1.0 dataset), and mIoU of semantic segmentation (Vaihingen dataset). As shown in Figure 7 and Figure 8 As shown in and, at a compression ratio of 4, the loss of semantic information is negligible, and the decline in downstream task performance is limited to within 7.5% when the compression ratio is less than 16.

[0145] In the embodiment, a digital acquisition system based on Metacam is also designed to simulate the acquisition end, thereby realizing the simulation process in the physical system. Due to its flexibility, DMD is retained in the prototype to adapt to experiments with different measurement matrix configurations. Please note that in actual applications, DMD can be replaced by a photomask to realize high-resolution imaging. Figure 11 is a physical layout diagram of the digital acquisition system in the embodiment. As shown in Figure 11As shown, the digital acquisition system includes an imaging lens assembly, an optical encoding unit, and a relay lens assembly. Through the camera lens (CHIOPT HC3505A), the scene image is projected onto a virtual plane for raw data acquisition. The relay lens (Thorlabs MAP10100100-A) focuses the target scene image onto the DMD (ViALUX V-9001, 2560x1600 resolution, 7.6 pm pixel pitch) to form a primary image. The DMD modulates the data by adjusting the amplitude of the light to achieve instantaneous encoding. The reflected light from the DMD is focused onto the image sensor (FLIR GS3-U3-120S6M-C, 4242x2830 resolution, 3.1 pm pixel pitch) through the zoom lens (Utron VTL0714V). The relay lens assembly is used to transfer the encoded data to the light-sensitive surface of the sensor, matching one DMD mirror to one image sensor pixel, ensuring accurate alignment of the DMD and the image sensor. It is worth noting that the system has flexible plug-and-play characteristics. The degree of hardware integration in the application has significant potential for enhancement. In addition, the system has the ability to adapt to various compression ratios, as the compression ratio is only dependent on the electrical domain calculation.

[0146] To verify the performance of the reconstruction algorithm in real-world scenarios, the test dataset is projected onto a surface under natural light conditions. The digital acquisition system captures the projected data to form measurements and inputs the encoded measurements into the decoding network BMNet for reconstruction. Due to the tilt hinge design of the digital micro-mirrors of the DMD, their rotational modulation occurs on an axis tilted 45° relative to the array direction. This results in a 45° rotation of the image after processing by the digital acquisition unit. To address this issue, image scaling is employed to magnify the projected image, followed by cropping to achieve a capture image size of 512x512, meeting the required size for reconstruction. In addition, we observe that due to the use of lenses in the alignment process, the calibrated mask will deviate from the binary encoding used in the simulation. To address this issue, the model is fine-tuned using the mask obtained in the real environment. Specifically, a uniform illumination source provided by an integrating sphere illuminates the entrance pupil. Then the binary mask is loaded onto the DMD. The resulting encoded light pattern serves as the actual experimental mask for actual testing.

[0147] Figure 12 is a comparison of the original image and the reconstructed image with a compression ratio of 16 in this embodiment. As shown in Figure 12 , although there are some differences between the reconstructed image and the original image, the features are still well restored, and it can adapt to engineering practical applications.

[0148] While the foregoing specific embodiments of the application have been described in some detail to provide a clear understanding thereof, it will be apparent to those of ordinary skill in the art that numerous modifications can be made to the specific embodiments described without departing from the spirit and scope of the application defined by the appended claims.

Claims

1. A high-resolution remote sensing image compression method based on block modulation imaging, characterized in that, The method comprises the following steps: S1: generating a binary mask matrix wherein H and W represent the height and width of the optical signal matrix in the high-resolution remote sensing image acquisition process, and the binary mask matrix M is distributed to the acquisition end and the receiving end; S2: The acquisition end configures the optical mask device according to the binary mask matrix M, wherein when the element value of the binary mask matrix M is 1, the optical signal of the corresponding channel passes, and when the element value is 0, the optical signal of the corresponding channel does not pass; S3: the optical signal matrix of high-resolution remote sensing image of the target area is collected by the optical lens of the collection end After the optical mask processing by the optical mask device, the modulation is performed by the relay lens, and the electrical signal matrix is converted by the photoelectric sensor S4: The acquisition end divides the electrical signal matrix Z of the high-resolution remote sensing image into blocks to obtain N block electrical signal matrices z of the same size i , i = 1, 2, …, N, superimposes the N block electrical signal matrices, normalizes them to a preset value range to obtain a measurement value matrix Y, and sends the measurement value matrix Y to the receiving end; S5: After receiving the measurement value matrix Y, the receiving end extracts the block mask matrix m corresponding to each block from the binary mask matrix M according to the block mode of the acquisition end i and then extracts each block signal matrix from the measurement value matrix Y N block signal matrices are obtained A three-dimensional signal matrix is constructed h and w represent the height and width of the block signal matrix respectively, an input is pre-constructed and trained decoding network, and an optical signal matrix of a high-resolution remote sensing image is decoded 2. The high resolution remote sensing image compression method of claim 1, wherein, In the step S1, the binary mask matrix M is randomly generated, wherein the probability of each element taking 1 is 50%.

3. The high resolution remote sensing image compression method of claim 1, wherein, The decoding network in the step S5 comprises a K-level linear projection module, a 3D U-net network, a 3D-2D conversion module and a 2D U-net network, wherein: The k-level linear projection module and the 3D U-net network are used for k-level iterative decoding, wherein: The linear projection module is used for linear projection on the signal matrix The obtained feature matrix v (k) is sent to the 3D U-net network, wherein The formula of linear projection is as follows: where the superscript T denotes transpose, Φ = [d1, d2,..., d N ] T , d i = diag(vec(M i )), y = vec(Y), vec() denotes vectorization, η (k) denotes the regularization term at the kth level; 3D U-net network performs reconstruction operation on the feature matrix v (k) to obtain a signal matrix is output, wherein the signal matrix obtained by the first K-1 level 3D U-net network is output to the next linear projection module, k'=1, 2, …, K-1, and the signal matrix obtained by the Kth level 3D U-net network is output to the 3D-2D conversion module; The 3D-2D conversion module is configured to convert the signal matrix into N signal matrices, and then splice the N signal matrices into a two-dimensional signal matrix according to the blocking manner of the acquisition end and input the two-dimensional signal matrix into the 2D-Unet network. 2D U-net network is used to reconstruct the input two-dimensional signal matrix to obtain the restored high-resolution remote sensing image optical signal matrix 4. The high resolution remote sensing image compression method of claim 3, wherein, The three-dimensional convolution module in the 3D U-net network adopts a gated three-dimensional convolution module.

5. The high resolution remote sensing image compression method of claim 3, wherein, The 3D U-net network comprises an encoder, a cross-attention module and a decoder, wherein: The encoder is used to encode the feature matrix v (k) to generate the latent variable u (k) and then send it to the bidirectional cross-attention module; The bidirectional cross-attention module is used to adopt a bidirectional cross-attention mechanism to perform bidirectional information exchange between the hidden variable h (k-1) and the latent variable u (k) , generate an optimized latent variable u' (k) , and send it to the decoder to generate a hidden variable h (k) , and send it to the bidirectional cross-attention module of the next stage 3D U-net network, where h (0) =u (0) ; The decoder is configured to decode the optimized latent variable u' to generate a signal matrix (k) performing a decoding process to generate a signal matrix for output.

6. The high-resolution remote-sensing image compression method of claim 5, wherein, The bidirectional cross-attention module comprises a first three-dimensional convolution module, a second three-dimensional convolution module, a first matrix multiplication module, a first softmax module, a second matrix multiplication module, a second softmax module, a third matrix multiplication module, a fourth matrix multiplication module, a third three-dimensional convolution module, a fourth three-dimensional convolution module, a first Hadamard product module and a second Hadamard product module, wherein: The first three-dimensional convolution module is used for performing three-dimensional convolution operation on the hidden variable h generated by the previous stage (k-1) to generate a query matrix q h , a key matrix v h and a value matrix k h Then, the key matrix k h is sent to the first matrix product module, the query matrix q h is sent to the second matrix product module, and the value matrix v h is sent to the third matrix product module The second three-dimensional convolution module is configured to perform a three-dimensional convolution operation on the latent variable u (k) to generate a query matrix q u , a key matrix v u , and a value matrix k u The query matrix q u is then sent to the first matrix multiplication module, the key matrix k u is sent to the second matrix multiplication module, and the value matrix v u is sent to the fourth matrix multiplication module. The first matrix multiplication module is configured to calculate a matrix product of the key matrix k h and the query matrix q u , and send the obtained feature matrix to the first softmax module. The first softmax module is used for processing the received feature matrix by using a softmax function, and the obtained feature matrix f1 is sent to the third matrix multiplication module; The third matrix multiplication module is configured to calculate a matrix product of the value matrix v h and the feature matrix f1, and send the obtained feature matrix to a third three-dimensional convolution module. The third three-dimensional convolution module is used for performing three-dimensional convolution operation on the received feature matrix, and the obtained feature matrix g1 is sent to the second Hadamard product module; A second Hadamard product module is used to compute the Hadamard product of the feature matrix g1 and the latent variable u (k) resulting in the optimized latent variable u' (k) and output; The second matrix multiplication module is configured to calculate a matrix product of the key matrix k u and the query matrix q h , and send the obtained feature matrix to the second softmax module. The second softmax module is used for processing the received feature matrix by using a softmax function, and the obtained feature matrix f2 is sent to the fourth matrix multiplication module; The fourth matrix multiplication module is configured to calculate a matrix product of the value matrix v u and the feature matrix f2, and send the obtained feature matrix to the fourth three-dimensional convolution module. The fourth three-dimensional convolution module is used for performing three-dimensional convolution operation on the received feature matrix, and the obtained feature matrix g2 is sent to the first Hadamard product module; The first Hadamard product module is used to calculate the Hadamard product of the feature matrix g2 and the hidden variable h (k-1) to obtain the hidden variable h (k) and output.

7. A high-resolution remote sensing image compression system based on block-modulation imaging, characterized in that, The method comprises an acquisition end compression module and a receiving end decoding module, wherein: The acquisition end compression module comprises an optical lens, an optical mask device, a relay lens, a photoelectric sensor and a compression sending module, wherein: Optical lens for collecting a matrix of optical signals of high resolution remote sensing images of a target area Optical mask device according to a pre-generated binary mask matrix is configured such that a light signal of a corresponding channel is passed when the element value of the binary mask matrix M is 1 and the light signal of the corresponding channel is not passed when the element value of the binary mask matrix M is 0; and the optical mask device performs an optical masking process on the optical signal matrix X; The relay lens modulates the optical signal matrix processed by the optical mask; The photoelectric sensor converts the modulated optical signal matrix into an electrical signal matrix The compression sending module is configured to block the electrical signal matrix Z of the high-resolution remote sensing image to obtain N block electrical signal matrices z of the same size i , i = 1, 2, …, N, superimpose the N block electrical signal matrices, normalize to a preset value range to obtain a measurement value matrix Y, and send the measurement value matrix Y to the receiving end. The receiving end comprises a receiving module and a decoding network, wherein: The receiving module is configured to extract a block mask matrix m corresponding to each block from the binary mask matrix M according to a block mode of the acquisition end i extracting each block signal matrix from the received measurement value matrix Y N block signal matrices constitute a three-dimensional signal matrix h, w represent height and width of the block signal matrix, respectively The decoding network is used to decode a three-dimensional signal matrix S to obtain an optical signal matrix 8. The high resolution remote sensing image compression system of claim 7, wherein, The decoding network comprises a K-level linear projection module, a 3D U-net network, a 3D-2D conversion module and a 2D U-net network, wherein: The k-level linear projection module and the 3D U-net network are used for k-level iterative decoding, wherein: The linear projection module is configured to perform linear projection on the signal matrix to obtain a feature matrix v (k) and send the feature matrix v to the 3D U-net network, wherein the formula of the linear projection is as follows: where the superscript T denotes transpose, Φ = [d1, d2,..., d N ] T , d i = diag(vec(M i )), y = vec(Y), vec() denotes vectorization, and η (k) denotes the regularization term at the kth level. 3D U-net network performs reconstruction operation on the feature matrix v (k) to obtain a signal matrix is output, wherein the signal matrix obtained by the first K-1 level 3D U-net network is output to the next linear projection module, k'=1, 2, …, K-1, and the signal matrix obtained by the Kth level 3D U-net network is output to the 3D-2D conversion module; The 3D-2D conversion module is configured to convert the signal matrix into N signal matrices, and then splice the N signal matrices into a two-dimensional signal matrix according to the blocking manner of the acquisition end and input the two-dimensional signal matrix into the 2D-Unet network. 2D U-net network is used to reconstruct the input two-dimensional signal matrix to obtain the restored high-resolution remote sensing image optical signal matrix 9. The high resolution remote sensing image compression system of claim 8, wherein, The 3D U-net network comprises an encoder, a cross-attention module and a decoder, wherein: The encoder is used to encode the feature matrix v (k) to generate the latent variable u (k) and then send it to the bidirectional cross-attention module; The bidirectional cross-attention module is used to adopt a bidirectional cross-attention mechanism to perform bidirectional information exchange between the hidden variable h (k-1) and the latent variable u (k) , generate an optimized latent variable u' (k) , and send it to the decoder to generate a hidden variable h (k) , and send it to the bidirectional cross-attention module of the next stage 3D U-net network, where h (0) = u (0) ; The decoder is configured to decode the optimized latent variable u' to generate a signal matrix (k) performing a decoding process to generate a signal matrix for output.

10. The high resolution remote sensing image compression system of claim 9, wherein, The bidirectional cross-attention module comprises a first three-dimensional convolution module, a second three-dimensional convolution module, a first matrix multiplication module, a first softmax module, a second matrix multiplication module, a second softmax module, a third matrix multiplication module, a fourth matrix multiplication module, a third three-dimensional convolution module, a fourth three-dimensional convolution module, a first Hadamard product module and a second Hadamard product module, wherein: The first three-dimensional convolution module is used for performing three-dimensional convolution operation on the hidden variable h generated by the previous stage (k-1) to generate a query matrix q h , a key matrix v h , and a value matrix k h Then, the key matrix k h is sent to the first matrix product module, the query matrix q h is sent to the second matrix product module, and the value matrix v h is sent to the third matrix product module The second three-dimensional convolution module is configured to perform a three-dimensional convolution operation on the latent variable u (k) to generate a query matrix q u , a key matrix v u , and a value matrix k u The query matrix q u is then sent to the first matrix multiplication module, the key matrix k u is sent to the second matrix multiplication module, and the value matrix v u is sent to the fourth matrix multiplication module. The first matrix multiplication module is configured to calculate a matrix product of the key matrix k h and the query matrix q u , and send the obtained feature matrix to the first softmax module. The first softmax module is configured to process the received feature matrix by using a softmax function, and send the obtained feature matrix f1 to the third matrix product module; The third matrix multiplication module is configured to calculate a matrix product of the value matrix v h and the feature matrix f1, and send the obtained feature matrix to a third three-dimensional convolution module. The third three-dimensional convolution module is configured to perform a three-dimensional convolution operation on the received feature matrix, and send the obtained feature matrix g1 to the second Hadamard product module; A second Hadamard product module is used to compute the Hadamard product of the feature matrix g1 and the latent variable u (k) resulting in an optimized latent variable u' (k) and output; The second matrix multiplication module is configured to calculate a matrix product of the key matrix k u and the query matrix q h , and send the obtained feature matrix to the second softmax module. The second softmax module is configured to process the received feature matrix by using a softmax function, and send the obtained feature matrix f2 to the fourth matrix product module; The fourth matrix multiplication module is configured to calculate a matrix product of the value matrix v u and the feature matrix f2, and send the obtained feature matrix to the fourth three-dimensional convolution module. The fourth three-dimensional convolution module is configured to perform a three-dimensional convolution operation on the received feature matrix, and send the obtained feature matrix g2 to the first Hadamard product module; The first Hadamard product module is used to calculate the Hadamard product of the feature matrix g2 and the hidden variable h (k-1) to obtain the hidden variable h (k) and output.