Limited communication image transmission method and device based on image mask reconstruction

Through the image mask reconstruction method, the problems of semantic loss and high latency of traditional image compression under extremely low bandwidth are solved, and efficient image transmission under low bandwidth conditions is achieved, with strong adaptability and robustness and bandwidth cost savings.

CN120499381AActive Publication Date: 2025-08-15BEIJING INST OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510710428.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-15
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Traditional image and video compression methods perform poorly under extremely low bandwidth conditions, resulting in semantic loss and high latency. Deep learning-based methods are computationally expensive, difficult to deploy, and high device performance requirements.

Method used

The image mask reconstruction method is used to downsample and segment the original image into tiles, calculate the mean and standard deviation matrix, generate the mask matrix, and reconstruct the image using a masked autoencoder at the receiving end to reduce the amount of data transmitted.

Benefits of technology

Significantly reduce bandwidth requirements, adapt to low bandwidth and unstable network environments, efficiently reconstruct images, maintain high quality, reduce transmission delay and save bandwidth costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499381A_ABST
    Figure CN120499381A_ABST
Patent Text Reader

Abstract

The invention provides a limited communication image transmission method and device based on image mask reconstruction, and the method is applied to an image transmitting end and an image receiving end, and comprises the steps: the image transmitting end carries out the processing of a to-be-transmitted original image, and obtains frame data, a mean value matrix, a standard deviation matrix, and a mask matrix; the image sending end packages the timestamp, the frame data, the statistic matrix and the mask matrix into a protocol data unit, and sends the protocol data unit to an image receiving end; and the image receiving end processes the received protocol data unit to obtain a reconstructed image. According to the invention, by reducing the transmission data volume, the bandwidth requirement can be obviously reduced, and the method is suitable for a low-bandwidth and unstable network environment; according to the invention, the missing image part can be reconstructed efficiently, the high image quality is maintained, the transmission delay is reduced, and the bandwidth cost is saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image transmission, and in particular to a method and device for limited communication image transmission based on image mask reconstruction. Background Art

[0002] In the field of image and video transmission, traditional image and video compression technologies primarily achieve efficient compression by reducing redundant information and optimizing encoding strategies. These technologies generally present no significant issues under favorable network conditions. However, with the increasing number of wireless devices and the increasing scarcity of spectrum resources, it is difficult to guarantee low latency and high bandwidth transmission conditions for all connections when transmitting over long distances or with a high number of concurrent communications within a region. Traditional image and video compression still primarily prioritizes fidelity, but its performance has significant limitations under extremely low bandwidth conditions. When an extremely low bitrate is specified as an optimization target, significant semantic loss (such as color variations, object outlines, and texture details) often occurs, significantly reducing the ability to understand and recognize the content of the transmitted image.

[0003] Existing compression algorithms are mainly divided into two types: one is a compression method that uses manually designed features and rule sets, and the other is a compression method based on deep learning. The former is mainly represented by H.264, H.265, etc. The main idea is based on block-oriented motion compensation. By dividing the content of a single frame into macroblocks of different sizes, compressing the area within the block, and then removing redundant information in the video in time and space based on the relationship between frames. This method has slightly lower computational complexity and is supported by hardware encoders. It can encode and decode more efficiently with lower power consumption and performance requirements. However, the compression rate is still limited, and it lacks adaptability to specific textures and motion scenes. It is more prone to block effects and blurring at low bit rates. The latter can automatically learn features through neural networks and can be optimized for specific scenarios, showing strong adaptability. For example, ROI-DVC uses the same specifications to input the region of interest (ROI) mask into the encoder, allowing the model to focus on areas that are important to the user. Taking into account the transmission characteristics, the VQ-DeepVSC framework uses adaptive keyframe extraction and difference modules, combined with semantic vector quantization codecs and multi-channel spatial conditional vector quantization methods, to achieve significant improvements in multi-scale structural similarity (MS-SSIM) and learning-perceptual image block similarity (LPIPS) compared to H.265. However, compression methods based on deep learning also have the characteristics of large computational workload and high deployment difficulty. Generally, they need to rely on specific hardware modules such as NPU to run the compression algorithm at the sending end, which has high requirements for device performance and low cost-effectiveness. Summary of the Invention

[0004] In view of this, the present application provides a method and device for limited communication image transmission based on image mask reconstruction to solve the above technical problems.

[0005] In a first aspect, an embodiment of the present application provides a restricted communication image transmission method based on image mask reconstruction, which is applied to an image transmitter and an image receiver, comprising:

[0006] The image sending end processes the original image to be sent to obtain frame data, a mean matrix, a standard deviation matrix and a mask matrix;

[0007] The image sending end packages the timestamp, frame data, statistic matrix and mask matrix into a protocol data unit, and sends the protocol data unit to the image receiving end;

[0008] The image receiving end processes the received protocol data unit to obtain a reconstructed image.

[0009] In one possible implementation, the image sending end processes the original image to be sent to obtain frame data, a mean matrix, a standard deviation matrix, and a mask matrix, including:

[0010] Downsample the original image I1 to obtain an image I2 of size 224×224;

[0011] The image I2 is divided into 14×14 tiles of size 16×16 pixels, each tile is represented as Where i and j represent the row and column indices of the tile in the image, 1≤i≤14, 1≤j≤14;

[0012] Calculate the mean vector of each tile and generate the mean matrix P based on the mean vectors of all tiles mean ;

[0013] Calculate the standard deviation vector of each tile and generate the standard deviation matrix P based on the standard deviation vectors of all tiles std ;

[0014] Mask the tiles according to the set masking ratio, and generate a mask matrix and an unmasked image;

[0015] Based on the unmasked image, frame data is generated.

[0016] In one possible implementation, the mean vector for each tile is calculated, including:

[0017] tiles The mean vector of for:

[0018]

[0019] in, is the mean of the R channel:

[0020]

[0021] in, For tiles The R channel value at pixel (x, y);

[0022] is the mean value of the G channel:

[0023]

[0024] in, For tiles The G channel value at pixel (x,y);

[0025] is the mean value of channel B:

[0026]

[0027] in, For tiles The B channel value at pixel (x,y).

[0028] In one possible implementation, the standard deviation vector for each tile is calculated, including:

[0029] tiles The standard deviation vector for:

[0030]

[0031] in, is the standard deviation of the R channel:

[0032]

[0033] is the standard deviation of the G channel:

[0034]

[0035] is the standard deviation of channel B:

[0036]

[0037] In one possible implementation, masking the tile according to a set masking ratio to generate a mask matrix and an unmasked image includes:

[0038] According to the predetermined masking rate, calculate the number of tiles N that need to be retained:

[0039] N=round((1-mask ratio )×14×14)

[0040] Where, mask_ratio is the masking ratio; round(·) is rounded down;

[0041] Based on the number of tiles N, the positions of the tiles to be retained are determined according to a pre-defined strategy, and a mask matrix M of size 14×14 is generated, where each element is 1 or 0, where 1 indicates that the tile at the corresponding position is retained, and 0 indicates that the tile at the corresponding position is masked;

[0042] The retained tiles are assembled into the unmasked image I in horizontal order umasked , where the height is 16 and the width is 16×N.

[0043] In one possible implementation, generating frame data based on the unmasked image includes:

[0044] The unmasked image I is encoded by a standard encoder umasked Perform binary encoding to obtain frame data in binary form.

[0045] In one possible implementation, the image receiving end processes the received protocol data unit to obtain a reconstructed image, including:

[0046] Unpack the received protocol data unit and extract the timestamp, frame data, and mean matrix P mean , standard deviation matrix P std and the mask matrix M;

[0047] The frame data is decoded by a standard decoder to obtain the unmasked image I umasked ;

[0048] Based on the mask matrix M and the unmasked image I umasked , generate a partially masked image I of size 224×224 partial ; Among them, image I partial The unmasked part of image I umasked , image I partial The masked part of is empty;

[0049] Use the pre-trained masked autoencoder to image I partial Processing is performed to obtain the first reconstructed image I reconstructed , the first reconstructed image I reconstructed The masked part is the reconstructed image I masked , while the unmasked part remains unchanged;

[0050] The first reconstructed image Ireconstructed Multiply by the standard deviation matrix P std Then, with the mean matrix P mean Add and obtain the second reconstructed image I with normal color reconstructed ;

[0051] The second reconstructed image I of normal color reconstructed Up-sample to obtain a third reconstructed image I with the same resolution as the original image final .

[0052] In one possible implementation, the method further includes:

[0053] In the order of timestamp, images I final Arrange them in sequence to generate a complete video.

[0054] In a second aspect, an embodiment of the present application provides a limited communication image transmission device based on image mask reconstruction, comprising: an image transmitting end and an image receiving end,

[0055] The image sending end is used to process the original image to be sent to obtain frame data, a mean matrix, a standard deviation matrix and a mask matrix; package the timestamp, frame data, statistic matrix and mask matrix into a protocol data unit, and send the protocol data unit to the image receiving end;

[0056] The image receiving end is used to process the received protocol data unit to obtain a reconstructed image.

[0057] By reducing the amount of transmitted data, this application can significantly reduce bandwidth requirements and adapt to low-bandwidth and unstable network environments; it can efficiently reconstruct missing image parts, maintain high image quality, reduce transmission delays, and save bandwidth costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific implementation methods or the description of the prior art. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0059] Figure 1 A flowchart of a method for limited communication image transmission based on image mask reconstruction provided in an embodiment of the present application;

[0060] Figure 2 This is a functional structure diagram of the restricted communication image transmission device based on image mask reconstruction provided in an embodiment of the present application. DETAILED DESCRIPTION

[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Generally, the components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.

[0062] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for protection, but merely represents selected embodiments of the present application. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments in the present application without creative work are within the scope of protection of the present application.

[0063] The technical solutions provided in the embodiments of the present application are described below.

[0064] like Figure 1 As shown, the embodiment of the present application provides a restricted communication image transmission method based on image mask reconstruction, which is applied to an image sending end and an image receiving end, including:

[0065] Step 101: The image sending end processes the original image to be sent to obtain frame data, a mean matrix, a standard deviation matrix, and a mask matrix;

[0066] Step 102: The image sending end packages the timestamp, frame data, statistic matrix and mask matrix into a protocol data unit, and sends the protocol data unit to the image receiving end;

[0067] Step 103: The image receiving end processes the received protocol data unit to obtain a reconstructed image.

[0068] In the embodiment of the present application, the original video frame is masked and reassembled at the image sending end, and is sent to a standard encoder for encoding and encapsulation before being transmitted to the image receiving end through a protocol. The image receiving end decodes and sends the masked autoencoder (MAE) to reconstruct the masked part, and finally forms a visual video stream through post-processing and upsampling. This can solve the problems of the feasibility of video data transmission and the high semantic loss of traditional video compression methods under extremely low bandwidth (5 to 100 kbps).

[0069] The method of the present invention significantly reduces bandwidth requirements by reducing the amount of transmitted data, making it suitable for low-bandwidth and unstable network environments. It can efficiently reconstruct missing image portions, maintain high image quality, reduce transmission latency, and save bandwidth costs. Furthermore, the method exhibits good adaptability and robustness, making it suitable for a variety of application scenarios.

[0070] In some embodiments, the image sending end processes the original image to be sent to obtain frame data, a mean matrix, a standard deviation matrix, and a mask matrix, including:

[0071] Downsample the original image I1 to obtain an image I2 of size 224×224;

[0072] The image I2 is divided into 14×14 tiles of size 16×16 pixels, each tile is represented as Where i and j represent the row and column indices of the tile in the image, 1≤i≤14, 1≤j≤14;

[0073] Calculate the mean vector of each tile and generate the mean matrix P based on the mean vectors of all tiles mean ;

[0074] Calculate the standard deviation vector of each tile and generate the standard deviation matrix P based on the standard deviation vectors of all tiles std ;

[0075] Mask the tiles according to the set masking ratio, and generate a mask matrix and an unmasked image;

[0076] Based on the unmasked image, frame data is generated.

[0077] In some embodiments, calculating the mean vector for each tile includes:

[0078] tiles The mean vector of for:

[0079]

[0080] in, is the mean of the R channel:

[0081]

[0082] in, For tiles The R channel value at pixel (x, y);

[0083] is the mean value of the G channel:

[0084]

[0085] in, For tiles The G channel value at pixel (x,y);

[0086] is the mean value of channel B:

[0087]

[0088] in, For tiles The B channel value at pixel (x,y).

[0089] In some embodiments, calculating the standard deviation vector for each tile includes:

[0090] tiles The standard deviation vector for:

[0091]

[0092] in, is the standard deviation of the R channel:

[0093]

[0094] is the standard deviation of the G channel:

[0095]

[0096] is the standard deviation of channel B:

[0097]

[0098] In some embodiments, masking the tile according to a set masking ratio to generate a mask matrix and an unmasked image includes:

[0099] According to the predetermined masking rate, calculate the number of tiles N that need to be retained:

[0100] N=round((1-mask ratio )×14×14)

[0101] Where, mask_ratio is the masking ratio; round(·) is rounded down;

[0102] Based on the number of tiles N, the positions of the tiles to be retained are determined according to a pre-defined strategy, and a mask matrix M of size 14×14 is generated, where each element is 1 or 0, where 1 indicates that the tile at the corresponding position is retained, and 0 indicates that the tile at the corresponding position is masked;

[0103] The retained tiles are assembled into the unmasked image I in horizontal order umasked , where the height is 16 and the width is 16×N.

[0104] In some embodiments, generating frame data based on the unmasked image includes:

[0105] The unmasked image I is encoded by a standard encoder umasked Perform binary encoding to obtain frame data in binary form.

[0106] In some embodiments, the received protocol data unit is unpacked to extract the timestamp, frame data, and mean matrix P. mean , standard deviation matrix P std and the mask matrix M;

[0107] The frame data is decoded by a standard decoder to obtain the unmasked image I umasked ;

[0108] Based on the mask matrix M and the unmasked image I umasked , generate a partially masked image I of size 224×224 partial ; Among them, image I partial The unmasked part of image I umasked , image I partial The masked part of is empty;

[0109] Use the pre-trained masked autoencoder to image I partial Processing is performed to obtain the first reconstructed image I reconstructed , the first reconstructed image I reconstructed The masked part is the reconstructed image I masked , while the unmasked part remains unchanged;

[0110] The first reconstructed image I reconstructed Multiply by the standard deviation matrix P std Then, with the mean matrix P mean Add and obtain the second reconstructed image I with normal color reconstructed ;

[0111] The second reconstructed image I of normal color reconstructed Up-sample to obtain a third reconstructed image I with the same resolution as the original image final .

[0112] In addition, the second reconstructed image I reconstructed Gaussian blur is applied to the edge of the completed tile to smooth the transition between the completed area and the original area; the kernel size and standard deviation of the Gaussian blur are adjusted according to the actual effect.

[0113] In some embodiments, the method further comprises:

[0114] In the order of timestamp, images I final Arrange them in sequence to generate a complete video.

[0115] Based on the above embodiments, the present application provides a limited communication image transmission device based on image mask reconstruction, see Figure 2 As shown, the limited communication image transmission device 200 based on image mask reconstruction provided by the embodiment of the present application includes at least: an image sending end and an image receiving end, wherein,

[0116] The image sending end is used to process the original image to be sent to obtain frame data, a mean matrix, a standard deviation matrix and a mask matrix; package the timestamp, frame data, statistic matrix and mask matrix into a protocol data unit, and send the protocol data unit to the image receiving end;

[0117] The image receiving end is used to process the received protocol data unit to obtain a reconstructed image.

[0118] It should be noted that the principle of solving the technical problem by the restricted communication image transmission device 200 based on image mask reconstruction provided in the embodiment of the present application is similar to the method provided in the embodiment of the present application. Therefore, the implementation of the restricted communication image transmission device 200 based on image mask reconstruction provided in the embodiment of the present application can refer to the implementation of the method provided in the embodiment of the present application, and the repeated parts will not be repeated.

[0119] It should be noted that although several units or subunits of the device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, depending on the embodiment of the application, the features and functions of two or more units described above can be embodied in a single unit. Conversely, the features and functions of a single unit described above can be further divided and embodied by multiple units.

[0120] Furthermore, although the operations of the method of the present application are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in this particular order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0121] Finally, it should be noted that the above embodiments are intended only to illustrate the technical solutions of this application and are not intended to limit the scope of the present invention. Although this application has been described in detail with reference to the embodiments, it should be understood by those skilled in the art that modifications or equivalent substitutions to the technical solutions of this application do not depart from the spirit and scope of the technical solutions of this application and should be encompassed by the claims of this application.

Claims

1. A constrained communication image transmission method based on image mask reconstruction, applied to an image transmitter and an image receiver, characterized in that: include: The image sending end processes the original image to be sent to obtain frame data, a mean matrix, a standard deviation matrix and a mask matrix; The image sending end packages the timestamp, frame data, statistic matrix and mask matrix into a protocol data unit, and sends the protocol data unit to the image receiving end; The image receiving end processes the received protocol data unit to obtain a reconstructed image.

2. The method according to claim 1, characterized in that The image transmitter processes the original image to be sent and obtains frame data, mean matrix, standard deviation matrix and mask matrix, including: Downsample the original image I1 to obtain an image I2 of size 224×224; The image I2 is divided into 14×14 tiles of size 16×16 pixels, each tile is represented as Where i and j represent the row and column indices of the tile in the image, 1≤i≤14, 1≤j≤14; Calculate the mean vector of each tile and generate the mean matrix P based on the mean vectors of all tiles mean ; Calculate the standard deviation vector of each tile and generate the standard deviation matrix P based on the standard deviation vectors of all tiles std ; Mask the tiles according to the set masking ratio, and generate a mask matrix and an unmasked image; Based on the unmasked image, frame data is generated.

3. The method according to claim 2, characterized in that Calculate the mean vector for each tile, including: tiles The mean vector of for: in, is the mean of the R channel: in, For tiles The R channel value at pixel (x, y); is the mean value of the G channel: in, For tiles The G channel value at pixel (x,y); is the mean value of channel B: in, For tiles The B channel value at pixel (x,y).

4. The method according to claim 3, characterized in that Calculate the standard deviation vector for each tile, including: tiles The standard deviation vector for: in, is the standard deviation of the R channel: is the standard deviation of the G channel: is the standard deviation of channel B:

5. The method according to claim 4, characterized in that Mask the tile according to the set masking ratio to generate a mask matrix and an unmasked image, including: According to the predetermined masking rate, calculate the number of tiles N that need to be retained: N=round((1-mask ratio )×14×14) Where, mask_ratio is the masking ratio; round(·) is rounded down; Based on the number of tiles N, the positions of the tiles to be retained are determined according to a pre-defined strategy, and a mask matrix M of size 14×14 is generated, where each element is 1 or 0, where 1 indicates that the tile at the corresponding position is retained, and 0 indicates that the tile at the corresponding position is masked; The retained tiles are assembled into the unmasked image I in horizontal order umasked , where the height is 16 and the width is 16×N.

6. The method according to claim 5, characterized in that Based on the unmasked image, generate frame data, including: The unmasked image I is encoded by a standard encoder umasked Perform binary encoding to obtain frame data in binary form.

7. The method according to claim 1, characterized in that The image receiving end processes the received protocol data unit to obtain a reconstructed image, including: Unpack the received protocol data unit and extract the timestamp, frame data, and mean matrix P mean , standard deviation matrix P std and the mask matrix M; The frame data is decoded by a standard decoder to obtain the unmasked image I umasked ; Based on the mask matrix M and the unmasked image I umasked , generate a partially masked image I of size 224×224 partial ; Among them, image I partial The unmasked part of image I umasked , image I partial The masked part of is empty; Use the pre-trained masked autoencoder to image I partial Processing is performed to obtain the first reconstructed image I reconstructed , the first reconstructed image I reconstructed The masked part is the reconstructed image I masked , while the unmasked part remains unchanged; The first reconstructed image I reconstructed Multiply by the standard deviation matrix P std Then, with the mean matrix P mean Add and obtain the second reconstructed image I with normal color reconstructed ; The second reconstructed image I of normal color reconstructed Up-sample to obtain a third reconstructed image I with the same resolution as the original image final .

8. The method according to claim 7, characterized in that The method further comprises: In the order of timestamp, the images I final Arrange them in sequence to generate a complete video.

9. A limited communication image transmission device based on image mask reconstruction, characterized in that: include: Image sending end and image receiving end, The image sending end is used to process the original image to be sent to obtain frame data, a mean matrix, a standard deviation matrix and a mask matrix; Packing the timestamp, frame data, statistic matrix and mask matrix into a protocol data unit, and sending the protocol data unit to the image receiving end; The image receiving end is used to process the received protocol data unit to obtain a reconstructed image.

Citation Information

Patent Citations

  • Image recovery method and device in video transmission and storage medium

    CN117812273A

  • Video snapshot compression imaging method and system based on reconstruction difficulty perception

    CN117834899A

  • Generation-entropy estimation combined limit image compression and decompression method and system

    CN117857795A

  • Multi-view missing data complementing method and device, electronic equipment and storage equipment

    CN118364234A

  • Methods for delineating cellular regions and classifying regions of histopathology and microanatomy

    US20150110381A1