T2X satellite-ground cooperation-oriented image reconstruction method, device and equipment
By adding masks to the original image and performing feature fusion and entropy encoding, the image transmission problem under the bandwidth limitation of satellite communication is solved, and the image reconstruction quality under extremely low data volume constraints is improved.
Patent Information
- Application Number
- CN202510379377.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-27
AI Technical Summary
In areas where cellular networks are insufficiently covered, satellite communication bandwidth resources are limited, resulting in serious limitations on the amount of data back-passed by the original image of the vehicle driver monitoring service, affecting the real-time and effectiveness of data transmission.
By adding a mask to the original image, global features and mask features are extracted, and fused in the channel dimension and spatial dimensions, a probability distribution is generated using the entropy model for encoding, and a bit stream is sent for image reconstruction.
Under the extremely low data volume constraint, the local peak signal-to-noise ratio of the reconstructed image is improved, the amount of data of hidden space characteristics is reduced, the clarity of the reconstructed image is improved, and information redundancy is reduced.
Smart Images

Figure CN120302015A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technologies, and in particular, to an image reconstruction method, device, and equipment for vehicle-to-everything (V2X) satellite-terrestrial collaboration. Background Art
[0002] In the existing vehicle transportation safety management system, the real-time monitoring of drivers' abnormal behaviors has become a core link in ensuring freight transportation safety. However, in areas with insufficient cellular network coverage such as mountainous regions, signal blind spots are likely to occur, resulting in data transmission interruption or delay. This makes it impossible to provide seamless network services for vehicles relying solely on the cellular network. In this case, satellite communication, as a technology capable of providing all-weather communication services, combined with the cellular network, can effectively make up for this defect and provide seamless network coverage for vehicles.
[0003] Nevertheless, when the vehicle is in an area where the cellular network is unavailable, it can only rely on satellite communication for data transmission. However, the problem is that the bandwidth resources of satellite communication are very limited, which leads to a serious limitation on the amount of original image backhaul data in the vehicle driver monitoring service under the condition of bandwidth limitation. Therefore, there is a contradiction between the high-bandwidth demand of the vehicle driver monitoring service and the low bandwidth of satellite communication under the constraint of extremely low data volume, which directly affects the real-time and effectiveness of data transmission. Summary of the Invention
[0004] In view of this, the present application provides an image reconstruction method, device, and equipment for vehicle-to-everything (V2X) satellite-terrestrial collaboration to solve the deficiencies in the related technologies.
[0005] In the first aspect of the present application, an image reconstruction method for vehicle-to-everything (V2X) satellite-terrestrial collaboration is provided. The method includes:
[0006] Adding a mask to the target area in the original image to obtain a masked image, and respectively performing feature extraction on the original image and the masked image to obtain global features and mask features;
[0007] Fusing the global features and the mask features respectively in the channel dimension and the spatial dimension, and fusing the fusion results from the channel dimension and the spatial dimension according to the fusion weights to obtain the latent space representation of the original image;
[0008] Performing quantization processing on the latent space representation to obtain a quantized latent space representation, and inputting the quantized latent space representation and the latent space representation into a preset entropy model to obtain the probability distribution of the quantized latent space representation;
[0009] Encoding the quantized latent space representation based on the probability distribution, and sending the encoding result to the receiving end in the form of a bit stream, so that the receiving end decodes the bit stream to obtain a reconstructed image.
[0010] According to an embodiment of the present application, the original image is a driver monitoring image, and the method further includes:
[0011] Performing semantic segmentation on the original image, and determining an image region with a semantic category of non-driver as the target region.
[0012] According to an embodiment of the present application, the fusing the global feature and the mask feature in the channel dimension and the spatial dimension respectively includes:
[0013] Performing feature splicing and dimension adjustment on the global feature and the mask feature in the channel dimension to obtain a first feature;
[0014] Extracting local spatial features from the first feature and adding position encoding information.
[0015] According to an embodiment of the present application, the inputting the quantized latent space representation and the latent space representation into a preset entropy model to obtain the probability distribution of the quantized latent space representation includes:
[0016] Dividing the quantized latent space representation into multiple slices along the channel dimension;
[0017] Inputting the multiple slices and the latent space representation into the entropy model, so that the entropy model sequentially performs serial processing on each slice to obtain the probability distribution of each slice;
[0018] The encoding the quantized latent space representation based on the probability distribution includes:
[0019] For each slice, encoding the slice based on the probability distribution of the slice.
[0020] According to an embodiment of the present application, based on the input latent space representation, generating hyperprior context information through a hyperprior network;
[0021] For the currently input slice, based on the currently input slice and all slices input before the currently input slice, generating channel context information, local context information, intra-slice global context information, and inter-slice global context information of the currently input slice through a context network;
[0022] Based on the hyperprior context information, the channel context information, the local context information, the intra-slice global context information, and the inter-slice global context information of the current slice, through a probability prediction sub-network g ep Generating the probability distribution of the current slice.
[0023] According to an embodiment of the present application, sending the coding result to a receiving end in the form of a bitstream so that the receiving end decodes the bitstream to obtain a reconstructed image includes:
[0024] Sending the coding result to a receiving end in the form of a bitstream so that the receiving end decodes each slice in the multiple slices in sequence based on the same probability distribution as the sending end, reconstructs the quantized latent space representation, and generates a reconstructed image.
[0025] In a second aspect of the present application, there is provided an image reconstruction device for T2X satellite-ground cooperation, and the device includes:
[0026] An extraction unit, configured to add a mask to a target region in an original image to obtain a masked image, and respectively perform feature extraction on the original image and the masked image to obtain a global feature and a masked feature;
[0027] A fusion unit, configured to fuse the global feature and the masked feature respectively in the channel dimension and the spatial dimension, and fuse the fusion result in the channel dimension and the spatial dimension according to a fusion weight to obtain a latent space representation of the original image;
[0028] A processing unit, configured to perform quantization processing on the latent space representation to obtain a quantized latent space representation, and input the quantized latent space representation and the latent space representation into a preset entropy model to obtain a probability distribution of the quantized latent space representation;
[0029] An encoding unit, configured to encode the quantized latent space representation based on the probability distribution, and send the encoding result to a receiving end in the form of a bitstream so that the receiving end decodes the bitstream to obtain a reconstructed image.
[0030] In a third aspect of the present application, there is provided an electronic device, including a processor and a memory, where the memory stores machine-executable instructions that can be executed by the processor, and the processor is configured to execute the machine-executable instructions to implement the steps of the method proposed in the above embodiment.
[0031] In a fourth aspect of the present application, there is provided a machine-readable storage medium, where machine-executable instructions are stored in the machine-readable storage medium, and when the machine-executable instructions are executed by a processor, the steps of the method proposed in the above embodiment are implemented.
[0032] In a fifth aspect of the present application, there is provided a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method proposed in the above embodiment are implemented.
[0033] As can be seen from the above technical solutions, a masked image is obtained by adding a mask to the target region in the original image, and global features and masked features are respectively extracted from the original image and the masked image; the global features and the masked features are fused in the channel dimension and the spatial dimension respectively, and the fused results are fused from the channel dimension and the spatial dimension according to the fusion weights to obtain the latent space representation of the original image; the latent space representation is quantized to obtain a quantized latent space representation, and the quantized latent space representation and the latent space representation are input into a preset entropy model to obtain the probability distribution of the quantized latent space representation; the quantized latent space representation is encoded based on the probability distribution and the encoding result is sent to the receiving end in the form of a bit stream, so that the receiving end decodes the bit stream to obtain a reconstructed image. By using the mask to sparsify the original image, the sparsity of the extracted latent space features is stronger, the data volume of the latent space features is reduced, and thus the local peak signal-to-noise ratio of the reconstructed image is improved under the constraint of extremely low data volume, and the clarity of the reconstructed image at an extremely low bit rate is enhanced. In addition, by fusing features from the channel dimension and the spatial dimension according to the fusion weights, the information redundancy caused by masked feature extraction can be reduced, and the model performance can be improved.
[0034] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a schematic flowchart of an image reconstruction method for T2X satellite-ground cooperation provided by an embodiment of this application;
[0036] Figure 2 is a schematic diagram of a truck satellite-ground cooperation communication system provided by an embodiment of this application;
[0037] Figure 3 is a schematic flowchart of an image reconstruction method for T2X satellite-ground cooperation provided by an embodiment of this application;
[0038] Figure 4 is a performance comparison diagram provided by an embodiment of this application;
[0039] Figure 5 is a visualization comparison diagram of latent space features and bit allocation provided by an embodiment of this application;
[0040] Figure 6 is a visualization comparison diagram of latent space features and bit allocation of the 5 channels with the maximum entropy provided by an embodiment of this application;
[0041] Figure 7 is a schematic structural diagram of an image reconstruction device for T2X satellite-ground cooperation provided by an embodiment of this application;
[0042] Figure 8 It is a schematic diagram of the hardware structure of an electronic device shown in an exemplary embodiment of the present application. Detailed implementation manners
[0043] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0044] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the" and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0045] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application and make the above objects, features and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the drawings.
[0046] In the existing vehicle transportation safety management system, the real-time monitoring of drivers' abnormal behaviors has become the core link to ensure freight safety. Taking trucks as an example, during the highway transportation of trucks, continuous communication needs will be generated. The communication needs of various truck operations are defined as Truck-to-everything (T2X). For example, truck drivers need to communicate with transportation companies, customers, and other vehicles in a timely and reliable manner to ensure that the goods are safely transported to the destination. At the same time, truck drivers need to obtain real-time road condition information and accurate navigation and positioning information to improve transportation efficiency and reduce delays and losses. For freight transportation companies and customers, it is very important to track and monitor the location and status of goods. Trucks need to communicate in real time to feedback the status and location of goods to ensure the safe transportation of goods. Logistics companies need to carry out the scheduling, allocation, and management of transportation tasks, as well as the management of drivers and vehicles. Trucks need to receive real-time tracking and management information of transportation tasks. Management departments need to monitor the driving behaviors of truck drivers in real time and intervene and warn them in a timely manner for abnormal driving behaviors of truck drivers.
[0047] However, in areas with insufficient cellular network coverage such as mountainous regions, signal blind spots are likely to occur, resulting in data transmission interruption or delay, which makes it impossible to provide seamless network services for vehicles relying solely on the cellular network. In this case, satellite communication, as a technology capable of providing all-weather communication services, combined with the cellular network, can effectively make up for this defect and provide seamless network coverage for vehicles.
[0048] Nevertheless, when the vehicle is in an area where the cellular network is unavailable, it can only rely on satellite communication for data transmission. However, the problem is that the bandwidth resources of satellite communication are very limited, which leads to a serious limitation on the amount of original image backhaul data of the vehicle driver monitoring service under the condition of bandwidth limitation. Therefore, there is a contradiction between the high bandwidth demand of the vehicle driver monitoring service and the low bandwidth of satellite communication under the constraint of extremely low data volume, which directly affects the real-time and effectiveness of data transmission.
[0049] In view of this, the embodiments of the present application disclose an image reconstruction method for T2X satellite-terrestrial collaboration to solve the deficiencies in the related art.
[0050] As Figure 1 shown, Figure 1 is a schematic flowchart of an image reconstruction method for T2X satellite-terrestrial collaboration provided by the embodiments of the present application. This method can be executed by a computing device. Exemplarily, the computing device can be an embedded system, a distributed computing node, or an intelligent terminal device with data processing capabilities, etc. The present application does not limit the type of the computing device. Exemplarily, the computing device can be mounted on vehicles, drones and other carriers. The present application does not limit the deployment method of the computing device.
[0051] Exemplarily, the image reconstruction method for T2X satellite-terrestrial collaboration can be applied in scenarios with extremely low data volume constraints, such as satellite-terrestrial collaborative communication, drone emergency communication scenarios, etc. Taking the truck satellite-terrestrial collaborative communication as an example, as Figure 2 shown, Figure 2 is a schematic diagram of a truck satellite-terrestrial collaborative communication system provided by the embodiments of the present application. Its core architecture consists of a ground station (truck), a geostationary satellite and a receiving end. The truck equipped with a satellite antenna uploads the image data of the driver monitoring service to the geostationary satellite through a satellite link, and is relayed by the geostationary satellite to the receiving end to complete image reconstruction. However, the bandwidth resources of satellite communication are usually extremely limited, and there is a contradiction between the high bandwidth demand required by the truck driver monitoring service and the low bandwidth limitation of satellite communication under the extremely low data volume constraint.
[0052] The image reconstruction method for T2X satellite-terrestrial collaboration may include the following steps:
[0053] S201: Add a mask to the target region in the original image to obtain a masked image, and perform feature extraction on the original image and the masked image respectively to obtain global features and masked features.
[0054] Add a mask to the target region in the obtained original image to obtain a masked image. After obtaining the masked image, process it in two parallel pipelines. One pipeline is used to perform feature extraction on the original image to obtain global features, and the other pipeline is used to perform feature extraction on the masked image to obtain masked features.
[0055] In some embodiments, semantic segmentation can be performed on the original image to determine the target region in the original image, and then a mask is added to the target region in the original image to obtain a masked image. Exemplarily, the target region can be a specific region of interest or non - interest in the original image.
[0056] In some embodiments, the original image can be a vehicle driver monitoring image, and the target region can be a specific region of non - interest in the original image, such as the non - driver region.
[0057] Specifically, semantic segmentation can be performed on the vehicle driver monitoring image to classify each pixel point of the vehicle driver monitoring image into specific semantic categories, such as driver, window, steering wheel, etc.; after completing semantic segmentation, identify which pixel points belong to the semantic category of the driver, and thus attribute the other pixel points to the non - driver semantic category; determine the region where the pixel points with the semantic category of the driver are located as the driver region, and determine the region where the pixel points with the semantic category of the non - driver are located as the non - driver region.
[0058] After determining the non - driver region, a mask matrix can be generated based on the semantic classification result. In this mask matrix, the matrix element value of the driver region is 1, and the matrix element value of the non - driver region is 0. Then, perform element - by - element multiplication of the mask matrix and the vehicle driver monitoring image to obtain a masked image.
[0059] In this embodiment, by adding a mask to the non - driver region, the original image is sparsified, and further the latent space features are sparsified, thereby reducing the amount of latent space feature data and improving the model performance.
[0060] In some embodiments, for the model used to extract masked features, during the training phase of the model, the image regions that the model can learn can be explicitly specified through mask constraints, guiding the learning behavior of the model during the training phase, forcing the model to learn the regional distribution information of the image, learning which regions require more bits to represent and which regions can be allocated fewer bits, so as to optimize the bit allocation region and ratio of the image.
[0061] For example, in a driver monitoring image, it can be divided into a driver area and a non-driver area. A mask is added to the non-driver area, and the pixel values of these areas are set to 0 (inactive state), while the pixel values of the driver area remain unchanged. Through the mask constraint, the model can only learn the features of the driver area during the training phase, and the features of the non-driver area are effectively masked, thus ensuring that the mask feature extraction is strictly limited to the area delimited by the mask constraint.
[0062] Exemplarily, to implement the above mask constraint, it can be flexibly set at different levels such as the input layer or intermediate feature layer of the model. For example, in the convolutional layer, the activation values of the masked area are suppressed through per-channel multiplication operations. Specifically, the mask is multiplied with the feature map in an element-wise multiplication manner, so that the feature values outside the masked area are set to zero, thereby implementing the constraint on the learning behavior of the model.
[0063] S202: Fuse the global feature and the mask feature respectively in the channel dimension and the spatial dimension, and fuse the fusion results from the channel dimension and the spatial dimension according to the fusion weights to obtain the latent space representation of the original image.
[0064] The latent space representation is an intermediate representation generated by the neural network. The latent space representation has a more compact dimension and can capture the key features of the data.
[0065] There is a lack of information interaction between the global feature extracted based on the original image and the mask feature extracted based on the masked image, resulting in information redundancy between the mask feature and the global feature. In the embodiments of the present application, the global feature and the mask feature are fused respectively in the channel dimension and the spatial dimension, and then the fusion results are feature-fused from both the channel dimension and the spatial dimension according to the fusion weights to obtain the latent space representation of the original image, so that the features can be adaptively fused from the channel dimension and the spatial dimension, reducing the information redundancy between the features and improving the model performance.
[0066] In some embodiments, fusing the global feature and the mask feature respectively in the channel dimension and the spatial dimension includes: performing feature concatenation and dimension adjustment on the global feature and the mask feature in the channel dimension to obtain a first feature; extracting local spatial features from the first feature and adding position encoding information.
[0067] Specifically, the global feature and the mask feature can be first concatenated along the channel dimension. That is, the two feature tensors are merged into a new feature tensor. During this process, the width (W) and height (H) of the feature tensor remain unchanged, while the number of channels (C) is added. For example, the global feature: with dimension C1×H×W, and the mask feature: with dimension C2×H×W. Concatenating the two feature tensors along the channel dimension results in a feature tensor with dimension (C1 + C2)×H×W. This new feature tensor contains all the channels of the global feature and all the channels of the mask feature, while the height and width remain unchanged, enabling the model to utilize the information in both the global feature and the mask feature simultaneously, thereby improving its representation ability and performance.
[0068] Adjust the dimension of the concatenated feature in the channel dimension to obtain the first feature. The concatenated feature can be processed by a 1x1 convolutional kernel to adjust its number of channels to a preset number of channels M. The essence of a 1x1 convolutional kernel is to perform a linear combination of the channels at each spatial position (H×W). Since the size of the convolutional kernel is 1x1, it does not change the spatial dimension of the feature map (i.e., height H and width W), but it will fuse and reorganize the information of different channels in the channel dimension. To adjust the number of channels of the concatenated feature tensor to the preset number of channels M, M 1x1 convolutional kernels can be used for processing, and the dimension of the resulting feature tensor becomes M×H×W.
[0069] Extract local spatial features from the first feature. The local spatial features can be extracted from the first feature through a 3×3 depthwise separable convolution. Depthwise separable convolution is an efficient convolution operation. Each input channel is assigned a corresponding convolutional kernel. Assuming the dimension of the feature tensor is M×H×W, then the depthwise convolution will use M 3×3 convolutional kernels, and each convolutional kernel only acts on the corresponding input channel. Through the 3×3 depthwise convolutional kernel, local spatial information within a 3×3 neighborhood around each pixel, such as edges and corners, can be captured.
[0070] Add position encoding information. The position encoding information can be directly added to the feature through a 1x1 convolution and / or a 3×3 depthwise separable convolution operation. Exemplarily, the convolution operation can implicitly add position encoding information through zero padding and boundary effects. Exemplarily, for the position information of each spatial position, it can be absolute position information (such as the row and column positions of a pixel in an image) or relative position information (such as the relative distance between a pixel and other pixels). In this embodiment, by adding position encoding information, the model can more effectively process and understand spatial position relationships.
[0071] In an embodiment of the present application, after fusing the global feature and the mask feature in the channel dimension and the spatial dimension respectively, the fused result will also be fused from the channel dimension and the spatial dimension according to the fusion weights.
[0072] The fusion weights include the fusion weights of different channel dimensions and the fusion weights of different spatial dimensions. Regarding the fusion weights of the channel dimension, the weight of each channel represents the contribution degree of the channel feature to the current task. For example, in driver behavior monitoring, channels related to the driver, such as channels related to the driver's face, hands, etc., will be given higher weights, while channels related to background interference will be given lower weights. Regarding the fusion weights of the spatial dimension, the fusion weight of each spatial position represents the importance of this area to the current task. For example, in driver behavior monitoring, areas related to the driver, such as the face area, the area where the hand interacts with the steering wheel, etc., will be given higher weights, while non-driver areas such as the background area will be given lower weights.
[0073] In some embodiments, the fusion weights can be fixed or dynamically calculated.
[0074] In some embodiments, a specified module can be used to dynamically calculate the fusion weights, and the fused result is fused from the channel dimension and the spatial dimension according to the fusion weights. Exemplarily, the specified module can be Swin Transformer, Swin Transformer v2, Residual Swin Transformer v2, etc., and the embodiments of the present application do not limit this.
[0075] Taking Residual Swin Transformer v2 as an example, based on Multi-head Self-Attention (MSA), Multilayer Perceptron (MLP) layer, Patch Merging, etc., the adjustment of the fusion weights of different channel dimensions can be realized. At the same time, based on mechanisms such as Window-based Self-Attention and Shifted Window, the adjustment of the fusion weights of different spatial dimensions can be realized.
[0076] S203: Quantize the latent space representation to obtain a quantized latent space representation, and input the quantized latent space representation and the latent space representation into a preset entropy model to obtain the probability distribution of the quantized latent space representation.
[0077] Quantization is the process of mapping continuous numerical values or vectors to discrete values or intervals, thereby reducing the representation complexity, storage requirements, or computational amount of data. The discrete form of the latent space representation obtained through quantization is the quantized latent space representation.
[0078] The entropy model is a model used to estimate the probability distribution of data. In the embodiments of the present application, the pre-trained entropy model can output the probability distribution of the values of the quantized latent space representation according to the input latent space representation and the quantized latent space representation.
[0079] In some embodiments, inputting the quantized latent space representation and the latent space representation into a preset entropy model to obtain the probability distribution of the quantized latent space representation includes: dividing the quantized latent space representation into multiple slices along the channel dimension; inputting the multiple slices and the latent space representation into the entropy model so that the entropy model processes each slice serially in turn to obtain the probability distribution of each slice.
[0080] Specifically, the quantized latent space representation can be evenly divided along the channels to obtain multiple slices. For example, if the quantized latent space representation has 320 channels, it can be evenly divided along the channels to successively obtain slice 1 to slice 10, and each slice has 32 channels.
[0081] Inputting the multiple divided slices and the latent space representation into the entropy model, the entropy model will process each slice one by one, and the processing of each slice depends on other slices that have been processed. For example, inputting slice 1 to slice 10 and the latent space representation into the entropy model, the entropy model processes slice 1 to slice 10 in turn to obtain the probability distribution of each slice.
[0082] In some embodiments, the preset entropy model is specifically used for:
[0083] S2031: Generate hyperprior context information through a hyperprior network based on the input latent space representation.
[0084] Specifically, the hyperprior network includes a hyperprior analysis transformation network h a and a hyperprior synthesis transformation network h s . Input the latent space representation into the hyperprior analysis transformation network h a , obtain the auxiliary information z, after quantization to obtain the quantized auxiliary information z, encode the quantized auxiliary information z through an arithmetic encoder AE (Arithmetic Encoder) to obtain a bitstream, decode it through an arithmetic decoder AD (Arithmetic decoder) and then pass it through the hyperprior synthesis transformation network h s to obtain the hyperprior context information.
[0085] S2032: For the currently input slice, based on the current slice and all slices input before the current slice, generate the channel context information, local context information, intra-slice global context information, and inter-slice global context information of the current slice through a context network.
[0086] Specifically, the context network can be composed of a channel context module, a local context module, an intra-slice global context module, and an inter-slice global context module.
[0087] For the current slice, the channel context module can capture the channel context information from all slices input before the current slice. Each slice can include an anchor part and a non-anchor part. The local context module can generate local context information based on the anchor part of the current slice. The intra-slice global context module can generate intra-slice global context information based on the correlation between the anchor part of the current slice and the anchor part and non-anchor part of the previous slice. The inter-slice global context module can capture the global correlation between the anchor part of the current slice and the previous slice and generate inter-slice global context information.
[0088] S2033: Based on the hyperprior context information, as well as the channel context information, local context information, intra-slice global context information, and inter-slice global context information of the current slice, through the probability prediction sub-network g ep Generate the probability distribution of the current slice.
[0089] The probability prediction sub-network g ep Can be a pre-trained neural network that can generate the probability distribution of the current slice according to the input context information.
[0090] For the current slice, input the channel context information, local context information, intra-slice global context information, and inter-slice global context information of the current slice, as well as the hyperprior context information, into the probability prediction sub-network g ep , and generate the probability distribution of the current slice.
[0091] S204: Encode the quantization latent space representation based on the probability distribution, and send the encoding result to the receiving end in the form of a bitstream, so that the receiving end decodes the bitstream to obtain a reconstructed image.
[0092] Entropy-encode the quantization latent space representation based on the probability distribution generated by the entropy model, convert the encoded data into the form of a bitstream and send it to the receiving end. After receiving the bitstream, the receiving end uses the same probability distribution as the sending end for decoding, reconstructs the quantization latent space representation, and generates a reconstructed image.
[0093] In some embodiments, the quantization latent space representation is divided into multiple slices along the channel dimension.
[0094] Then, the quantized latent space representation is encoded based on the probability distribution, including: for each slice, encoding the slice based on the probability distribution of the slice.
[0095] The encoding result is sent to the receiving end in the form of a bitstream, so that the receiving end decodes the bitstream to obtain a reconstructed image, including:
[0096] The encoding result is sent to the receiving end in the form of a bitstream, so that the receiving end decodes each slice in multiple slices in turn based on the same probability distribution as the sending end, reconstructs the quantized latent space representation, and generates a reconstructed image.
[0097] Specifically, the decoding of each slice depends on other slices that have been decoded.
[0098] In the embodiment of the present application, a mask image is obtained by adding a mask to the target region in the original image, and global features and mask features are respectively extracted from the original image and the mask image; the global features and the mask features are fused in the channel dimension and the spatial dimension respectively, and the fusion result is fused from the channel dimension and the spatial dimension according to the fusion weight to obtain the latent space representation of the original image; the latent space representation is quantized to obtain a quantized latent space representation, and the quantized latent space representation and the latent space representation are input into a preset entropy model to obtain the probability distribution of the quantized latent space representation; the quantized latent space representation is encoded based on the probability distribution and the encoding result is sent to the receiving end in the form of a bitstream, so that the receiving end decodes the bitstream to obtain a reconstructed image. By using the mask to sparsify the original image, the sparsity of the extracted latent space features is stronger, the data volume of the latent space features is reduced, and thus the local peak signal-to-noise ratio of the reconstructed image is improved under the constraint of extremely low data volume, and the clarity of the reconstructed image at an extremely low bit rate is enhanced. In addition, fusing features from the channel dimension and the spatial dimension according to the fusion weight can reduce the information redundancy caused by mask feature extraction and improve the model performance.
[0099] As Figure 3 shown, Figure 3 is a schematic flowchart of an image reconstruction method for T2X satellite-ground cooperation provided by an embodiment of the present application. The image reconstruction method for T2X satellite-ground cooperation can be applied to a satellite-ground cooperation communication scenario with extremely low data volume constraints as Figure 2 shown, and this method can be executed by a computing device installed on a vehicle.
[0100] Perform semantic segmentation on the driver monitoring service image x, classifying each pixel point of the driver monitoring service image x into specific semantic categories, such as the driver, window, steering wheel, etc.; after completing the semantic segmentation, identify which pixel points belong to the semantic category of the driver, and thus classify the other pixel points into the non-driver semantic category; determine the area where the pixel points with the non-driver semantic category are located as the non-driver area.
[0101] Add a mask to the non-driver area to obtain the masked image x masked 。
[0102] Input the driver monitoring service image x into a pre-trained global feature extraction module (GlobalExtraction) for processing to obtain global features, and input the masked image x masked into a pre-trained masked feature extraction module (Masked Extraction) for processing to obtain masked features.
[0103] Among them, the residual blocks "Residual,N,↑ / ↓" in the global feature extraction module and the masked feature extraction module contain the following components:
[0104] 1. Conv3×3,N,↑ / ↓: Represents a 3×3 convolutional layer, where "N" represents the number of output channels, "↑" represents upsampling, and "↓" represents downsampling.
[0105] 2. GELU: GELU (Gaussian Error Linear Unit) represents the Gaussian error linear unit, which is an activation function used to introduce non-linearity.
[0106] 3. Conv3×3,N, / : Represents a 3×3 convolutional layer, where "N" represents the number of output channels, and " / " is a parameter separator.
[0107] 4. GDN / IGDN: GDN (Generalized Divisive Normalization) represents generalized divisive normalization. By transforming the image data into a form closer to the Gaussian distribution, it reduces the correlation between different data points. Gaussianization is crucial for image data compression because the Gaussian distribution has simpler statistical properties, and the compression model can more easily learn the structure of the data; IGDN (Inverse Generalized Divisive Normalization) represents inverse generalized divisive normalization, which is the inverse process of GDN. In the compression and decompression process, using IGDN can effectively restore the compressed Gaussianized data to the original data, ensuring the reversibility and high fidelity of the compression and decompression process.
[0108] 5. Addition operator (+): Represents a Residual Connection, which directly adds the input to the output of the module, helping to alleviate the vanishing gradient problem in deep networks and enhancing the training effect of the model.
[0109] The residual block "Residual,N" in the global feature extraction module and the masked feature extraction module contains the following components: "Conv3×3,N, / ", "GELU", and the addition operator (+). This will not be elaborated here.
[0110] For the masked feature extraction module, during the training phase of the model, the image regions that the model can learn (i.e., the driver region) are explicitly specified through mask constraints, guiding the learning behavior of the model during the training phase, forcing the model to learn the regional distribution information of the image, learning which regions require more bits to represent and which regions can be allocated fewer bits, thereby optimizing the bit allocation regions and ratios of the image.
[0111] Exemplarily, to implement the above mask constraints, flexible settings can be made at different levels such as the input layer or intermediate feature layers of the model. For example, in the convolutional layer, the activation values of the masked regions are suppressed through per-channel multiplication operations. Specifically, the mask is multiplied with the feature map in an element-wise multiplication manner, causing the feature values outside the masked region to be set to zero, thereby implementing the constraint on the learning behavior of the model.
[0112] After obtaining the global feature and the masked feature, there is a lack of information interaction between the global feature and the masked feature, resulting in information redundancy between the masked feature and the global feature. The global feature and the masked feature can be input into the Masked-Global Fusion module. Inside the Masked-Global Fusion module, first, the global feature and the masked feature are concatenated along the channel dimension, that is, the masked feature and the global feature are concatenated along the channel dimension; then the number of channels of the feature is adjusted to M through a 1x1 convolutional kernel, and then local spatial features are extracted through a 3×3 depthwise separable convolution, and positional encoding information is added. Finally, through the residual Swin Transformer v2 module, the features are adaptively fused from the channel dimension and the spatial dimension according to the fusion weights to obtain the latent space representation y of the driver monitoring service image x.
[0113] Among them, the residual Swin Transformer v2 module contains the following components: Feature Embedding FE, multiple Swin Transformer v2, Feature Unembedding FU, and the addition operator (+).
[0114] Input the latent space representation y into the Entropy Model, and pass it through the hyperprior analysis transformation network h inside the entropy model a to obtain the auxiliary information z, and after quantization, obtain the quantized auxiliary information Encode the quantized auxiliary information through the arithmetic encoder AE On the one hand, send the encoding result to the receiving end in the form of a bitstream. On the other hand, the sending end itself decodes through the arithmetic decoder AD, and then passes it through the hyperprior synthesis transformation network h s to obtain the hyperprior context information
[0115] Quantize the latent space representation y into an integer to obtain a discrete form of the latent space representation, that is, the quantized latent space representation For the quantized latent space representation Equally divide it along the channel dimension to obtain slices
[0116] Input all slices into the entropy model, so that the entropy model sequentially processes each slice in series to obtain the probability distribution of each slice. Specifically, for the current slice Through the Channel-wise Context, Local Context, Intra-Global Context, and Inter-Global Context inside the entropy model, obtain the current slice of the channel context information, local context information, intra-slice global context information, and inter-slice global context information. For the current slice of the channel context information, local context information, intra-slice global context information, and inter-slice global context information, as well as the hyperprior context information, input it into the probability prediction sub-network g ep to generate the probability distribution of the current slice
[0117] For each slice, encode the slice based on the probability distribution of the slice. Finally, send the encoding result to the receiving end in the form of a bitstream
[0118] The receiving end decodes each slice in turn based on the same probability distribution as the sending end to reconstruct the quantized latent space representation and generate the reconstructed image x
[0119] Regarding decoding, the receiving end is configured with an entropy model, which has the same model parameters as the entropy model used by the sending end. The receiving end can receive the auxiliary information The corresponding bitstream is decoded by the arithmetic decoder AD inside the entropy model, and then passed through the hyperprior synthesis transformation network h s to obtain hyperprior context information.
[0120] For the current slice, channel context information, local context information, intra-slice global context information, and inter-slice global context information are generated based on other slices that have been decoded. For the generation methods of the above context information, refer to S2032, which will not be elaborated here. It should be noted that the anchor part of the current slice used in the process of generating the above context information can be generated based on the hyperprior context information and the channel context information.
[0121] The hyperprior context information, channel context information, local context information, intra-slice global context information, and inter-slice global context information are input into the probability prediction sub-network g ep , to generate the probability distribution of the current slice, and the current slice is decoded based on the probability distribution of the current slice.
[0122] In the embodiments of the present application, the original image is sparsified through a mask, so that the sparsity of the extracted latent space features is stronger, the data volume of the latent space features is reduced, and thus the local peak signal-to-noise ratio of the reconstructed image is improved under the constraint of extremely low data volume, and the clarity of the reconstructed image at an extremely low bit rate is enhanced. In addition, features are fused from the channel dimension and the spatial dimension according to the fusion weight, so that the information redundancy caused by mask feature extraction can be reduced, and the model performance can be improved.
[0123] The image reconstruction method for T2X satellite-ground cooperation in the present application is applied to a satellite-ground cooperation communication scenario with extremely low data volume constraints as shown in Figure 2 . Assuming that the resolution of the original service image sent by the driver monitoring service is 720P, the local region peak signal-to-noise ratio (Peak signal to noise ratio, PNSR) of the reconstructed image is simulated at different bit rates (unit: bit per pixel, bpp), and the performance of this scheme is compared with the scheme of global feature extraction without mask guidance training (MGALIC w / omask), as well as the existing MLIC, ELIC, WACNN, and STF schemes. The comparison results are as shown in Figure 4 . The MGALIC curve marked with an asterisk is the simulation result of this technical scheme. It can be seen that this technical scheme can effectively reduce the data volume of the image latent space features and improve the local PSNR of the reconstructed image.
[0124] As shown in Figure 5 Figure 5Visual comparisons of the latent space features and bit allocation for this technical solution and other solutions were conducted. It can be seen that, compared with other solutions, the sparsity of the latent space features extracted by this technical solution is stronger. In addition, the latent space features extracted by this solution can effectively eliminate the spatial correlation of the features and reduce the spatial information redundancy of the features. Compared with other solutions, the bit allocation area of this solution is more concentrated and efficient.
[0125] As Figure 6 shown, Figure 6 the role of mask-guided model training was analyzed by visualizing the latent space features and bit allocation of the 5 channels with the maximum entropy. As can be seen from Figure 6 this, the mask fully constrains the learning behavior of the model. The mask feature extraction is completely restricted within the mask constraint area. The original input is sparsified using the mask, effectively sparsifying the mask features. The global features are non-sparse and there is spatial information redundancy between the global features and the mask features. The mask-global feature fusion module effectively fuses the features in the channel dimension and spatial dimension, reduces the spatial information redundancy, and at the same time sparsifies the latent space features. The mask enables the model to learn the regional distribution information of the image during the training stage, making the bit allocation concentrated within the mask constraint area.
[0126] The above content describes the method provided by this application. Next, the device provided by this application will be described:
[0127] Please refer to Figure 7 , which is a schematic structural diagram of an image reconstruction device for T2X satellite-ground collaboration provided by an embodiment of this application.
[0128] As Figure 7 shown, the device may include:
[0129] An extraction unit 710, configured to add a mask to a target area in an original image to obtain a masked image, and respectively perform feature extraction on the original image and the masked image to obtain global features and mask features;
[0130] A fusion unit 720, configured to fuse the global features and the mask features in the channel dimension and spatial dimension respectively, and fuse the fusion results in the channel dimension and spatial dimension according to a fusion weight to obtain a latent space representation of the original image;
[0131] A processing unit 730, configured to perform quantization processing on the latent space representation to obtain a quantized latent space representation, and input the quantized latent space representation and the latent space representation into a preset entropy model to obtain a probability distribution of the quantized latent space representation;
[0132] An encoding unit 740, configured to encode the quantized latent space representation based on the probability distribution and send the encoding result to a receiving end in the form of a bitstream, so that the receiving end decodes the bitstream to obtain a reconstructed image.
[0133] Optionally, the original image is a driver monitoring image, and the extraction unit 710 is further configured to:
[0134] Perform semantic segmentation on the original image and determine an image region with a semantic category of non-driver as the target region.
[0135] Optionally, the fusion unit 720 is specifically configured to:
[0136] Perform feature concatenation and dimension adjustment on the global feature and the mask feature in the channel dimension to obtain a first feature;
[0137] Extract local spatial features from the first feature and add position encoding information.
[0138] Optionally, the processing unit 730 is specifically configured to:
[0139] Divide the quantized latent space representation into multiple slices along the channel dimension;
[0140] Input the multiple slices and the latent space representation into the entropy model, so that the entropy model sequentially performs serial processing on each slice to obtain the probability distribution of each slice;
[0141] The encoding of the quantized latent space representation based on the probability distribution includes:
[0142] For each slice, encode the slice based on the probability distribution of the slice.
[0143] Optionally, the entropy model is specifically configured to:
[0144] Generate hyperprior context information through a hyperprior network based on the input latent space representation;
[0145] For the currently input slice, generate channel context information, local context information, intra-slice global context information, and inter-slice global context information of the current slice through a context network based on the current slice and all slices input before the current slice;
[0146] Based on the hyperprior context information, the channel context information, the local context information, the intra-slice global context information, and the inter-slice global context information of the current slice, through a probability prediction sub-network g ep Generate the probability distribution of the current slice.
[0147] Optionally, the encoding result can be sent to the receiving end in the form of a bit stream, so that the receiving end decodes each of the multiple slices in sequence based on the same probability distribution as the sending end, reconstructs the quantized latent space representation, and generates a reconstructed image.
[0148] For the implementation processes of the functions and roles of each unit in the above device, please refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.
[0149] The embodiment of the present application also provides a hardware structure. Refer to Figure 8 , Figure 8 , which is the structural diagram of the electronic device provided by the embodiment of the present application. As Figure 8 shown, the hardware structure may include: a processor and a machine-readable storage medium, and the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the method disclosed in the above examples of the present application.
[0150] Based on the same application concept as the above method, the embodiment of the present application also provides a machine-readable storage medium, and a number of computer instructions are stored on the machine-readable storage medium. When the computer instructions are executed by a processor, the method disclosed in the above examples of the present application can be implemented.
[0151] Exemplarily, the above machine-readable storage medium can be any electronic, magnetic, optical or other physical storage device, which can contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium can be: RAM (Random Access Memory, random access memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or a combination thereof.
[0152] It should be noted that in this article, relational terms such as target and target are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0153] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. An image reconstruction method for T2X space-ground cooperation, characterized in that The method includes: Adding a mask to the target region in the original image to obtain a masked image, and respectively extracting features from the original image and the masked image to obtain global features and masked features; Fusing the global features and the masked features respectively in the channel dimension and the spatial dimension, and fusing the fusion results from the channel dimension and the spatial dimension according to the fusion weights to obtain the latent space representation of the original image; Performing quantization processing on the latent space representation to obtain a quantized latent space representation, and inputting the quantized latent space representation and the latent space representation into a preset entropy model to obtain the probability distribution of the quantized latent space representation; Encoding the quantized latent space representation based on the probability distribution, and sending the encoding result to the receiving end in the form of a bit stream, so that the receiving end decodes the bit stream to obtain a reconstructed image.
2. The method according to claim 1, wherein The original image is a driver monitoring image, and the method further includes: Performing semantic segmentation on the original image, and determining the image region with the semantic category of non-driver as the target region.
3. The method according to claim 1, wherein The fusing of the global features and the masked features respectively in the channel dimension and the spatial dimension includes: Performing feature concatenation and dimension adjustment on the global features and the masked features in the channel dimension to obtain a first feature; Extracting local spatial features from the first feature and adding position encoding information.
4. The method according to claim 1, wherein The inputting of the quantized latent space representation and the latent space representation into a preset entropy model to obtain the probability distribution of the quantized latent space representation includes: Dividing the quantized latent space representation into multiple slices along the channel dimension; Inputting the multiple slices and the latent space representation into the entropy model, so that the entropy model sequentially performs serial processing on each slice to obtain the probability distribution of each slice; The encoding of the quantized latent space representation based on the probability distribution includes: For each slice, encoding the slice based on the probability distribution of the slice.
5. The method according to claim 4, wherein The entropy model is specifically used for: Generating hyperprior context information through a hyperprior network based on the input latent space representation; For the currently input slice, generating channel context information, local context information, global context information within the slice, and global context information between slices of the current slice through a context network based on the current slice and all slices input before the current slice; Based on the super prior context information, as well as the channel context information, local context information, global context information within the slice, and global context information between slices of the current slice, generate the probability distribution of the current slice through the probability prediction sub-network g ep Generate the probability distribution of the current slice.
6. The method according to claim 4, characterized in that The sending of the encoding result to the receiving end in the form of a bit stream, so that the receiving end decodes the bit stream to obtain a reconstructed image, includes: Sending the encoding result to the receiving end in the form of a bit stream, so that the receiving end decodes each slice in the multiple slices in sequence based on the same probability distribution as the sending end, reconstructs the quantized latent space representation, and generates a reconstructed image.
7. An image reconstruction device for T2X space-ground cooperation, characterized in that, The apparatus includes: An extraction unit, configured to add a mask to the target region in the original image to obtain a masked image, and respectively extract features from the original image and the masked image to obtain global features and masked features; A fusion unit, configured to fuse the global feature and the mask feature in the channel dimension and the spatial dimension respectively, and fuse the fusion results in the channel dimension and the spatial dimension according to a fusion weight to obtain a latent space representation of the original image; A processing unit, configured to perform quantization processing on the latent space representation to obtain a quantized latent space representation, and input the quantized latent space representation and the latent space representation into a preset entropy model to obtain a probability distribution of the quantized latent space representation; An encoding unit, configured to encode the quantized latent space representation based on the probability distribution, and send the encoding result to a receiving end in the form of a bit stream, so that the receiving end decodes the bit stream to obtain a reconstructed image.
8. An electronic device, characterized in that, Comprising a processor and a memory, the memory stores machine-executable instructions that can be executed by the processor, and the processor is configured to execute the machine-executable instructions to implement the method according to any one of claims 1-6.
9. A machine-readable storage medium, characterized in that, Machine-executable instructions are stored in the machine-readable storage medium, and when the machine-executable instructions are executed by a processor, the method according to any one of claims 1-6 is implemented.
10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by a processor, the steps of the method according to any one of claims 1-6 are implemented.
Citation Information
Patent Citations
Image compression method and system based on adaptive channel and space window entropy model
CN116567240A
Learning image compression method and device for image sparse mask window attention
CN118368431A
Image coding method and device based on deep learning
CN119135910A
Image feature processing methods and decoding device
WO2024250872A1
Cited By
Deep learning progressive out-of-order transmission image compression method and device for satellite communication
CN120711154A
Satellite communication-oriented deep learning progressive out-of-order transmission image compression method and device
CN120711154B
Hybridflow image high-quality compression method under extremely low bit
CN121217923A