Emergency monitoring image semantic coding method, system and device based on illumination feature decoupling and medium
The emergency monitoring image semantic coding method, which decouples the image from illumination features, decomposes the target image into illumination features and content features for compression and transmission, and then reconstructs the image at the receiving end. This solves the problems of insufficient image compression efficiency and poor reconstruction quality at extremely low bit rates in existing technologies, and achieves efficient and reliable emergency monitoring image transmission and reconstruction.
Patent Information
- Application Number
- CN202511756765.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-11-27
AI Technical Summary
Existing emergency monitoring image compression methods are inefficient at extremely low bit rates, resulting in artifacts and semantic distortion in reconstructed images. Furthermore, existing LIC methods lack interpretability, making it difficult to ensure the effective transmission and reconstruction of critical information in complex mountain monitoring scenarios.
An emergency monitoring image semantic coding method that decouples images based on illumination features decomposes the target image and the reference image into illumination features and content features. It uses a latent variable encoder and an entropy coding model for compression and transmission, and performs illumination synthesis and reconstruction at the receiving end, thereby achieving efficient semantic compression and high-fidelity reconstruction of the image.
It maintains the consistency of image structural information and detailed features under extremely low bit rate conditions, has the ability to adapt to complex lighting changes, significantly improves the reconstruction accuracy and transmission efficiency of emergency monitoring images, and avoids image quality degradation caused by unstable ambient lighting.
Smart Images

Figure CN121239797B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of communication and image processing technology, and in particular to a method, system, device and medium for semantic coding of emergency monitoring images based on illumination feature decoupling. Background Technology
[0002] In recent years, natural disasters such as flash floods and wildfires have occurred frequently, characterized by their suddenness and unpredictability, often causing enormous loss of life and property. To achieve real-time disaster monitoring and emergency response, satellite cloud imagery and drone monitoring have been widely used, but these methods suffer from high costs and limited coverage. With the development of IoT technology, edge devices such as cameras have been widely deployed in remote areas, providing abundant on-site image information for emergency scenarios. However, due to insufficient communication network coverage in remote mountainous areas, real-time transmission of image data mainly relies on BeiDou satellite short message communication. While the BeiDou system boasts full-time and all-space coverage, its short message communication service has limited bandwidth and message length, placing extremely stringent requirements on image compression and transmission—that is, image encoding and transmission must be completed under extremely low bitrate conditions.
[0003] Existing image compression methods, including traditional encoders such as JPEG, JPEG2000, BPG, and VVC, while performing well within a general bitrate range, generally suffer from insufficient compression efficiency, noticeable artifacts, and semantic distortion in reconstructed images at ultra-low bitrates. In recent years, Learnable Image Compression (LIC) methods, leveraging the nonlinear modeling capabilities of deep neural networks, have demonstrated superiority in reconstructing images at extremely low bitrates. However, most existing LIC methods operate as black-box models, lacking interpretability for the extracted latent features, and are not optimized for emergency monitoring scenarios. In complex mountain monitoring footage, due to the indistinct foreground entities, semantic coding methods relying on target detection or semantic segmentation have limited applicability, making it difficult to guarantee the effective transmission and reconstruction of critical monitoring information.
[0004] Therefore, a new semantic coding method for emergency monitoring images is urgently needed. Summary of the Invention
[0005] In view of the above problems, embodiments of this application provide an emergency monitoring image semantic coding method, system, device and medium based on illumination feature decoupling, so as to overcome the above problems or at least partially solve the above problems.
[0006] In a first aspect, this application provides a semantic coding method for emergency monitoring images based on illumination feature decoupling, the method comprising:
[0007] The transmitting end decomposes the target image and the reference image in the feature domain through the illumination decomposition network in the illumination decomposition-synthesis module, and obtains the illumination features of the target image and the reference image, as well as the content features of the target image and the reference image.
[0008] The transmitting end uses the encoder in the illumination feature compression module to obtain a compressed bitstream using the illumination features of the target image and the reference image, as well as the content features of the target image and the reference image, and sends the compressed bitstream to the receiving end. The encoder includes: a latent variable encoder, a super prior encoder, and an entropy coding model.
[0009] The receiving end uses the received compressed bitstream and the decoder in the illumination feature compression module to recover the illumination features and content features of the target image. The decoder includes: a latent variable decoder, a super prior decoder, and an entropy decoding model.
[0010] The receiving end uses the illumination synthesis network in the illumination decomposition-synthesis module to reconstruct the image using the illumination features and content features recovered from the target image, thereby obtaining the reconstructed target image.
[0011] Optionally, the transmitting end decomposes the target image and the reference image in the feature domain using the illumination decomposition network in the illumination decomposition-synthesis module, to obtain the illumination features of the target image and the reference image respectively, and the content features of the target image and the reference image respectively, including:
[0012] The transmitting end performs pixel rearrangement on the first image through the pixel rearrangement module in the illumination decomposition network to obtain the pixel rearrangement result. The first image is the target image or the reference image.
[0013] The transmitting end obtains a first intermediate result by using the pixel rearrangement result through the first convolutional module, the depth convolutional module stacked N times in sequence, and the second convolutional module in the illumination decomposition network. The depth convolutional module includes a first branch and a second branch. The first branch is deployed with channel-wise convolutional units and point-wise convolutional units, and the second branch is deployed with point-wise convolutional units.
[0014] The sending end uses the segmentation module in the illumination decomposition network to uniformly segment the first intermediate result from the channel dimension to obtain a first part of intermediate result and a second part of intermediate result, and uses the first part of intermediate result as the content feature of the first image.
[0015] The sending end performs max pooling processing on the second part of the intermediate results through the max pooling module in the illumination decomposition network to obtain the illumination features of the first image.
[0016] Optionally, the receiving end uses the illumination synthesis network in the illumination decomposition-synthesis module to perform image reconstruction using the illumination features and content features recovered from the target image, to obtain a reconstructed target image of the target image, including:
[0017] The receiving end performs an element-wise multiplication operation on the illumination features and content features recovered from the target image through the element-wise multiplication module in the illumination synthesis network to obtain the element-wise multiplication result.
[0018] The receiving end obtains a second intermediate result through the third convolution module, the depth convolution module stacked N times in sequence, and the fourth convolution module in the illumination synthesis network. The depth convolution module includes a first branch and a second branch. The first branch is deployed with channel-wise convolution units and point-wise convolution units, and the second branch is deployed with point-wise convolution units. N is an integer greater than or equal to 3.
[0019] The receiving end performs a pixel inverse rearrangement operation on the second intermediate result through the pixel inverse rearrangement module in the illumination synthesis network to obtain the reconstructed target image of the target image. The pixel inverse rearrangement operation is the inverse operation of the pixel rearrangement operation.
[0020] Optionally, the transmitting end obtains a compressed bitstream by using the encoder in the illumination feature compression module, utilizing the illumination features of the target image and the reference image, as well as the content features of the target image and the reference image, respectively, including:
[0021] The transmitting end concatenates the illumination features and content features of the target image to obtain the illumination-content concatenated features of the target image, and concatenates the illumination features and content features of the reference image to obtain the illumination-content concatenated features of the reference image.
[0022] The transmitting end inputs the illumination content splicing features of the reference image and the target image into the latent variable encoder in the encoder;
[0023] The transmitting end processes the illumination content splicing features of the reference image and the target image through the fifth convolution module, the depth convolution module stacked N times in sequence, and the sixth convolution module in the latent variable encoder to obtain latent variables. The latent variables are then processed by the super prior encoder and the entropy coding model to obtain the compressed bitstream, where N is an integer greater than or equal to 3.
[0024] The receiving end uses the received compressed bitstream and, through the decoder in the illumination feature compression module, recovers the illumination features and content features from the target image, including:
[0025] The receiving end processes the received compressed bitstream through the super prior decoder and the entropy decoding model to obtain the recovered latent variables;
[0026] The receiving end processes the recovered latent variables through the seventh convolution module, pixel rearrangement module, depth convolution module stacked M times in sequence in the latent variable decoder to obtain the illumination features and content features recovered for the target image, where M is an integer greater than or equal to 6; the depth convolution module includes: a first branch and a second branch, the first branch is deployed with channel-wise convolution units and point-wise convolution units, and the second branch is deployed with point-wise convolution units.
[0027] Optionally, the emergency monitoring image semantic coding method based on illumination feature decoupling is implemented through an emergency monitoring image semantic coding model based on illumination feature decoupling. The training process of the emergency monitoring image semantic coding model based on illumination feature decoupling includes:
[0028] Multiple high-low illumination image pairs are obtained as a training set, and each high-low illumination image pair shares the same content features;
[0029] Using the training set, the illumination decomposition-synthesis module and the illumination feature compression module to be trained are trained according to the target loss function to obtain the emergency monitoring image semantic coding model based on illumination feature decoupling.
[0030] The target loss function includes at least one of the following:
[0031] The first loss function value is used to guide the illumination feature compression module to recover illumination features and content features from the sample image and align them with the illumination features and content features decomposed by the illumination decomposition network in the illumination decomposition-synthesis module, so as to minimize the difference between the image reconstructed by the illumination decomposition-synthesis network and the sample image. The sample image is a high-illumination image or a low-illumination image included in each high-low illumination image pair.
[0032] The second loss function value is used to guide the illumination decomposition network in the illumination decomposition-synthesis module to learn the same content features shared by each high-low illumination image pair.
[0033] The third loss function value is used to guide the illumination decomposition network in the illumination decomposition-synthesis module to learn the different illumination features in each high-low illumination image pair, including the high illumination image and the low illumination image.
[0034] Optionally, the target loss function further includes at least one of the following:
[0035] The fourth loss function value is used to minimize the difference between the content features recovered by the illumination feature compression module to be trained for the first high illumination image and the content features recovered for the first low illumination image. The first high illumination image and the first low illumination image are two images in a high-low illumination image pair.
[0036] Compression ratio, which is determined based on the original number of bits of the sample image and the number of bits of the compressed bitstream for the sample image, the sample image comprising either a high-light image or a low-light image for each high-low illumination image pair.
[0037] Optionally, the transmitting end obtains a compressed bitstream by using the encoder in the illumination feature compression module, utilizing the illumination features of the target image and the reference image, as well as the content features of the target image and the reference image, respectively, including:
[0038] The transmitting end compares the illumination features of the target image with the illumination features of the reference image, and compares the content features of the target image with the content features of the reference image;
[0039] The transmitting end determines whether there are any changed content features and whether there are any changed illumination features;
[0040] When the sending end determines that the target image and the reference image share the same content features, it uses the encoder in the illumination feature compression module to obtain a compressed bitstream using the changed illumination features and the same content features; or
[0041] When the transmitting end determines that there are changed content features and changed illumination features, it uses the encoder in the illumination feature compression module to obtain a compressed bitstream using the changed illumination features and the changed content features.
[0042] The method further includes:
[0043] When the target image and the reference image share the same content features, the receiving end uses the pre-cached content features recovered from the reference image as the content features recovered from the target image.
[0044] When there are changed content features, the sending end obtains the content features recovered for the target image based on the received changed content features and the pre-cached content features recovered for the reference image.
[0045] A second aspect of this application provides an emergency monitoring image semantic coding system based on illumination feature decoupling, the system comprising:
[0046] The feature domain decomposition module is used to decompose the target image and the reference image in the feature domain through the illumination decomposition network in the illumination decomposition-synthesis module, so as to obtain the illumination features of the target image and the reference image respectively, and the content features of the target image and the reference image respectively.
[0047] An illumination feature compression module is used by the transmitting end to obtain a compressed bitstream by using the illumination features of the target image and the reference image, and the content features of the target image and the reference image, respectively, through an encoder, and to send the compressed bitstream to the receiving end; and by the receiving end to recover the illumination features and content features of the target image through a decoder. The encoder includes: a latent variable encoder, a super-prior encoder, and an entropy coding model; the decoder includes: a latent variable decoder, a super-prior decoder, and an entropy decoding model.
[0048] The illumination decomposition-composite module is used to perform image reconstruction by using the illumination features and content features recovered from the target image through an illumination synthesis network, so as to obtain the reconstructed image of the target image.
[0049] A third aspect of this application provides an electronic device, including a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the steps of the emergency monitoring image semantic coding method based on illumination feature decoupling as described in the first aspect of this application.
[0050] A fourth aspect of this application provides a readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the emergency monitoring image semantic coding method based on illumination feature decoupling as described in the first aspect of this application.
[0051] The beneficial effects of this application are:
[0052] This application provides a semantic coding method for emergency monitoring images based on illumination feature decoupling. The method includes: a transmitting end decomposing a target image and a reference image in the feature domain using an illumination decomposition network in an illumination decomposition-synthesis module to obtain illumination features and content features of the target image and the reference image respectively; the transmitting end using an encoder in an illumination feature compression module to obtain a compressed bitstream using the illumination features and content features of the target image and the reference image respectively, and sending the compressed bitstream to a receiving end; the encoder includes a latent variable encoder, a super-prior encoder, and an entropy coding model; the receiving end using the received compressed bitstream and a decoder in the illumination feature compression module to recover the illumination features and content features of the target image; the decoder includes a latent variable decoder, a super-prior decoder, and an entropy decoding model; and the receiving end using an illumination synthesis network in the illumination decomposition-synthesis module to perform image reconstruction using the recovered illumination features and content features of the target image to obtain a reconstructed target image.
[0053] This application achieves efficient semantic compression and high-fidelity image reconstruction in bandwidth-constrained emergency monitoring scenarios by decomposing the target image and reference image into illumination features and content features at the transmitting end, and jointly compressing the decomposed features in the encoder to generate a compressed bitstream. The receiving end decodes the compressed bitstream and uses an illumination synthesis network to complete the reconstruction of the target image. It can maintain the consistency of image structure information and detail features under extremely low bit rate conditions, has adaptive processing capability for complex illumination changes, and effectively avoids the problem of image quality degradation caused by unstable ambient lighting. Thus, it significantly improves the reconstruction accuracy and transmission efficiency of emergency monitoring images under low bandwidth transmission conditions. Attached Figure Description
[0054] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0055] Figure 1 This is a flowchart illustrating the steps of an emergency monitoring image semantic coding method based on illumination feature decoupling provided in an embodiment of this application.
[0056] Figure 2 This is a schematic diagram of the network structure of an illumination decomposition network provided in an embodiment of this application;
[0057] Figure 3 This is a schematic diagram of the network structure of an illumination synthesis network provided in an embodiment of this application;
[0058] Figure 4 This is a schematic diagram of the network structure of a depthwise convolutional module provided in an embodiment of this application;
[0059] Figure 5 This is a schematic diagram of the network structure of a latent variable encoder provided in an embodiment of this application;
[0060] Figure 6 This is a schematic diagram of the network structure of a latent variable decoder provided in an embodiment of this application;
[0061] Figure 7 This is a schematic diagram of a network structure for a priori analysis transform and composition transform provided in an embodiment of this application;
[0062] Figure 8 This is a schematic diagram of the overall structure of an emergency monitoring image semantic coding model based on illumination feature decoupling provided in an embodiment of this application;
[0063] Figure 9 This is a schematic diagram of the network structure of an entropy coding model provided in an embodiment of this application;
[0064] Figure 10 This is a schematic diagram of an emergency monitoring image semantic coding system based on illumination feature decoupling provided in an embodiment of this application;
[0065] Figure 11 This is a schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0066] Exemplary embodiments of this application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.
[0067] In a first aspect, this application provides an emergency monitoring image semantic coding method based on illumination feature decoupling, such as... Figure 1 As shown, the method includes:
[0068] S101, the transmitting end decomposes the target image and the reference image in the feature domain through the illumination decomposition network in the illumination decomposition-synthesis module, and obtains the illumination features of the target image and the reference image respectively, as well as the content features of the target image and the reference image respectively.
[0069] In this step, the sending end first needs to acquire the target image and the reference image, and then input the acquired target image and reference image into the illumination decomposition network in the illumination decomposition-synthesis module. In some cases, before inputting the target image and reference image into the illumination decomposition network, it is also necessary to perform operations such as pixel block rearrangement on the target image and reference image.
[0070] Subsequently, the processed target image and reference image are decomposed by an illumination decomposition network to obtain the illumination features of the target image and the reference image, as well as the content features of the target image and the reference image.
[0071] S102, the transmitting end uses the encoder in the illumination feature compression module to obtain a compressed bitstream using the illumination features of the target image and the reference image, as well as the content features of the target image and the reference image, and sends the compressed bitstream to the receiving end. The encoder includes: a latent variable encoder, a super prior encoder, and an entropy coding model.
[0072] In this step, the sending end takes the illumination features and content features obtained from the illumination decomposition network of the target image and the reference image as input and inputs them into the encoder in the illumination feature compression module. The encoder in the illumination feature compression module performs nonlinear mapping and compression representation on the above illumination features and content features to obtain the corresponding latent variables. Subsequently, the latent variables are processed by the super-prior encoder and the entropy coding model to obtain the compressed bitstream. The compressed bitstream contains compressed information of the illumination features and content features of both the target image and the reference image.
[0073] Finally, the sending end transmits the compressed bit stream to the receiving end through the communication link for subsequent decoding and image reconstruction. In this application, the encoder in the illumination feature compression module includes: a latent variable encoder, a super-prior encoder, and an entropy coding model.
[0074] S103, the receiving end uses the received compressed bitstream and the decoder in the illumination feature compression module to recover the illumination features and content features of the target image. The decoder includes: a latent variable decoder, a priori decoder, and an entropy decoding model.
[0075] In this step, the receiving end receives the compressed bitstream transmitted by the sending end and inputs it into the decoder in the illumination feature compression module. First, the received compressed bitstream is processed by a priori decoder and an entropy decoding model to obtain the recovered latent variables. Then, the recovered latent variables are decoded by a latent variable decoder to obtain the illumination features and content features of the target image in the feature domain. During the decoding process, the decoder can reconstruct and restore the features in the latent space, thereby obtaining the illumination features and content features of the target image in the feature domain. These decoded illumination features reflect the illumination conditions in the monitoring scene, while the content features preserve the structural and semantic information of the image. In this application, the decoder includes a latent variable decoder, a priori decoder, and an entropy decoding model.
[0076] S104, the receiving end uses the illumination synthesis network in the illumination decomposition-synthesis module to perform image reconstruction using the illumination features and content features recovered from the target image, and obtains the reconstructed target image of the target image.
[0077] In this step, after obtaining the illumination features and content features of the target image, the receiving end inputs both into the illumination synthesis network in the illumination decomposition-synthesis module. In some cases, the illumination synthesis network adopts a structure symmetrical to the illumination decomposition network, which can perform element-wise fusion of illumination features and content features, effectively combining illumination information and content information through pointwise multiplication and nonlinear mapping operations.
[0078] After fusion, the illumination synthesis network further transforms and reconstructs the processed results, gradually restoring the representation of the target image in the pixel domain. The final image not only realistically reflects the structure and semantic content of the original scene but also maintains the consistency of illumination conditions, thereby generating a reconstructed target image that is highly similar to the original input. This provides reliable support for image transmission in emergency monitoring scenarios. For monitoring scenarios under low-light conditions (such as nighttime monitoring), the network can input content features extracted from the target image and reference illumination features extracted from a normally illuminated image (i.e., the reference image) to generate a visually attractive image with improved illumination conditions.
[0079] The proposed semantic coding method for emergency monitoring images based on illumination feature decoupling decomposes the target image and reference image into illumination features and content features at the transmitting end. The encoder then performs joint compression of the decomposed features to generate a compressed bitstream. At the receiving end, the compressed bitstream is decoded and the target image is reconstructed using an illumination synthesis network. This method achieves efficient semantic compression and high-fidelity image reconstruction in bandwidth-constrained emergency monitoring scenarios. It also maintains the consistency of image structural information and detailed features under extremely low bitrate conditions and has adaptive processing capabilities for complex illumination changes. This effectively avoids the image quality degradation problem caused by unstable ambient lighting, thereby significantly improving the reconstruction accuracy and transmission efficiency of emergency monitoring images under low bandwidth transmission conditions.
[0080] In one embodiment, the transmitting end decomposes the target image and the reference image in the feature domain using the illumination decomposition network in the illumination decomposition-synthesis module, respectively, to obtain the illumination features of the target image and the reference image, and the content features of the target image and the reference image, including:
[0081] The transmitting end performs pixel rearrangement on the first image through the pixel rearrangement module in the illumination decomposition network to obtain the pixel rearrangement result. The first image is the target image or the reference image.
[0082] The transmitting end obtains a first intermediate result by using the pixel rearrangement result through the first convolutional module, the depth convolutional module stacked N times in sequence, and the second convolutional module in the illumination decomposition network. The depth convolutional module includes a first branch and a second branch. The first branch is deployed with channel-wise convolutional units and point-wise convolutional units, and the second branch is deployed with point-wise convolutional units.
[0083] The sending end uses the segmentation module in the illumination decomposition network to uniformly segment the first intermediate result from the channel dimension to obtain a first part of intermediate result and a second part of intermediate result, and uses the first part of intermediate result as the content feature of the first image.
[0084] The sending end performs max pooling processing on the second part of the intermediate results through the max pooling module in the illumination decomposition network to obtain the illumination features of the first image.
[0085] In this embodiment, refer to as follows Figure 2 The diagram shows the network structure of the illumination decomposition network. The transmitting end uses the illumination decomposition network in the illumination decomposition-synthesis module to decompose the target image and reference image in the feature domain, obtaining their respective illumination features and content features. Specifically, the steps include:
[0086] First, the sending end performs pixel rearrangement processing on the input first image through the pixel rearrangement module in the illumination decomposition network to obtain the pixel rearrangement result. In this embodiment, the first image can be either the target image or a reference image. In some cases, the pixel rearrangement operation can use a preset downsampling factor to locally rearrange the original image, reorganizing the spatial information of adjacent pixel blocks into different channels, so that the subsequent network structure can extract deep feature information more efficiently.
[0087] Furthermore, the sending end inputs the pixel rearrangement result into the first convolutional module, the depthwise convolutional module stacked N times in sequence, and the second convolutional module in the illumination decomposition network to obtain the first intermediate result. The depthwise convolutional module includes a first branch and a second branch: the first branch includes channel-wise convolutional units and pointwise convolutional units, used to extract spatial and channel features respectively; the second branch includes pointwise convolutional units, used to further improve feature representation capabilities. Through multi-layer convolution and non-linear feature extraction operations, the structural and semantic information of the input image can be fully extracted. In this embodiment, N is an integer greater than or equal to 3.
[0088] Furthermore, the transmitting end uses the segmentation module in the illumination decomposition network to uniformly segment the first intermediate result from the channel dimension, obtaining a first part of the intermediate result and a second part of the intermediate result. The first part of the intermediate result is used to characterize the content features of the first image, while in some cases, the second part of the intermediate result contains feature information related to illumination changes in the image.
[0089] Finally, the sending end performs max pooling on the intermediate results of the second part through the max pooling module in the illumination decomposition network to extract the illumination features of the first image. In this embodiment, the max pooling operation can effectively compress the feature dimension, highlight key information such as illumination intensity and distribution, and reduce redundant content in the illumination features, thereby providing a more compact and efficient feature representation for subsequent encoding and compression.
[0090] In one embodiment, the receiving end uses the illumination synthesis network in the illumination decomposition-synthesis module to perform image reconstruction using the illumination features and content features recovered from the target image, to obtain a reconstructed target image of the target image, including:
[0091] The receiving end performs an element-wise multiplication operation on the illumination features and content features recovered from the target image through the element-wise multiplication module in the illumination synthesis network to obtain the element-wise multiplication result.
[0092] The receiving end obtains a second intermediate result through the third convolution module, the depth convolution module stacked N times in sequence, and the fourth convolution module in the illumination synthesis network. The depth convolution module includes a first branch and a second branch. The first branch is deployed with channel-wise convolution units and point-wise convolution units, and the second branch is deployed with point-wise convolution units. N is an integer greater than or equal to 3.
[0093] The receiving end performs a pixel inverse rearrangement operation on the second intermediate result through the pixel inverse rearrangement module in the illumination synthesis network to obtain the reconstructed target image of the target image. The pixel inverse rearrangement operation is the inverse operation of the pixel rearrangement operation.
[0094] In this embodiment, refer to as follows Figure 3 The diagram shows the network structure of the illumination synthesis network. At the receiving end, the illumination synthesis network in the illumination decomposition-synthesis module uses the illumination features and content features recovered from the target image to perform image reconstruction, obtaining the reconstructed target image. Specifically, it includes:
[0095] First, the receiving end performs element-wise multiplication on the recovered illumination features and content features of the target image through the element-wise multiplication module in the illumination synthesis network to obtain the element-wise multiplication result. In this embodiment, the element-wise multiplication operation can be based on Retinex theory to effectively fuse illumination information and content information, laying the foundation for subsequent image reconstruction.
[0096] Furthermore, the receiving end sequentially inputs the element-wise product result into the third convolutional module cascaded in series, the depthwise convolutional module stacked N times in sequence, and the fourth convolutional module in the illumination synthesis network to obtain the second intermediate result. The depthwise convolutional module includes two branches: the first branch consists of channel-wise convolutional units and pointwise convolutional units, used to simultaneously capture spatial and channel features; the second branch includes pointwise convolutional units, used to further enhance feature representation capabilities. Through multi-layer convolution and non-linear feature extraction, the structural information and detailed features of the original image can be effectively recovered, where N is an integer greater than or equal to 3.
[0097] Furthermore, the receiving end performs a pixel inverse rearrangement operation on the second intermediate result through the pixel inverse rearrangement module in the illumination synthesis network to obtain the reconstructed target image of the target image. The pixel inverse rearrangement operation is the inverse operation of the pixel rearrangement operation at the sending end, which can remap the recombined information in the feature domain back to the pixel domain, thereby obtaining a reconstruction result consistent with the structure of the input image.
[0098] In this embodiment, through the aforementioned illumination synthesis process, the receiving end can not only recover a near-original high-quality image under extremely low bitrate conditions, but also maintain the structural integrity and illumination consistency of the image, thus improving the visibility of nighttime surveillance images. Furthermore, the element-wise multiplication operation effectively fuses illumination features and content features, the convolutional module and the depthwise convolutional module ensure the detail fidelity of the image reconstruction, and the pixel inverse rearrangement operation further ensures the correspondence between the final image and the original input in pixel space.
[0099] In one embodiment, reference is made to, as shown in... Figure 4 The diagram shows the network structure of the deep convolutional module, as follows: Figure 4 As shown, the deep convolutional module consists of two parallel branches, which are used to capture the spatial correlation and channel correlation of the input features, respectively, so as to improve the network's multi-scale feature representation ability and feature reconstruction accuracy.
[0100] Specifically, the left branch consists of a first convolutional layer, a ReLU activation unit, a grouped convolutional layer, and a second convolutional layer, all connected in series. The first convolutional layer uses a 1×1 kernel for linear channel mapping and compression of the input features. The grouped convolutional layer employs a depthwise separable convolutional structure with a 3×3 kernel, significantly reducing network parameters and computational complexity while preserving local spatial information. The second convolutional layer also uses a 1×1 convolution for feature fusion and channel restoration. This sequential structure of the left branch effectively extracts local spatial patterns and enhances the inter-layer coupling of features.
[0101] The right branch consists of a third convolutional layer, a segmentation module, a ReLU activation unit, and a fourth convolutional layer, all connected in series. The segmentation module segments and reassembles the input feature channels, achieving feature partitioning. Subsequently, element-wise summation is used to fuse features between branches, enhancing cross-channel context awareness. The ReLU activation unit and the fourth convolutional layer further optimize the non-linear representation and dimensionality mapping of the features, significantly improving feature extraction capabilities while maintaining computational efficiency.
[0102] In one embodiment, the transmitting end obtains a compressed bitstream by using the encoder in the illumination feature compression module, utilizing the illumination features of the target image and the reference image, as well as the content features of the target image and the reference image, respectively, including:
[0103] The transmitting end concatenates the illumination features and content features of the target image to obtain the illumination-content concatenated features of the target image, and concatenates the illumination features and content features of the reference image to obtain the illumination-content concatenated features of the reference image.
[0104] The transmitting end inputs the illumination content splicing features of the reference image and the target image into the latent variable encoder in the encoder;
[0105] The transmitting end processes the illumination content splicing features of the reference image and the target image through the fifth convolution module, the depth convolution module stacked N times in sequence, and the sixth convolution module in the latent variable encoder to obtain latent variables. The latent variables are then processed by the super prior encoder and the entropy coding model to obtain the compressed bitstream, where N is an integer greater than or equal to 3.
[0106] The receiving end uses the received compressed bitstream and, through the decoder in the illumination feature compression module, recovers the illumination features and content features from the target image, including:
[0107] The receiving end processes the received compressed bitstream through the super prior decoder and the entropy decoding model to obtain the recovered latent variables;
[0108] The receiving end processes the recovered latent variables through the seventh convolution module, pixel rearrangement module, depth convolution module stacked M times in sequence in the latent variable decoder to obtain the illumination features and content features recovered for the target image, where M is an integer greater than or equal to 6; the depth convolution module includes: a first branch and a second branch, the first branch is deployed with channel-wise convolution units and point-wise convolution units, and the second branch is deployed with point-wise convolution units.
[0109] In this embodiment, the transmitting end uses the encoder in the illumination feature compression module to generate a compressed bitstream using the illumination features and content features of the target image and the reference image, and then sends it to the receiving end. Specifically, this includes the following steps:
[0110] The sending end concatenates the illumination features and content features of the target image along the channel dimension to obtain the illumination-content concatenated features of the target image. Simultaneously, the same concatenation operation is performed on the illumination features and content features of the reference image to obtain the illumination-content concatenated features of the reference image. Through these concatenation operations, illumination and content information can be fused in a unified feature space, providing input for subsequent learning of latent variable representations.
[0111] Furthermore, the transmitting end inputs the illumination content splicing features of the target image and the reference image into the latent variable encoder. In this embodiment, the network structure diagram of the latent variable encoder is shown below. Figure 5As shown, the latent variable encoder sequentially includes a fifth convolutional module, a depthwise convolutional module stacked N times, and a sixth convolutional module, where N is an integer greater than or equal to 3. In this embodiment, the depthwise convolutional module in the latent variable encoder also consists of two branches: the first branch includes channel-wise convolutional units and pointwise convolutional units, used to extract spatial features and channel features respectively; the second branch includes pointwise convolutional units, used to supplement feature representation capabilities. Through the above multi-layer convolution and nonlinear processing, the latent variable representation is obtained. Then, the sending end further performs entropy modeling and compression encoding on the latent variables using a super-prior encoder and an entropy coding model to generate the final compressed bitstream, which is then sent to the receiving end.
[0112] At the receiving end, the received compressed bitstream is used to recover the illumination and content features of the target image through the decoder in the illumination feature compression module. The receiving end first decodes the compressed bitstream to obtain the recovered latent variables. Specifically, the received compressed bitstream is processed by the super-prior decoder and entropy decoding model in the illumination feature compression module to obtain the recovered latent variables. Then, the recovered latent variables are input into the latent variable decoder. The network structure diagram of the latent variable decoder is shown below. Figure 6 As shown, the latent variable decoder sequentially includes a seventh convolutional module, a pixel rearrangement module, a depthwise convolutional module stacked M times, and an eighth convolutional module, where M is an integer greater than or equal to 6. The latent variable decoder recovers the illumination and content features of the target image. Similarly, the depthwise convolutional module in the latent variable decoder also includes two branches: the first branch consists of channel-wise convolutional units and pointwise convolutional units, used to efficiently extract multi-layer features of latent variables; the second branch includes pointwise convolutional units, used to enhance feature reconstruction capabilities.
[0113] This embodiment obtains a compressed bitstream by concatenating the illumination features and content features of the target image and the reference image at the sending end and inputting them into the encoder. This allows for the full fusion of scene structure information and illumination information in a unified feature space, avoiding redundancy caused by the independent processing of illumination and content in traditional methods.
[0114] At the receiving end, the compressed bitstream is decoded and the pixels are reverse rearranged by the decoder, which can not only accurately recover the illumination features and content features of the target image, but also ensure the consistency of features in the spatial domain and semantic domain.
[0115] In one embodiment, the emergency monitoring image semantic coding method based on illumination feature decoupling is implemented through an emergency monitoring image semantic coding model based on illumination feature decoupling. The training process of the emergency monitoring image semantic coding model based on illumination feature decoupling includes:
[0116] Multiple high-low illumination image pairs are obtained as a training set, and each high-low illumination image pair shares the same content features;
[0117] Using the training set, the illumination decomposition-synthesis module and the illumination feature compression module to be trained are trained according to the target loss function to obtain the emergency monitoring image semantic coding model based on illumination feature decoupling.
[0118] The target loss function includes at least one of the following:
[0119] The first loss function value is used to guide the illumination feature compression module to recover illumination features and content features from the sample image and align them with the illumination features and content features decomposed by the illumination decomposition network in the illumination decomposition-synthesis module, so as to minimize the difference between the image reconstructed by the illumination decomposition-synthesis network and the sample image. The sample image is a high-illumination image or a low-illumination image included in each high-low illumination image pair.
[0120] The second loss function value is used to guide the illumination decomposition network in the illumination decomposition-synthesis module to learn the same content features shared by each high-low illumination image pair.
[0121] The third loss function value is used to guide the illumination decomposition network in the illumination decomposition-synthesis module to learn the different illumination features in each high-low illumination image pair, including the high illumination image and the low illumination image.
[0122] In this embodiment, the emergency monitoring image semantic coding method based on illumination feature decoupling is implemented through an emergency monitoring image semantic coding model based on illumination feature decoupling. The training process of this model includes the following steps:
[0123] Multiple pairs of images with high and low illumination are obtained as a training set. Each pair of high and low illumination images maintains consistency in scene content, that is, they share the same content features, and differ only in illumination conditions.
[0124] The aforementioned training set was used to jointly train the illumination decomposition-synthesis module and the illumination feature compression module. During training, the network parameters were optimized according to a preset target loss function, thereby obtaining a complete semantic coding model for emergency monitoring images based on illumination feature decoupling.
[0125] In this embodiment, the target loss function includes at least one of the following:
[0126] The first loss function value guides the illumination feature compression module to align the illumination features and content features recovered from the sample image with the illumination features and content features decomposed by the illumination decomposition network, while minimizing the difference between the image reconstructed by the illumination synthesis network and the sample image. In this embodiment, the sample image is any high-illumination image or low-illumination image from each pair of high- and low-illumination images in the training set.
[0127] The value of the first loss function is shown in the following formula:
[0128] (1)
[0129] in, This represents the original input specular image; This represents the original low-light image as input; This represents the content features of a highlight image; Indicates the content features of a low-light image; This represents the lighting characteristics of a highlight image; This represents the lighting characteristics of a low-light image.
[0130] The second loss function value is used to guide the illumination decomposition network in the illumination decomposition-synthesis module to learn the same content features shared by each high-low illumination image pair.
[0131] The value of the second loss function is shown in the following formula:
[0132] (2)
[0133] The third loss function value is used to guide the illumination decomposition network in the illumination decomposition-synthesis module to learn the different illumination features in each high-low illumination image pair, including the high illumination image and the low illumination image.
[0134] The value of the third loss function is shown in the following formula:
[0135] (3)
[0136] This embodiment, through the aforementioned training process, can distinguish and decouple illumination features from content features within the feature domain, ensuring accurate representation of scene structure and semantics under different lighting conditions. The joint constraints of multiple loss functions guarantee high fidelity of the reconstructed image while improving the separability of illumination and content features. This enables efficient compression and high-quality reconstruction under extremely low bitrate conditions in emergency monitoring scenarios, significantly enhancing the model's applicability and robustness in complex lighting environments.
[0137] In one embodiment, the target loss function further includes at least one of the following:
[0138] The fourth loss function value is used to minimize the difference between the content features recovered by the illumination feature compression module to be trained for the first high illumination image and the content features recovered for the first low illumination image. The first high illumination image and the first low illumination image are two images in a high-low illumination image pair.
[0139] Compression ratio, which is determined based on the original number of bits of the sample image and the number of bits of the compressed bitstream for the sample image, the sample image comprising either a high-light image or a low-light image for each high-low illumination image pair.
[0140] In this embodiment, the target loss function further includes:
[0141] The fourth loss function value is used to minimize the difference between the content features extracted by the illumination feature compression module during the restoration process for the two images in the same high- and low-illumination image pair. Specifically, for a high- and low-illumination image pair, the first high-illumination image and the first low-illumination image should be consistent or aligned in terms of scene content. Therefore, by constraining the content features recovered from the first high-illumination image to be as close as possible to or aligned with the content features recovered from the first low-illumination image, the model's ability to extract stable and consistent content features under different illumination conditions can be further enhanced.
[0142] The value of the fourth loss function is shown in the following formula:
[0143] (4)
[0144] Compression ratio is used to measure the bitrate performance of a model during the compression phase. It is calculated based on the number of bits in a sample image and the number of bits in the corresponding compressed bitstream. The sample images are either the high-light or low-light image from each pair of high-light / low-light images in the training set. By incorporating the compression ratio metric into the objective loss function, the size of the compressed bitstream can be effectively reduced while maintaining image reconstruction quality, thereby improving compression efficiency.
[0145] The compression ratio is shown in the following formula:
[0146] (5)
[0147] in, Indicates bitrate; Representing an image compressed bitstream Length; Representing an image Height; Representing an image The width.
[0148] In this application, in order to optimize the distortion metric, ensure the fidelity of the reconstructed image, and maintain the model's reconstruction capability, the first loss function value, the second loss function value, the third loss function value, and the fourth loss function value are fused to obtain the cross-image reconstruction loss function as shown in the following formula:
[0149] (6)
[0150] In one embodiment, the transmitting end obtains a compressed bitstream by using the encoder in the illumination feature compression module, utilizing the illumination features of the target image and the reference image, as well as the content features of the target image and the reference image, respectively, including:
[0151] The transmitting end compares the illumination features of the target image with the illumination features of the reference image, and compares the content features of the target image with the content features of the reference image;
[0152] The transmitting end determines whether there are any changed content features and whether there are any changed illumination features;
[0153] When the sending end determines that the target image and the reference image share the same content features, it uses the encoder in the illumination feature compression module to obtain a compressed bitstream using the changed illumination features and the same content features; or
[0154] When the transmitting end determines that there are changed content features and changed illumination features, it uses the encoder in the illumination feature compression module to obtain a compressed bitstream using the changed illumination features and the changed content features.
[0155] The method further includes:
[0156] When the target image and the reference image share the same content features, the receiving end uses the pre-cached content features recovered from the reference image as the content features recovered from the target image.
[0157] When there are changed content features, the sending end obtains the content features recovered for the target image based on the received changed content features and the pre-cached content features recovered for the reference image.
[0158] In this embodiment, the transmitting end generates a compressed bitstream using the encoder in the illumination feature compression module, utilizing the illumination features and content features of the target image and the reference image respectively. Specifically, this includes the following steps:
[0159] First, the sending end compares the illumination features of the target image and the reference image, and at the same time compares the content features of the target image and the reference image to determine the differences between them in terms of illumination conditions and content structure.
[0160] Furthermore, based on the comparison results, the sending end determines whether there are changes in the content features and the illumination features between the target image and the reference image.
[0161] When the target image and the reference image share the same content features, it indicates that the main structure of the image scene has not changed. At this time, the sending end generates a compressed bitstream based on the changed illumination features in the target image and the same content features shared by both, thereby avoiding unnecessary repeated encoding of content features, reducing data redundancy, and realizing efficient semantic compression and high-fidelity image reconstruction in bandwidth-constrained emergency monitoring scenarios. Moreover, it can maintain the consistency of image structural information and detailed features under extremely low bitrate conditions.
[0162] When both the target image and the reference image have changes in content features and illumination features, it indicates that the image has changed in both illumination conditions and scene content. In this case, the sending end generates a compressed bit stream based on the changed illumination features and the changed content features to ensure that the transmitted data can fully cover the new scene information.
[0163] In addition, the method also includes the following process: at the receiving end, when the target image and the reference image share the same content features, the receiving end does not need to re-decode the content features of the target image, but directly calls and uses the pre-cached content features of the reference image as the content features of the target image, thereby reducing the decoding overhead.
[0164] Furthermore, when the target image exhibits changes in its content features, the receiving end fuses or updates the received changed content features with the pre-cached reference image content features to obtain the corresponding content features of the target image. This ensures the timeliness and accuracy of the content features while reducing the amount of duplicate data transmitted.
[0165] In this embodiment, through the above processing method, the sending end can selectively compress and transmit illumination features and content features according to the differences between the target image and the reference image, thereby avoiding redundant information transmission and improving compression efficiency. At the same time, the receiving end uses the cached reference image features to participate in the feature recovery of the target image, further reducing the bit rate requirement. It can significantly reduce the size of the compressed bit stream while ensuring the image reconstruction quality, making it suitable for bandwidth-constrained emergency monitoring communication scenarios.
[0166] In one embodiment, such as Figure 7 The diagram shows a schematic representation of the network structure for the advanced prior analysis transformation and the composition transformation.
[0167] Specifically, such as Figure 7 As shown, a is a schematic diagram of the network structure of the super-prior analysis transform, and b is a schematic diagram of the network structure of the super-prior synthesis transform. The network structure of the super-prior analysis transform includes multiple convolutional modules and depthwise convolutional modules connected in series. The kernel size and stride of the first convolutional module are both set to 1, the kernel size and compensation of the third and fifth convolutional modules are both set to 2, and the second, fourth and sixth convolutional modules are depthwise convolutional modules. The network structure of the super-prior synthesis transform includes a convolutional module, a pixel rearrangement module, and a depthwise convolutional module connected in series. The kernel size and stride of the first and fourth convolutional modules are both set to 1. The second and fifth modules are pixel rearrangement modules, and the third and sixth modules are depthwise convolutional modules. The depthwise convolutional module in the sixth position is a depthwise convolutional module stacked three times in sequence. In this embodiment, by introducing the network structure of super-prior analysis transform and synthesis transform, this embodiment not only improves the accuracy and stability of illumination feature compression, but also significantly reduces the bit rate requirement while ensuring reconstruction quality. This achieves efficient and low-latency semantic compression transmission, which is particularly suitable for real-time image transmission tasks with limited bandwidth in emergency monitoring scenarios.
[0168] In an alternative embodiment, to address the computational intensiveness and slow speed issues faced by the model when processing large images and high-dimensional data, this application also introduces a chessboard-based context model. This model can significantly improve encoding and decoding speed while maintaining near-optimal encoding performance. This context model uses a chessboard pattern to divide latent variables into two complementary subsets: anchor points and non-anchor points. Anchor point positions consist of elements in the latent variables whose row and column indices sum to an odd number, and their mean and scale distribution parameters are predicted solely from global probability distribution parameter features by a parametric inference network. Non-anchor point positions consist of elements whose row and column indices sum to an even number, and their mean and scale distribution parameters are derived by combining parametric features extracted from the prior variables and the output of the context prediction model. In this process, using the elements of the anchor point positions in the latent variables as input, the distribution parameters of the complete latent variables can be derived by fusing the mean and scale information of the anchor and non-anchor points, and the obtained distribution parameters are then input into the Gaussian conditional model.
[0169] In the decoding phase, the model employs a two-stage parallel decoding strategy. The first stage independently decodes the latent elements at anchor positions based on global parameter features. At this stage, the decoding process does not depend on the values of neighboring latent elements, thus achieving efficient parallel processing of anchor positions and significantly improving decoding speed. The second stage, building upon this foundation, decodes latent elements at non-anchor positions. During decoding, non-anchor positions can access local contextual information from surrounding anchor positions to fully utilize spatial dependencies, thereby ensuring the accuracy and consistency of latent variable reconstruction.
[0170] In one embodiment, such as Figure 8 The diagram shown illustrates the overall structure of the semantic coding model for emergency monitoring images based on illumination feature decoupling provided in this application. The model mainly includes an illumination decomposition-synthesis module and an illumination feature compression module. The illumination decomposition-synthesis module includes an illumination encoder-decoder, which comprises an illumination decomposition network and an illumination synthesis network. The illumination feature compression module includes a latent variable encoder-decoder and a super-prior encoder-decoder. The latent variable encoder-decoder comprises a latent analysis transform network and a latent synthesis transform network. The super-prior encoder-decoder comprises a super-prior analysis transform network and a super-prior synthesis transform network.
[0171] Specifically, firstly, the input target image and reference image are fed into an illumination decomposition network, where illumination features and content features are separated within the feature domain. Next, the illumination features and content features from the illumination decomposition network are input into an illumination feature compression module for joint feature compression. This module includes a latent analysis transformation network, used to map high-dimensional features to low-dimensional latent variables, achieving dimensionality reduction and compact representation of the feature domain. Subsequently, the latent variables are input into a priori analysis transformation network to generate auxiliary latent variables, used to characterize the scale dependencies of the latent variables in spatial and channel dimensions. These auxiliary latent variables are further input into a priori synthesis transformation network to predict the distribution parameters of the latent variables. Based on this, the latent variables and their distribution parameters are fed into an entropy coding model to generate a compressed bitstream, which is then transmitted to the receiving end via a transmission channel. In this embodiment, the network structure of the entropy coding model is as follows: Figure 9 As shown, it includes: a sequentially connected chessboard-based context model, a parametric inference network, a Gaussian tuning model, and an arithmetic encoder.
[0172] At the receiving end, the approximate reconstruction result of the latent variables is recovered through the corresponding entropy decoding model, and then input into the latent synthesis transformation network in the latent variable codec for inverse feature reconstruction to obtain the decoded versions of the illumination features and content features of the target image. The entropy decoding model includes: a sequentially connected checkerboard-based context model, a parameter inference network, a Gaussian modulation model, and an arithmetic decoder.
[0173] Finally, the reconstructed illumination and content features are input into the illumination synthesis network. Through element-wise multiplication, convolution, and pixel inverse rearrangement, the feature domain information is restored to the image domain representation, outputting the reconstructed target image. This achieves end-to-end illumination feature decoupling and semantic compression from the input image to the reconstructed image. The illumination decomposition-synthesis module separates content features from illumination features, enabling the network to maintain a consistent understanding of the scene structure under conditions of strong light, weak light, or localized illumination changes. Joint modeling of latent variables and super-prior allows for accurate probability estimation of latent features, thereby improving compression ratio and reconstruction quality. Finally, the illumination synthesis network achieves adaptive illumination compensation, ensuring that the reconstructed image maintains visual consistency with the original image in terms of brightness, contrast, and detail. This enables efficient transmission and high-fidelity restoration of emergency monitoring images under low bitrate conditions, demonstrating significant real-time and robust advantages.
[0174] In this embodiment, the overall loss function of the emergency monitoring image semantic coding model based on illumination feature decoupling is shown in the following formula:
[0175] (7)
[0176] in, and All are hyperparameters, when When a higher value is assigned, the model can achieve a lower bitrate. It was assigned the value 0.0001.
[0177] Based on the same inventive concept, a second aspect of the embodiments of this application provides an emergency monitoring image semantic coding system based on illumination feature decoupling, such as... Figure 10 As shown, the system includes:
[0178] The feature domain decomposition module 201 is used to decompose the target image and the reference image in the feature domain through the illumination decomposition network in the illumination decomposition-synthesis module, so as to obtain the illumination features of the target image and the reference image respectively, and the content features of the target image and the reference image respectively.
[0179] The illumination feature compression module 202 is used by the transmitting end to obtain a compressed bitstream by using the illumination features of the target image and the reference image, and the content features of the target image and the reference image, respectively, through an encoder, and to send the compressed bitstream to the receiving end; and by the receiving end to recover the illumination features and content features of the target image through a decoder. The encoder includes a latent variable encoder, a super-prior encoder, and an entropy coding model. The decoder includes a latent variable decoder, a super-prior decoder, and an entropy decoding model.
[0180] The illumination decomposition-composite module 203 is used to perform image reconstruction by using the illumination features and content features recovered from the target image through an illumination synthesis network, so as to obtain the reconstructed target image of the target image.
[0181] Based on the same inventive concept, a third aspect of the embodiments of this application provides a method as follows: Figure 11 The electronic device 100 shown includes a processor 120, a memory 110, and a program or instructions stored in the memory 110 and executable on the processor 120. When the program or instructions are executed by the processor 120, they implement the steps of the emergency monitoring image semantic coding method based on illumination feature decoupling as described in the first aspect of this application.
[0182] Based on the same inventive concept, in a fourth aspect of this application, a readable storage medium is provided, on which a program or instructions are stored, and when the program or instructions are executed by a processor, the steps of the emergency monitoring image semantic coding method based on illumination feature decoupling as described in the first aspect of this application are implemented.
[0183] Each embodiment in this specification focuses on the differences from other embodiments. For the same or similar parts between the embodiments, please refer to each other.
[0184] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0185] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0186] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0187] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0188] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.
[0189] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0190] The above provides a detailed description of the emergency monitoring image semantic coding method, system, device, and medium based on illumination feature decoupling. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A semantic coding method for emergency monitoring images based on illumination feature decoupling, characterized in that, The method includes: The transmitting end decomposes the target image and the reference image in the feature domain through the illumination decomposition network in the illumination decomposition-synthesis module, and obtains the illumination features of the target image and the reference image, as well as the content features of the target image and the reference image. The transmitting end decomposes the target image and the reference image in the feature domain through the illumination decomposition-synthesis module, respectively, to obtain the illumination features of the target image and the reference image, and the content features of the target image and the reference image, including: The transmitting end performs pixel rearrangement operation on the first image through the pixel rearrangement module in the illumination decomposition network to obtain the pixel rearrangement result. The first image is the target image or the reference image. The transmitting end obtains a first intermediate result by using the pixel rearrangement result through a first convolutional module, a depth convolutional module stacked N times in sequence, and a second convolutional module in the illumination decomposition network. The depth convolutional module includes a first branch and a second branch. The first branch is deployed with channel-wise convolutional units and point-wise convolutional units, and the second branch is deployed with point-wise convolutional units. The transmitting end uses the segmentation module in the illumination decomposition network to uniformly segment the first intermediate result from the channel dimension to obtain a first part of intermediate result and a second part of intermediate result, and uses the first part of intermediate result as the content feature of the first image. The transmitting end performs max pooling processing on the second part of the intermediate results through the max pooling module in the illumination decomposition network to obtain the illumination features of the first image; The transmitting end uses the encoder in the illumination feature compression module to obtain a compressed bitstream using the illumination features of the target image and the reference image, as well as the content features of the target image and the reference image, and sends the compressed bitstream to the receiving end. The encoder includes: a latent variable encoder, a super prior encoder, and an entropy coding model. The receiving end uses the received compressed bitstream and the decoder in the illumination feature compression module to recover the illumination features and content features of the target image. The decoder includes: a latent variable decoder, a super prior decoder, and an entropy decoding model. The receiving end uses the illumination synthesis network in the illumination decomposition-synthesis module to reconstruct the image using the illumination features and content features recovered from the target image, and obtains the reconstructed target image of the target image. The receiving end, through the illumination synthesis network in the illumination decomposition-synthesis module, performs image reconstruction using the illumination features and content features recovered from the target image to obtain a reconstructed target image, including: The receiving end performs an element-wise multiplication operation on the illumination features and content features recovered from the target image through the element-wise multiplication module in the illumination synthesis network to obtain the element-wise multiplication result. The receiving end obtains a second intermediate result through the third convolution module, the depth convolution module stacked N times in sequence, and the fourth convolution module in the illumination synthesis network. The depth convolution module includes a first branch and a second branch. The first branch is deployed with channel-wise convolution units and point-wise convolution units, and the second branch is deployed with point-wise convolution units. N is an integer greater than or equal to 3. The receiving end performs a pixel inverse rearrangement operation on the second intermediate result through the pixel inverse rearrangement module in the illumination synthesis network to obtain the reconstructed target image of the target image. The pixel inverse rearrangement operation is the inverse operation of the pixel rearrangement operation.
2. The emergency monitoring image semantic coding method based on illumination feature decoupling according to claim 1, characterized in that, The transmitting end uses the encoder in the illumination feature compression module to obtain a compressed bitstream using the illumination features of the target image and the reference image, as well as the content features of the target image and the reference image, including: The transmitting end concatenates the illumination features and content features of the target image to obtain the illumination-content concatenated features of the target image, and concatenates the illumination features and content features of the reference image to obtain the illumination-content concatenated features of the reference image. The transmitting end inputs the illumination content splicing features of the reference image and the target image into the latent variable encoder in the encoder; The transmitting end processes the illumination content splicing features of the reference image and the target image through the fifth convolution module, the depth convolution module stacked N times in sequence, and the sixth convolution module in the latent variable encoder to obtain latent variables. The latent variables are then processed by the super prior encoder and the entropy coding model to obtain the compressed bitstream, where N is an integer greater than or equal to 3. The receiving end uses the received compressed bitstream and, through the decoder in the illumination feature compression module, recovers the illumination features and content features from the target image, including: The receiving end processes the received compressed bitstream through the super prior decoder and the entropy decoding model to obtain the recovered latent variables; The receiving end processes the recovered latent variables through the seventh convolution module, pixel rearrangement module, depth convolution module stacked M times in sequence in the latent variable decoder to obtain the illumination features and content features recovered for the target image, where M is an integer greater than or equal to 6; the depth convolution module includes: a first branch and a second branch, the first branch is deployed with channel-wise convolution units and point-wise convolution units, and the second branch is deployed with point-wise convolution units.
3. The emergency monitoring image semantic coding method based on illumination feature decoupling according to claim 1, characterized in that, The emergency monitoring image semantic coding method based on illumination feature decoupling is implemented through an emergency monitoring image semantic coding model based on illumination feature decoupling. The training process of the emergency monitoring image semantic coding model based on illumination feature decoupling includes: Multiple high-low illumination image pairs are obtained as a training set, and each high-low illumination image pair shares the same content features; Using the training set, the illumination decomposition-synthesis module and the illumination feature compression module to be trained are trained according to the target loss function to obtain the emergency monitoring image semantic coding model based on illumination feature decoupling. The target loss function includes at least one of the following: The first loss function value is used to guide the illumination feature compression module to recover illumination features and content features from the sample image and align them with the illumination decomposition network in the illumination decomposition-synthesis module to decompose the sample image, so as to minimize the difference between the image reconstructed by the illumination decomposition-synthesis network in the illumination decomposition-synthesis module and the sample image. The sample image is a high-illumination image or a low-illumination image included in each high-low illumination image pair. The second loss function value is used to guide the illumination decomposition network in the illumination decomposition-synthesis module to learn the same content features shared by each high-low illumination image pair. The third loss function value is used to guide the illumination decomposition network in the illumination decomposition-synthesis module to learn the different illumination features in each high-low illumination image pair, including the high illumination image and the low illumination image.
4. The emergency monitoring image semantic coding method based on illumination feature decoupling according to claim 3, characterized in that, The target loss function also includes at least one of the following: The fourth loss function value is used to minimize the difference between the content features recovered by the illumination feature compression module to be trained for the first high illumination image and the content features recovered for the first low illumination image. The first high illumination image and the first low illumination image are two images in a high-low illumination image pair. Compression ratio, which is determined based on the original number of bits of the sample image and the number of bits of the compressed bitstream for the sample image, the sample image comprising either a high-light image or a low-light image for each high-low illumination image pair.
5. The emergency monitoring image semantic coding method based on illumination feature decoupling according to claim 1, characterized in that, The transmitting end uses the encoder in the illumination feature compression module to obtain a compressed bitstream using the illumination features of the target image and the reference image, as well as the content features of the target image and the reference image, including: The transmitting end compares the illumination features of the target image with the illumination features of the reference image, and compares the content features of the target image with the content features of the reference image; The transmitting end determines whether there are any changed content features and whether there are any changed illumination features; When the sending end determines that the target image and the reference image share the same content features, it uses the encoder in the illumination feature compression module to obtain a compressed bitstream using the changed illumination features and the same content features; or When the transmitting end determines that there are changed content features and changed illumination features, it uses the encoder in the illumination feature compression module to obtain a compressed bitstream using the changed illumination features and the changed content features. The method further includes: When the target image and the reference image share the same content features, the receiving end uses the pre-cached content features recovered from the reference image as the content features recovered from the target image. When there are changed content features, the sending end obtains the content features recovered for the target image based on the received changed content features and the pre-cached content features recovered for the reference image.
6. An emergency monitoring image semantic coding system based on illumination feature decoupling, characterized in that, The system includes: The feature domain decomposition module is used to decompose the target image and the reference image in the feature domain through the illumination decomposition network in the illumination decomposition-synthesis module, so as to obtain the illumination features of the target image and the reference image respectively, and the content features of the target image and the reference image respectively. The illumination decomposition network decomposes the target image and the reference image in the feature domain to obtain the illumination features of the target image and the reference image respectively, as well as the content features of the target image and the reference image respectively, including: The pixel rearrangement module in the illumination decomposition network performs a pixel rearrangement operation on the first image to obtain a pixel rearrangement result, wherein the first image is the target image or the reference image. The illumination decomposition network consists of a first convolutional module, a depth convolutional module stacked N times in sequence, and a second convolutional module connected in series. The first intermediate result is obtained by using the pixel rearrangement result. The depth convolutional module includes a first branch and a second branch. The first branch is deployed with channel-wise convolutional units and point-wise convolutional units, and the second branch is deployed with point-wise convolutional units. The segmentation module in the illumination decomposition network uniformly segments the first intermediate result from the channel dimension to obtain a first part of intermediate result and a second part of intermediate result, and uses the first part of intermediate result as the content feature of the first image. The max pooling module in the illumination decomposition network performs max pooling processing on the second part of the intermediate results to obtain the illumination features of the first image. The illumination feature compression module is used by the transmitting end to obtain a compressed bitstream by using the illumination features of the target image and the reference image, and the content features of the target image and the reference image, respectively, through an encoder, and to send the compressed bitstream to the receiving end; and by the receiving end to recover the illumination features and content features of the target image through a decoder. The encoder includes: a latent variable encoder, a super-prior encoder, and an entropy coding model. The decoder includes: a latent variable decoder, a super-prior decoder, and an entropy decoding model. The illumination decomposition-synthesis module is used to perform image reconstruction by using the illumination features and content features recovered from the target image through an illumination synthesis network, so as to obtain the reconstructed target image of the target image; The illumination synthesis network utilizes the illumination features and content features recovered from the target image to perform image reconstruction, obtaining a reconstructed target image, including: The element-wise multiplication module in the illumination synthesis network performs an element-wise multiplication operation on the illumination features and content features recovered from the target image to obtain the element-wise multiplication result. The third convolutional module, the depth convolutional module stacked N times in sequence, and the fourth convolutional module in the illumination synthesis network are connected in series to obtain the second intermediate result. The depth convolutional module includes a first branch and a second branch. The first branch is deployed with channel-wise convolutional units and point-wise convolutional units, and the second branch is deployed with point-wise convolutional units. N is an integer greater than or equal to 3. The pixel inverse rearrangement module in the illumination synthesis network performs a pixel inverse rearrangement operation on the second intermediate result to obtain the reconstructed target image of the target image. The pixel inverse rearrangement operation is the inverse operation of the pixel rearrangement operation.
7. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein when the program or instructions are executed by the processor, they implement the steps of the emergency monitoring image semantic coding method based on illumination feature decoupling as described in any one of claims 1-5.
8. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions, which, when executed by a processor, implement the steps of the emergency monitoring image semantic coding method based on illumination feature decoupling as described in any one of claims 1-5.
Citation Information
Patent Citations
Progressive pixel-level adjustment low-illumination image enhancement method
CN117593222A
Low-illumination image semantic segmentation method based on multi-scale style decoupling network MSD-Net
CN118710907A