A semantic transmission method for visible light and infrared fusion images of nighttime vehicle-road environment

By collecting and enhancing visible light and infrared images in a night vehicle-road environment, performing global and local fusion, and performing end-to-end semantic transmission through a potential diffusion model, the problem of insufficient image fusion capability in a night vehicle-road environment in the existing technology is solved, and efficient environmental perception and image transmission are achieved.

CN119360341BActive Publication Date: 2025-05-23CHINA JILIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411919761.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-23
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

The prior art is difficult to effectively fuse visible light and infrared images in night vehicle-road environments, resulting in limited perception capabilities. Traditional image fusion technology ignores the semantic information of the image, limiting the application potential of advanced semantic analysis tasks.

Method used

A semantic transmission method for visible light and infrared images in a night vehicle-road environment is proposed. Through image enhancement, image fusion and semantic communication technology, visible light and infrared images are collected and enhanced, global and local fusion is performed, semantic features are extracted and evaluated, semantic feature groups are generated, and end-to-end semantic transmission is carried out through the potential diffusion model.

Benefits of technology

It realizes efficient integration of visible light and infrared images in a night vehicle-road environment, improves environmental perception ability and image transmission efficiency, optimizes resource allocation, and improves image reconstruction efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360341B_ABST
    Figure CN119360341B_ABST
Patent Text Reader

Abstract

The present invention discloses a semantic transmission method of visible light and infrared fusion images of a nighttime vehicle-road environment. First, visible light and infrared images are collected during vehicle driving, and a visible light image enhancement network is used to improve the visibility of the image under low light conditions; second, the enhanced visible light image is fused with the infrared image; then, the fused image is semantically guided, the target semantic information is identified, the semantic features are extracted, and the importance of the feature vector is evaluated, thereby generating a semantic feature group; finally, the transmitting end encodes the fused image into a latent space representation, and transmits it to a conditional signal through channel coding and compression, and transmits it to the receiving end through an additive white Gaussian noise channel. The receiving end uses a conditional perception neural network to dynamically adjust the denoiser weight based on a potential diffusion model, control the generation process, and finally generate a restored image through a semantic decoder. This method is of great significance for improving the reliability and safety of the autonomous driving system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image fusion, semantic communication, etc., and in particular to a semantic transmission method of a visible light and infrared fused image of a vehicle-road environment at night. Background Art

[0002] With the development of autonomous driving technology, it is particularly important for vehicles to obtain and perceive environmental information. In the field of intelligent transportation systems, nighttime driving safety is a major challenge. Due to the poor lighting conditions at night, the visibility of visible light images is greatly reduced, resulting in limited performance of perception systems that rely on visible light images. Although infrared images can provide certain visual information at night, their low resolution and insufficient detail information make it difficult to meet the needs of accurate target recognition. Therefore, how to effectively fuse visible light and infrared images to improve the perception of nighttime road scenes has become a technical problem that needs to be solved urgently. At the same time, in the field of semantic communication, end-to-end semantic encoding and decoding technology is the key to achieving efficient information transmission. This technology involves encoding semantic information into a format that can be transmitted and restoring the original semantic information at the receiving end through a decoder.

[0003] Although the technology of visible light and infrared image fusion has made certain progress, the existing methods still have many shortcomings in practical applications. First, many fusion algorithms find it difficult to balance the detail information of visible light images and the contrast advantage of infrared images when processing night images, resulting in unsatisfactory visual quality and semantic information of the fused images. Second, existing fusion methods often lack adaptability to specific environmental conditions at night and find it difficult to maintain stable performance under dynamically changing lighting and weather conditions. In addition, traditional image fusion techniques mostly focus on pixel-level processing and ignore the semantic information of images, which limits the application potential of fused images in advanced semantic analysis tasks. Therefore, developing a new method that can effectively fuse visible light and infrared images in a nighttime vehicle-road environment and realize end-to-end semantic transmission is of great significance to improving the environmental perception capability of autonomous vehicles. Summary of the invention

[0004] In order to solve the above technical problems existing in the prior art, the present invention proposes a semantic transmission method of visible light and infrared fusion images of a nighttime vehicle-road environment, which can improve the environmental perception ability of an autonomous driving vehicle by using technologies such as image enhancement, image fusion, and semantic communication, and realize end-to-end semantic transmission of the vehicle, thereby improving nighttime driving safety and image transmission efficiency. The specific technical solution is as follows:

[0005] A semantic transmission method of visible light and infrared fusion images in a nighttime vehicle-road environment comprises the following steps:

[0006] Step 1: Collect visible light images and infrared images of the vehicle during driving, and use an unsupervised visible light image enhancement network to enhance the images;

[0007] Step 2: Fusing the enhanced visible light image and infrared image;

[0008] Step 3: Perform semantic guidance processing on the fused image, identify the target semantic information of the fused image, extract semantic features of the target in the image, and evaluate the importance of the semantic feature vector, and finally generate a semantic feature group;

[0009] Step 4: The transmitter encodes the fused image into a latent space representation. The receiver adds noise to the latent space through a denoiser based on a latent diffusion model to learn the data distribution. Finally, the denoised latent space representation is input into the semantic decoder to generate a restored image.

[0010] Furthermore, the step 1 specifically includes:

[0011] Step 1.1: Collect the original visible light image and infrared image, input the visible light image into a sub-network containing three convolutional attention modules and one convolutional layer, and input the infrared image into a sub-network containing one convolutional attention module and one convolutional layer;

[0012] Step 1.2: The visible light image and infrared image are output from their respective sub-networks and then merged through the convolutional layer to obtain a pixel-by-pixel brightness enhancement map. The generated pixel-by-pixel brightness enhancement map is applied to the original visible light image and enhanced using the brightness enhancement curve. The specific formula is:

[0013] FLE(Mvis) = Mvis + Mmap * Mvis * (1 - Mvis)

[0014] Among them, Mvis represents the original visible light image, Mmap represents the pixel-by-pixel brightness enhancement map, and FLE() represents the brightness enhancement curve.

[0015] Furthermore, the convolutional attention module in step 1.1 is based on Res-block.

[0016] Furthermore, the step 2 specifically includes:

[0017] Step 2.1: Send the enhanced visible light image and infrared image to the shared feature layer based on Transfermer to extract the basic features of the two modal images. The formula is expressed as:

[0018] ,

[0019] in, represents a shared encoder, represents an infrared image, represents the enhanced visible light image, and They represent the basic features of the extracted infrared image and the enhanced visible light image respectively;

[0020] Step 2.2: The features extracted in step 2.1 are simultaneously passed to the local encoder and the global encoder for further processing. The global encoder is used to capture global structural information, while the local encoder is used to extract detailed texture information to retain the unique characteristics of each modality. The specific formula is expressed as:

[0021] ,

[0022] ,

[0023] in, represents the global encoder, represents the local encoder, and Respectively represent the global features of the extracted infrared image and the enhanced visible light image, and Respectively represent the local features of the extracted infrared image and the enhanced visible light image;

[0024] Step 2.3: Extract the global features from step 2.2 , and local features , Send it to the global fusion layer and the local fusion layer for global and local fusion. Finally, the fused features , Input to the decoder to generate the fused image , the specific formula is:

[0025] ,

[0026] ,

[0027] ,

[0028] in, and represents the global fusion layer and the local fusion layer, Represents a decoder, , They represent the fused global features and local features respectively.

[0029] Furthermore, in step 2.2, the global encoder is constructed using a Restormer block, and the local encoder uses a reversible neural network.

[0030] Furthermore, the step 3 specifically includes:

[0031] Step 3.1: The target semantic information of the fused image is identified by a convolutional neural network, and the fused image of the infrared image and the visible light image is mapped to a semantic feature space. In the semantic feature space, the generated feature map retains the complete semantic information of the fused image of the infrared and visible light;

[0032] Step 3.2: The semantic segmentation network is used to act on the feature map generated in step 3.1 to generate a semantic information label map. The feature map and the semantic information label map are combined to obtain a feature map with semantic information labels.

[0033] Step 3.3: Generate a semantic importance map in the semantic feature space through semantic importance modeling, wherein the semantic importance map is used to reflect the importance score of each semantic feature vector, and combine the feature map marked with the semantic information with the semantic importance map to generate a semantic feature group.

[0034] Furthermore, the step 4 specifically includes:

[0035] Step 4.1: At the sending end, a semantic encoder is used to extract the latent space representation, i.e., semantic information, from the fused image.

[0036] Step 4.2: Channel encode the latent space representation based on the semantic importance guidance and further compress it into a conditional signal, which is then transmitted to the receiver through an additive white Gaussian noise channel;

[0037] Step 4.3: At the receiving end, a condition-aware neural network is used to generate dynamic weights according to the received conditional signal;

[0038] Step 4.4: The latent diffusion model-based denoiser generates a denoised latent space representation by inputting the conditioned signal in step 4.2 and the dynamic weights generated in step 4.3, as well as the randomly generated Gaussian noise introduced during the training phase to simulate the transmission process.

[0039] Furthermore, the denoiser based on the latent diffusion model comprises a diffusion model and a denoising network, and has two working stages: a forward diffusion process and a reverse reasoning process. In the forward diffusion process, the initial latent space representation Gaussian noise is progressively added to produce a noisy latent space representation , and then feed it to the denoiser network. The subscript T refers to the time index, which only occurs during the training phase. During the reverse inference process, the randomly generated Gaussian noise is fed to the denoising network, which in addition receives the noise-conditioned signal The dynamic weight W is used as input only to adjust the weight of the denoiser and does not directly participate in the forward diffusion or reverse reasoning process.

[0040] Furthermore, in the forward diffusion process, the noise condition signal Assume that it is a known condition and represent the initial latent space Seen as a target , the purpose of the forward diffusion process is to generate a noisy latent space representation , specifically, input At each time index t and the variance The Gaussian noise is gradually added, resulting in arrive A series of potential empty representations with noise, this process is denoted as q, and the mathematical expression is:

[0041] ,

[0042] ,

[0043] The reverse inference process performs the denoising function. During the training process, the denoising network learns to denoise the noisy latent space representation. , to compensate for the loss of semantic information caused by compression. In this process, the noise-conditioned signal and noisy latent space representation is fed into the denoising network, which gradually denoises , and finally obtain the denoised latent space representation The goal is to make Closely approximating the initial latent space representation , the mathematical expression of the reverse process is:

[0044] ,

[0045] where θ represents the learnable parameters of the reverse process, is a randomly generated Gaussian noise, modeled as , the denoising network is used to learn , let the denoising network function be , the loss function of the denoiser is expressed as:

[0046] ,

[0047] in represents a Gaussian noise variable.

[0048] Furthermore, the conditional-perceptual neural network is composed of multiple parallel fully-connected layers, each fully-connected layer generates unique dynamic weights for different hidden layers of the denoising network, each dynamic weight vector is unique for each sample, and for a single sample, the conditional-perceptual neural network generates N dynamic weight vectors, which are then applied to different hidden layers of the denoising network.

[0049] The advantages and beneficial effects of the present invention are:

[0050] 1. The present invention proposes a visible light image enhancement mechanism, and the enhanced visible light image can be better fused with the infrared image;

[0051] 2. The present invention proposes a multi-scale visible light and infrared image fusion mechanism. The images are fused at both global and local scales. This design can preserve both the global structure and local details of infrared and visible light images.

[0052] 3. The present invention proposes an image target semantic importance mechanism, which guides conditional channel coding according to the degree of importance, and can effectively improve transmission efficiency and optimize resource allocation;

[0053] 4. The present invention proposes a denoiser based on a diffusion model to perform a semantic communication mechanism for transmitting images, which can better reduce computing requirements and improve image reconstruction efficiency and quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a flow chart of a semantic transmission method of visible light and infrared fusion images in a nighttime vehicle-road environment according to an embodiment of the present invention.

[0055] Figure 2 It is a schematic diagram of the process of enhanced fusion of visible light image and infrared image in an embodiment of the present invention.

[0056] Figure 3 Schematic diagram of the structure of a diffusion model-based denoiser in an embodiment of the present invention. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical scheme and technical effect of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0058] like Figure 1 As shown, a semantic transmission method of a visible light and infrared fusion image in a nighttime vehicle-road environment of the present invention is intended to improve nighttime driving safety and image transmission efficiency, and includes the following steps:

[0059] Step 1: During the driving process of the vehicle, visible light images and infrared images are collected, and an unsupervised visible light image enhancement network is used to improve the visibility of the visible light image under low light conditions, that is, image enhancement is performed. The image enhancement, refer to Figure 2 , specifically including the following steps:

[0060] Step 1.1: Input the collected original visible light image and infrared image into different sub-networks respectively. The sub-network of visible light image contains three convolutional attention modules and one convolutional layer, while the sub-network of infrared image contains one convolutional attention module and one convolutional layer. The features of the source image are extracted through their respective independent sub-networks.

[0061] Furthermore, the convolutional attention module in step 1.1 can extract useful local features and suppress redundant features, thereby improving the effect of image enhancement. The convolutional attention module is based on Res-block and increases the information sharing of feature tensors in spatial and channel dimensions. These modules are designed to generate spatial and channel attention maps, and rescale the input feature maps through average pooling and global pooling. Then, the convolutional layer is used to further process the features.

[0062] Step 1.2: The visible light image and infrared image are output from the sub-networks and then merged through the convolutional layer to obtain a pixel-by-pixel brightness enhancement map. The generated pixel-by-pixel brightness enhancement map is applied to the original visible light image and enhanced using the brightness enhancement curve. The specific formula is:

[0063] FLE(Mvis) = Mvis + Mmap * Mvis * (1 - Mvis)

[0064] Among them, Mvis represents the original visible light image, Mmap represents the pixel-by-pixel brightness enhancement map, and FLE() represents the brightness enhancement curve. This formula ensures that the enhanced image improves brightness while maintaining contrast.

[0065] Step 2: Fuse the enhanced visible light image and infrared image. Figure 2 , specifically including the following steps:

[0066] Step 2.1: Send the enhanced visible light image and infrared image to the shared feature layer based on Transfermer to extract the basic features of the two modal images, providing a basis for subsequent feature alignment and fusion. The specific formula can be expressed as:

[0067] ,

[0068] in, represents a shared encoder, represents an infrared image, represents the enhanced visible light image, and They represent the basic features of the extracted infrared image and the enhanced visible light image respectively.

[0069] Step 2.2: The features extracted in step 2.1 are simultaneously passed to the local encoder and the global encoder for further processing. The global encoder is used to capture global structural information, while the local encoder is used to extract detailed texture information to retain the unique characteristics of each modality. The specific formula can be expressed as:

[0070] ,

[0071] ,

[0072] in, represents the global encoder, represents the local encoder, and Respectively represent the global features of the extracted infrared image and the enhanced visible light image, and They represent the local features of the extracted infrared image and the enhanced visible light image respectively.

[0073] Furthermore, the global encoder in step 2.2 is constructed using the Restormer block, which is a powerful Transformer variant specifically designed for image processing tasks. The Restormer block effectively combines the advantages of the self-attention mechanism and convolutional neural network by introducing a local enhanced Transformer module, thereby improving the model's ability to extract image features. This design enables the basic decoder to better understand and represent the global structural information in infrared and visible light images, providing a solid foundation for subsequent feature fusion. The local encoder uses a reversible neural network to extract detail features, which can effectively capture subtle changes in the image while maintaining the reversibility of the features.

[0074] Step 2.3: Extract the global features from step 2.2 , and local features , Send it to the global fusion layer and the local fusion layer for global and local fusion. Finally, the fused features , Input to the decoder to generate the fused image The specific formula can be expressed as:

[0075] ,

[0076] ,

[0077] ,

[0078] in, and represents the global fusion layer and the local fusion layer, Represents a decoder, , They represent the fused global features and local features respectively.

[0079] Step 3: Conduct semantic guidance on the fused visible light image and infrared image, evaluate the importance of each semantic feature vector, and guide the subsequent conditional channel coding design. First, identify the target semantic information of the fused image, extract semantic features of the target in the image, generate semantic information tags and segment them into multiple semantic streams, perform semantic importance modeling at the same time, evaluate the importance of the semantic feature vector in each semantic stream, and finally generate a semantic feature group. The specific steps include the following:

[0080] Step 3.1: The target semantic information is identified through a convolutional neural network, and the fused image of the infrared image and the visible light image is mapped to a semantic feature space. In the semantic feature space, the generated feature map retains the complete semantic information of the fused image of infrared and visible light.

[0081] Step 3.2: The semantic segmentation network is used to act on the feature map generated in step 3.1 to generate a semantic information label map, which defines multiple semantic regions. Combining the feature map and the semantic information label map, a feature map with semantic information labeling can be obtained.

[0082] Step 3.3: Next, a semantic importance map is generated on this semantic feature space through semantic importance modeling. The semantic importance map reflects the importance score of each semantic feature vector. The semantic information labeling map provides basic semantic information, while the semantic importance map further refines the semantic information and finally generates a semantic feature group together, which jointly indicates which areas are more important in communication.

[0083] Step 4: The transmitter encodes the fused image into a latent space representation, and the receiver adds noise to the latent space through a denoiser based on a latent diffusion model to learn the data distribution, and finally inputs the denoised latent space representation into the semantic decoder to generate a restored image. The specific steps include the following:

[0084] Step 4.1: At the sending end, a semantic encoder is used to extract latent space representations, i.e., semantic information, from the fused image. These latent codes represent the main content and structure of the image.

[0085] Step 4.2: Based on the semantic importance of the semantic feature group, the latent space representation is channel-coded to protect it from the influence of channel noise, and it is further compressed into a conditional signal, which is then transmitted to the receiver through an additive white Gaussian noise (AWGN) channel. At the receiver, the received conditional signal is affected by AWGN and thus becomes a noisy conditional signal.

[0086] Step 4.3: At the receiving end, the conditional perception neural network generates dynamic weights based on the received noisy conditional signal. These dynamic weights are used to effectively control the generation process of the diffusion model-based denoiser to enhance the denoising performance.

[0087] Further, the conditional-aware neural network consists of multiple parallel fully connected layers, each of which generates unique dynamic weights for different hidden layers of the denoising network. Each dynamic weight vector is unique for each sample. For a single sample, the conditional-aware network can generate N dynamic weight vectors, which are then applied to different hidden layers of the denoising network. Specifically, the dynamic weights are applied to selected fully connected layers and convolutional layers in the denoising network. The length of each dynamic weight vector matches the number of weights in the corresponding hidden layer of the denoising network.

[0088] Step 4.4: The denoiser based on the latent diffusion model generates a denoised latent space representation by inputting the noisy conditional signal in step 4.2 and the dynamic weights generated in step 4.3 as well as the randomly generated Gaussian noise. This Gaussian noise is used in the training phase to simulate the noise introduced during the transmission process so that the denoiser can be trained to effectively remove the noise and restore the latent space representation. See Figure 3 .

[0089] Furthermore, the diffusion model-based denoiser in step 4.4 is divided into two parts: the diffusion model and the denoising network, which works in two stages, the forward diffusion process and the reverse reasoning process. In the forward diffusion process, the initial latent space representation Gaussian noise is progressively added to produce a noisy latent space representation , which is then fed to the denoiser network, with the subscript T referring to the time index. This process only occurs during the training phase. During the reverse inference process, the randomly generated Gaussian noise is fed to the denoising network. In addition, the denoising network receives the noise-conditioned signal and dynamic weights W as input. The dynamic weights W are only used to adjust the weights of the denoiser and are not directly involved in the forward diffusion or reverse inference process.

[0090] In the forward diffusion process, the noise condition signal Assume that it is a known condition and represent the initial latent space Seen as a target The purpose of the forward diffusion process is to generate a noisy latent null representation Specifically, input At each time index t and the variance The Gaussian noise is gradually added, resulting in arrive A series of noisy latent space representations. This process, denoted as q, is mathematically expressed as:

[0091] ,

[0092] ,

[0093] The reverse inference process performs the denoising function. During the training process, the denoising network learns to denoise the noisy latent space representation , to compensate for the loss of semantic information caused by compression. The distribution and This is achieved by inferring and diffusing between the distributions of and noisy latent space representation is fed into the denoising network, which gradually denoises , and finally obtain the denoised latent space representation The goal is to make Closely approximating the initial latent space representation The reverse process can be expressed mathematically as:

[0094] ,

[0095] where θ represents the learnable parameters of the reverse process, is a randomly generated Gaussian noise, and the modeling can be recorded as , the denoising network is used to learn . Assume that the denoising network function is , the loss function of the denoiser can be expressed as:

[0096] ,

[0097] in represents a Gaussian noise variable.

[0098] Finally, after the forward diffusion process and the reverse inference process, the denoiser outputs the denoised latent space representation.

[0099] Step 4.5: Finally, the denoised latent space representation is fed into the semantic decoder to generate the restored image.

[0100] In summary, the semantic transmission method of visible light and infrared fusion images of nighttime vehicle-road environment of the present invention includes four key steps: first, collect visible light and infrared images during vehicle driving, and use visible light image enhancement network to improve the visibility of images under low light conditions; second, fuse the enhanced visible light image with the infrared image to obtain more comprehensive environmental information; then, perform semantic guidance processing on the fused image, identify target semantic information, extract semantic features, and evaluate the importance of feature vectors to generate semantic feature groups; finally, the transmitter encodes the fused image into a latent space representation, and transmits it to the receiver through channel coding and compression into a conditional signal through an additive white Gaussian noise (AWGN) channel. The receiver uses a conditional perception neural network to dynamically adjust the denoiser weights based on the potential diffusion model to more effectively control the generation process, and finally generates a restored image through a semantic decoder. This method is particularly suitable for nighttime vehicle-road environment, can effectively improve the transmission quality and reconstruction accuracy of semantic information of images, and is of great significance for improving the reliability and safety of autonomous driving systems.

[0101] The above is only a preferred implementation case of the present invention and does not limit the present invention in any form. Although the implementation process of the present invention is described in detail above, for those familiar with the art, they can still modify the technical solutions recorded in the above examples, or replace some of the technical features therein with equivalents. All modifications, equivalent replacements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A semantic transmission method of visible light and infrared fusion images in a nighttime vehicle-road environment, characterized in that: The steps include: Step 1: Collect visible light images and infrared images of the vehicle during driving, and use an unsupervised visible light image enhancement network to enhance the images; Step 2: Fusing the enhanced visible light image and infrared image; Step 3: Perform semantic guidance processing on the fused image, identify the target semantic information of the fused image, extract semantic features of the target in the image, and evaluate the importance of the semantic feature vector, and finally generate a semantic feature group, which includes: Step 3.1: The target semantic information of the fused image is identified by a convolutional neural network, and the fused image of the infrared image and the visible light image is mapped to a semantic feature space. In the semantic feature space, the generated feature map retains the complete semantic information of the fused image of the infrared and visible light; Step 3.2: The semantic segmentation network is used to act on the feature map generated in step 3.1 to generate a semantic information label map. The feature map and the semantic information label map are combined to obtain a feature map with semantic information labels. Step 3.3: Generate a semantic importance map in the semantic feature space through semantic importance modeling, wherein the semantic importance map is used to reflect the importance score of each semantic feature vector, and combine the feature map marked with the semantic information with the semantic importance map to generate a semantic feature group; Step 4: The transmitter encodes the fused image into a latent space representation. The receiver adds noise to the latent space through a denoiser based on a latent diffusion model to learn the data distribution. Finally, the denoised latent space representation is input into the semantic decoder to generate the restored image. Specifically, it includes: Step 4.1: At the sending end, a semantic encoder is used to extract the latent space representation, i.e., semantic information, from the fused image. Step 4.2: Based on the semantic importance of the semantic feature group, the latent space representation is channel-coded to protect it from the influence of channel noise, and it is further compressed into a conditional signal, which is then transmitted to the receiver through the additive white Gaussian noise (AWGN) channel; at the receiver, the received conditional signal is affected by the AWGN and thus becomes a noisy conditional signal; Step 4.3: At the receiving end, a condition-aware neural network is used to generate dynamic weights according to the received conditional signal; Step 4.4: The latent diffusion model-based denoiser generates a denoised latent space representation by inputting the conditioned signal in step 4.2 and the dynamic weights generated in step 4.3, as well as the randomly generated Gaussian noise introduced during the training phase to simulate the transmission process.

2. The semantic transmission method according to claim 1, characterized in that: The step 1 specifically includes: Step 1.1: Collect the original visible light image and infrared image, input the visible light image into a sub-network containing three convolutional attention modules and one convolutional layer, and input the infrared image into a sub-network containing one convolutional attention module and one convolutional layer; Step 1.2: The visible light image and infrared image are output from their respective sub-networks and then merged through the convolutional layer to obtain a pixel-by-pixel brightness enhancement map. The generated pixel-by-pixel brightness enhancement map is applied to the original visible light image and enhanced using the brightness enhancement curve. The specific formula is: FLE(Mvis) = Mvis + Mmap * Mvis * (1 - Mvis) Among them, Mvis represents the original visible light image, Mmap represents the pixel-by-pixel brightness enhancement map, and FLE() represents the brightness enhancement curve.

3. The semantic transmission method according to claim 2, characterized in that: The convolutional attention module in step 1.1 is based on Res-block.

4. The semantic transmission method according to claim 1, characterized in that: The step 2 specifically includes: Step 2.1: Send the enhanced visible light image and infrared image to the shared feature layer based on Transfermer to extract the basic features of the two modal images. The formula is expressed as: , in, represents a shared encoder, represents an infrared image, represents the enhanced visible light image, and They represent the basic features of the extracted infrared image and the enhanced visible light image respectively; Step 2.2: The features extracted in step 2.1 are simultaneously passed to the local encoder and the global encoder for further processing. The global encoder is used to capture global structural information, while the local encoder is used to extract detailed texture information to retain the unique characteristics of each modality. The specific formula is expressed as: , in, represents the global encoder, represents the local encoder, and Respectively represent the global features of the extracted infrared image and the enhanced visible light image, and Respectively represent the local features of the extracted infrared image and the enhanced visible light image; Step 2.3: Extract the global features from step 2.2 , and local features , Send it to the global fusion layer and the local fusion layer for global and local fusion. Finally, the fused features , Input to the decoder to generate the fused image , the specific formula is: , in, and represents the global fusion layer and the local fusion layer, Represents a decoder, , They represent the fused global features and local features respectively.

5. The semantic transmission method according to claim 4, characterized in that: In step 2.2, the global encoder is constructed using the Restormer block, and the local encoder uses a reversible neural network.

6. The semantic transmission method according to claim 1, characterized in that: The denoiser based on the latent diffusion model comprises a diffusion model and a denoising network, and has two working stages: a forward diffusion process and a reverse reasoning process. In the forward diffusion process, the initial latent space representation Gaussian noise is gradually added to produce a noisy latent space representation , which is then fed to the denoiser network. The subscript T refers to the time index, which only occurs during the training phase. During the reverse inference process, the latent space representation with noise is fed to the denoising network, which in addition receives the noise-conditioned signal and dynamic weights As input, dynamic weights It is only used to adjust the weights of the denoiser and does not directly participate in the forward diffusion or reverse inference process.

7. The semantic transmission method according to claim 6, characterized in that: In the forward diffusion process, the noise condition signal Assume that it is a known condition and represent the initial latent space Seen as a target , the purpose of the forward diffusion process is to generate a noisy latent space representation , specifically, input At each time index t and the variance The Gaussian noise is gradually added, resulting in arrive A series of potential empty representations with noise, this process is denoted as q, and the mathematical expression is: , The reverse reasoning process performs the denoising function. During the training process, the denoising network learns to denoise the latent space representation with noise. , to compensate for the loss of semantic information caused by compression. In this process, the noise-conditioned signal and the noisy latent space representation is fed into the denoising network, which gradually denoises , and finally obtain the denoised latent space representation The goal is to make Closely approximating the initial latent space representation , the mathematical expression of the reverse process is: , where θ represents the learnable parameters of the reverse process, is a latent space representation with noise, and the model is recorded as , the denoising network is used to learn , let the denoising network function be , the loss function of the denoiser is expressed as: , in represents a Gaussian noise variable.

8. The semantic transmission method according to claim 6, characterized in that: The conditional perceptual neural network consists of multiple parallel fully connected layers, each of which generates unique dynamic weights for different hidden layers of the denoising network. Each dynamic weight vector is unique for each sample. For a single sample, the conditional perceptual neural network generates N dynamic weight vectors, which are then applied to different hidden layers of the denoising network.

Citation Information

Patent Citations

  • Dark light image enhancement method based on fusion of visible light and near-infrared light

    CN113935935A

  • Night thermal infrared image semantic segmentation enhancement method based on improved ResNet

    CN115601723A