Coding Method and Device for a Remote Sensing and Communication Fusion System Oriented to Radar Images
By adopting a fusion method of semantic encoder and neural network channel encoder in radar image communication and perception system, the problem of separation of radar image communication and perception system design in the prior art is solved, and the effective fusion of radar image communication and perception system and the improvement of target recognition capabilities are achieved.
Patent Information
- Application Number
- CN202411498343.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-10-25
AI Technical Summary
In the prior art, radar image communication and perception system design are separated, and an integrated network architecture is lacking, resulting in the inability to effectively integrate detection and transmission.
Using a coding method of a synesthesia fusion system for radar images, the integration of object detection technology and semantic communication technology is achieved by establishing a semantic encoder model and a neural network channel encoder model. The specific steps include obtaining the synthetic aperture radar image dataset, extracting multi-layer depth feature information, performing multi-scale fusion and redundancy increase processing, and loading the feature information onto the physical channel for transmission.
It realizes the effective fusion of radar image communication and perception system, improves the stability and robustness of data under different signal-to-noise ratio conditions, and enhances the positioning and recognition capabilities of targets.
Smart Images

Figure CN119399588B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wireless communication, and particularly relates to a coding method and device for a communication and sensing fusion system for radar images. Background Art
[0002] Synthetic Aperture Radar (SAR) has unique imaging capabilities: it can provide high-resolution two-dimensional images without being affected by sunlight, cloud cover, and weather conditions. In recent years, SAR has important research significance in many fields such as marine environmental monitoring, resource exploration, and disaster emergency response, and object detection in its SAR images has broad application scenarios. The combination of SAR image target recognition and communication systems can achieve efficient and reliable detection and early warning, realizing information communication and sensing integration. The current research on SAR image recognition systems and communication systems is separated, and their combination is not effective enough. The emergence of semantic communication technology provides an effective design idea for the integration of SAR systems and communication systems.
[0003] Currently, the designs of radar image communication and sensing technologies based on object detection are all separate. They divide communication and sensing into two parts, either detecting first and then transmitting, or transmitting first and then detecting. They do not consider designing an integrated network architecture.
[0004] In the prior art, a patent for invention with a publication number of CN117036975A discloses a method and device for target recognition of SAR images. The method includes: obtaining a SAR image to be recognized; extracting quantitative semantic features and image features of the SAR image to be recognized; based on the pre-constructed semantic mapping association between quantitative semantic features and qualitative semantic features, mapping the quantitative semantic features of the SAR image to be recognized into qualitative semantic features of the SAR image to be recognized; encoding the image features, quantitative semantic features, and qualitative semantic features of the SAR image to be recognized respectively, and performing feature fusion on the respectively encoded features to output a target recognition result of the SAR image to be recognized according to the feature fusion result.
[0005] Another invention patent with the publication number CN118587439A discloses a multi-source remote sensing image semantic segmentation method and device based on a denoising diffusion probability model, including a dual-branch backbone network based on Mamba and a denoising diffusion probability model network; the paired original optical image and original synthetic aperture radar image are used to extract features and fuse features in four stages through the dual-branch backbone network based on Mamba to obtain multi-source fusion features in four stages, and then the multi-source fusion features in four stages are spliced into a multi-scale fusion feature; in the denoising diffusion probability model network, the original semantic segmentation label is denoised through the forward denoising module, and the result of splicing the multi-scale fusion feature and the denoised original semantic segmentation label is used as the output of the forward denoising module; the input feature is denoised through the noise decoder to predict the semantic segmentation label of the input feature.
[0006] Another invention patent with the publication number CN118540024A discloses a semantic communication encoding and decoding method, device, equipment and storage medium. The method includes: extracting semantic features of target data according to a semantic encoder; converting the semantic features into a preset data transmission format according to a channel encoder; transmitting the semantic features to a receiving end according to a channel transmission module; decoding the semantic features according to a channel decoder to obtain the target data; converting the target data back to the original data format according to a semantic decoder; wherein, the semantic encoder and the semantic decoder are trained according to a deep learning model, and the semantic encoder, channel encoder, channel transmission module, channel decoder and semantic encoder are taken as a whole, and a preset convolutional neural network deep learning model is trained according to the whole and the end-to-end training method. Summary of the Invention
[0007] A brief overview of the embodiments of the present invention is given below to provide a basic understanding of certain aspects of the present invention. It should be understood that the following overview is not an exhaustive overview of the present invention. It is not intended to identify the key or important parts of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is only to present certain concepts in a simplified form as a prelude to the more detailed description to be discussed later.
[0008] To solve the above problems, the present application provides an encoding method and device for a communication and sensing fusion system for radar images, which can form an integrated structure for transmission and detection.
[0009] According to one aspect of the present application, there is provided an encoding method for a communication and sensing fusion system for radar images, including the following steps:
[0010] Step S1: Obtain a data set, which includes a synthetic aperture radar image data set, test environment parameters, and signal-to-noise ratio parameters;
[0011] Step S2: Establish a semantic encoder model, which includes a Backbone module and a Neck module, for extracting multi-layer depth feature information of the synthetic aperture radar image dataset;
[0012] Step S3: Establish a neural network channel encoder model, which includes a multi-scale fusion module and a redundancy increase module, for processing the multi-layer depth feature information output by the semantic encoder model and outputting the feature information;
[0013] Step S4: Load the feature information output by the neural network channel encoder model onto the physical channel and send it.
[0014] As a specific solution, in the step S1, the synthetic aperture radar image dataset includes basic parameters such as size and input batch. In addition, the signal-to-noise ratio parameter can be selected according to the actual scenario for the target signal-to-noise ratio in the design process.
[0015] As a specific solution, in the step S2, the semantic encoder model is established based on the Backbone module and the Neck module of a single-object detection framework;
[0016] Among them, the Backbone module includes a first convolutional layer CBS, a second convolutional layer CBS, a first cross-stage partial bottleneck convolution C2f, a third convolutional layer CBS, a second cross-stage partial bottleneck convolution C2f, a fourth convolutional layer CBS, a third cross-stage partial bottleneck convolution C2f, a fifth convolutional layer CBS, a fourth cross-stage partial bottleneck convolution C2f, and a spatial pyramid pooling fast SPPF connected in sequence;
[0017] The Neck module includes a first upsampling operation Upsample, a first scale concatenation operation Concat, a fifth cross-stage partial bottleneck convolution C2f, a second upsampling operation Upsample, a second scale concatenation operation Concat, a sixth cross-stage partial bottleneck convolution C2f, a sixth convolutional layer CBS, a third scale concatenation operation Concat, a seventh cross-stage partial bottleneck convolution C2f, a seventh convolutional layer CBS, a fourth Concat, and an eighth cross-stage partial bottleneck convolution C2f connected in sequence.
[0018] In the above solution, the Backbone module adopts a series of convolutional and transposed convolutional layers (CBS), and at the same time uses residual connections and bottleneck structures to reduce the size of the network and improve performance. The Backbone module adopts the Cross Stage Partial Bottleneck Convolution (C2f) module based on the idea of Cross Stage Local Network, making the single-object detection framework more lightweight while being able to obtain richer gradient flow information. At the end of the Backbone module, a Spatial Pyramid Pooling Fast (SPPF) module is adopted. The SPPF module serially passes 3 Maxpools with a size of 5×5, and then concatenates each layer for multi-scale feature fusion, enabling the model to utilize both global and local information simultaneously, thereby improving the object localization and recognition capabilities while greatly reducing the computational overhead of the network.
[0019] For the application scenario of radar images, this application introduces the Spatial Pyramid Pooling Fast (SPPF) and the Cross Stage Partial Bottleneck Convolution (C2f), making the Backbone module more lightweight and obtaining richer feature information. Compared with the existing Spatial Pyramid Pooling (SPP), SPPF improves the model's ability to obtain local and global feature information while being more lightweight, greatly improving the object localization and recognition capabilities.
[0020] The Neck module uses the Path Aggregation Network Feature Pyramid Network (PAN-FPN) structure. PAN-FPN constructs a top-down and bottom-up network structure, achieving the complementarity of shallow location information and deep semantic information through feature fusion, and thus realizing the diversity and integrity of feature information, further enhancing the feature representation ability. For the current application scenario, the Neck module adopts the Path Aggregation Network Feature Pyramid structure (PAN-FPN). PAN-FPN is a network structure that combines top-down and bottom-up. Compared with the existing bottom-up (PAN) and top-down (FPN), PAN-FPN can achieve the complementarity of front position information and deep semantic information, thereby realizing the diversity and integrity of feature information and further enhancing the feature expression ability.
[0021] Among them, the structures of each convolutional layer CBS (the first convolutional layer CBS... the seventh convolutional layer CBS) are the same, including a convolution, a normalization process, and an activation function connected in sequence.
[0022] The Spatial Pyramid Pooling Fast (SPPF) includes a CBS, a first max pooling, a second max pooling, a third max pooling, a scale concatenation, and a CBS connected in sequence.
[0023] The structures of each cross-stage partial bottleneck convolution C2f (the first cross-stage partial bottleneck convolution C2f... the seventh cross-stage partial bottleneck convolution C2f) are the same, including CBS, scale separation, the first bottleneck layer, the second bottleneck layer, the third bottleneck layer, scale splicing, and CBS connected in sequence.
[0024] The upsampling operation of upsample and the scale feature splicing operation of Concat both adopt existing technologies, so they will not be elaborated here.
[0025] As a specific solution, in step S2, the extraction of multi-layer depth feature information from the synthetic aperture radar image dataset by the semantic encoder model is expressed by the following formula:
[0026] Assume that the input of the synthetic aperture radar image dataset is an RGB image , where M is the input picture, R is the RGB image in the synthetic aperture radar image dataset, H represents the height of the RGB image, W represents the width of the RGB image, and 3 represents that the number of channels of the RGB image is 3; the result after the input data passes through the semantic encoding model is S, and S is expressed as:
[0027] (1)
[0028] Among them, S is the output of the semantic encoder, represents the semantic encoding network, and its parameters are β. β includes the first convolutional layer CBS, the second convolutional layer CBS, the first cross-stage partial bottleneck convolution C2f, the third convolutional layer CBS, the second cross-stage partial bottleneck convolution C2f, the fourth convolutional layer CBS, the third cross-stage partial bottleneck convolution C2f, the fifth convolutional layer CBS, the fourth cross-stage partial bottleneck convolution C2f, and the spatial pyramid fast pooling SPPF of the backbone network module, as well as the first upsampling operation Upsample, the first scale splicing operation Concat, the fifth cross-stage partial bottleneck convolution C2f, the second upsampling operation Upsample, the second scale splicing operation Concat, the sixth cross-stage partial bottleneck convolution C2f, the sixth convolutional layer CBS, the third scale splicing operation Concat, the seventh cross-stage partial bottleneck convolution C2f, the seventh convolutional layer CBS, the fourth Concat, and the eighth cross-stage partial bottleneck convolution C2f of the neck network module. The synthetic aperture radar image dataset includes STCD dataset image datasets, etc.
[0029] As a specific solution, step S3 specifically includes:
[0030] Step S31: Establish a multi-scale fusion module to fuse the multi-scale feature information (M1, M2, M3) transmitted by the neck network (Neck) module to generate new feature information X1; the specific fusion process is as follows:
[0031] (2)
[0032] Among them is the upsampling operation; represents the concatenation operation; M1, M2, and M3 are the outputs of three different scales of the semantic encoding module Neck network, that is, the outputs of the sixth cross-stage partial bottleneck convolution C2f, the seventh cross-stage partial bottleneck convolution C2f, and the eighth cross-stage partial bottleneck convolution C2f;
[0033] Step S32: First pass the new feature information X1 generated in Step S31 through the redundancy increasing module, and then perform normalization processing to obtain the normalized feature information X, ensuring that the signal maintains its quality and integrity as much as possible during the transmission process; the specific process is as follows:
[0034] , (3)
[0035] In the formula, X is the normalized feature information, (•) represents the redundancy layer operation of the redundancy increasing module, and its parameter is ; BN(•) is the normalization operation.
[0036] The multi-scale fusion module is used to effectively realize the fusion and information transmission between the semantic encoder model and the neural network channel encoder model.
[0037] As a specific solution, Step S4 specifically includes:
[0038] Step S41: Establish a physical channel, and the physical channel uses a Gaussian white noise channel or a Rayleigh fading channel;
[0039] Step S42: Transmit the normalized feature information X through the physical channel established in Step S41, and the signal Y received at the receiving end is expressed as:
[0040] (4)
[0041] In the formula, Y is the signal received at the receiving end; h represents the Rayleigh fading channel gain, η represents Gaussian noise with a mean of 0 and a variance of σ 2 That is, η~(0, σ 2 )(Gaussian noise η follows a normal distribution with a mean of 0 and a variance of σ2); for AWGN channel is the channel model representation under the Gaussian channel, for fading channel is the channel model representation under the Rayleigh fading channel.
[0042] According to another aspect of the present application, there is provided an encoding device for a communication and sensing fusion system for radar images, which is implemented by using the above encoding method.
[0043] The present application adopts the above solution, and compared with the prior art, has the following advantages:
[0044] 1. Currently, the designs of radar image communication and sensing technologies based on target detection are all separate designs, and there is no integrated design yet.
[0045] 2. The semantic encoder of this encoding method combines two modules, the Backbone network and the Neck network, and can effectively realize the fusion of target detection technology and semantic communication technology.
[0046] 3. The neural network-based channel encoder designed by this encoding method consists of a multi-scale fusion module and a redundancy addition module. The designed multi-scale fusion module can realize the organic coupling of the semantic encoder and the channel encoder.
[0047] In addition, through the above improved algorithm, experiments prove that as the signal-to-noise ratio increases, the data becomes more and more stable, and the data has better robustness to noise interference. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The present invention can be better understood by referring to the description given below in conjunction with the accompanying drawings, in which the same or similar reference numerals are used in all the drawings to denote the same or similar components. The accompanying drawings, together with the following detailed description, are included in this specification and form a part of this specification, and are used to further illustrate the preferred embodiments of the present invention and to explain the principles and advantages of the present invention. In the drawings:
[0049] Figure 1 It is a structural diagram of the encoding method for a communication and sensing fusion system for radar images according to an embodiment of the present invention;
[0050] Figure 2 It is a structural diagram of the semantic encoder according to an embodiment of the present invention;
[0051] Figure 3 It is a structural diagram of the Backbone and Neck according to an embodiment of the present invention;
[0052] Figure 4 It is a structural diagram of the CBS module according to an embodiment of the present invention;
[0053] Figure 5 It is a structural diagram of the SPPF module according to an embodiment of the present invention;
[0054] Figure 6 It is a structural diagram of the C2f according to an embodiment of the present invention;
[0055] Figure 7 The PAN-FPN structure diagram of the embodiment of the present invention, where the direction indicated by the arrow is the data flow direction;
[0056] Figure 8 The structure diagram of the channel encoder based on neural network of the embodiment of the present invention;
[0057] Figure 9 The structure diagram of the redundancy addition layer of the embodiment of the present invention;
[0058] Figure 10 Ten groups of data under different SNRs of the embodiment of the present invention. Detailed implementation manners
[0059] Embodiments of the present invention will be described below with reference to the accompanying drawings. Elements and features described in one drawing or one embodiment of the present invention can be combined with elements and features shown in one or more other drawings or embodiments. It should be noted that for the sake of clarity, representations and descriptions of components and processes unrelated to the present invention and known to those of ordinary skill in the art are omitted in the drawings and the description.
[0060] The present invention provides an encoding method for a communication and sensing fusion system for radar images, characterized in that it includes the following steps:
[0061] Step S1: Initialization stage, including determining the synthetic aperture radar image data set, determining the test environment parameters, determining the signal-to-noise ratio parameters, and important parameters in the corresponding design process;
[0062] Step S2: Establish a semantic encoder model, which mainly includes a backbone network module and a neck network module, for extracting multi-layer depth feature information of radar image data;
[0063] Step S3: According to the multi-layer depth feature information obtained through the semantic encoding model, transfer it to a channel encoder model based on neural network, which mainly includes a multi-scale fusion module and a redundancy addition module;
[0064] Step S4: Load the feature information output by the channel encoder based on neural network onto the physical channel and send it.
[0065] Among them, step S1 is specifically:
[0066] Step S11: Determine the basic parameters of the synthetic aperture radar image data set, including the size, input batch, and important parameters in the corresponding design process.
[0067] Step S12: Determine the target signal-to-noise ratio in the design process according to the scenario.
[0068] Step S2 is specifically:
[0069] Based on the backbone module and the neck module of a single-object detection framework, a basic semantic encoder model is established. The framework diagram is shown in Figure 2 .
[0070] Among them, the detailed model diagrams of the backbone module and the neck module are shown in Figure 3 . The backbone module includes a first convolutional layer CBS, a second convolutional layer CBS, a first cross-stage partial bottleneck convolution C2f, a third convolutional layer CBS, a second cross-stage partial bottleneck convolution C2f, a fourth convolutional layer CBS, a third cross-stage partial bottleneck convolution C2f, a fifth convolutional layer CBS, a fourth cross-stage partial bottleneck convolution C2f, and a spatial pyramid pooling fast SPPF connected in sequence.
[0071] The neck module includes a first upsampling operation Upsample, a first scale concatenation operation Concat, a fifth cross-stage partial bottleneck convolution C2f, a second upsampling operation Upsample, a second scale concatenation operation Concat, a sixth cross-stage partial bottleneck convolution C2f, a sixth convolutional layer CBS, a third scale concatenation operation Concat, a seventh cross-stage partial bottleneck convolution C2f, a seventh convolutional layer CBS, a fourth Concat, and an eighth cross-stage partial bottleneck convolution C2f connected in sequence;
[0072] Among them, each convolutional layer CBS (the first convolutional layer CBS... the seventh convolutional layer CBS) has the same structure, including convolution, normalization processing, and activation function connected in sequence.
[0073] The spatial pyramid pooling fast SPPF includes a CBS, a first max pooling, a second max pooling, a third max pooling, scale concatenation, and a CBS connected in sequence.
[0074] Each cross-stage partial bottleneck convolution C2f (the first cross-stage partial bottleneck convolution C2f... the seventh cross-stage partial bottleneck convolution C2f) has the same structure, including a CBS, scale separation, a first bottleneck layer, a second bottleneck layer, a third bottleneck layer, scale concatenation, and a CBS connected in sequence.
[0075] The upsampling operation of Upsample and the scale feature concatenation operation of Concat both adopt existing technologies, so they will not be elaborated here.
[0076] The backbone module is responsible for feature extraction. A series of convolutional and transposed convolutional layers are used to perform five downsamplings on the input radar image, obtaining five different scale features. At the same time, residual connections and bottleneck structures are used to reduce the size of the network and improve performance. The backbone module adopts the Cross Stage Partial Bottleneck Convolution (C2f) module based on the idea of Cross Stage Local Network, making the single-object detection framework more lightweight while also being able to obtain richer gradient flow information. At the end of the backbone module, the Spatial Pyramid Pooling Fast (SPPF) module is adopted. The SPPF module serially passes through 3 Maxpools with a size of 5×5, and then concatenates each layer for multi-scale feature fusion, enabling the model to utilize both global and local information simultaneously, thereby improving the object localization and recognition ability while greatly reducing the computational overhead of the network.
[0077] Step S23: The Path Aggregation Network Feature Pyramid Network (PAN-FPN) structure is used in the design of the Neck module. PAN-FPN constructs a top-down and bottom-up network structure, achieving the complementarity of shallow location information and deep semantic information through feature fusion, and thus realizing the diversity and integrity of feature information and further enhancing the feature representation ability.
[0078] The processing process of Step S2 can be represented by Formula 1. Assume that the input of the STCD dataset image is an RGB image , and the result after the input data passes through the semantic encoding model is .
[0079] (1)
[0080] Among them, S is the output of the semantic encoder, and its position in Figure 3 is the output of the sixth Cross Stage Partial Bottleneck Convolution C2f, the seventh Cross Stage Partial Bottleneck Convolution C2f, and the eighth Cross Stage Partial Bottleneck Convolution C2f; Denote the semantic encoding network, whose parameter is β. β represents the convolutional layer, BN normalization, and activation function. Specifically, β is the superposition of modules such as the backbone network module (the first convolutional layer CBS, the second convolutional layer CBS, the first cross-stage partial bottleneck convolution C2f, the third convolutional layer CBS, the second cross-stage partial bottleneck convolution C2f, the fourth convolutional layer CBS, the third cross-stage partial bottleneck convolution C2f, the fifth convolutional layer CBS, the fourth cross-stage partial bottleneck convolution C2f, and the spatial pyramid pooling fast SPPF) and the neck network module (the first upsampling operation Upsample, the first scale concatenation operation Concat, the fifth cross-stage partial bottleneck convolution C2f, the second upsampling operation Upsample, the second scale concatenation operation Concat, the sixth cross-stage partial bottleneck convolution C2f, the sixth convolutional layer CBS, the third scale concatenation operation Concat, the seventh cross-stage partial bottleneck convolution C2f, the seventh convolutional layer CBS, the fourth Concat, and the eighth cross-stage partial bottleneck convolution C2f), such as the convolutional layer CBS, the cross-stage partial bottleneck convolution C2f, and the spatial pyramid pooling fast SPPF in these modules. The bold S on the left side of the equation is the output of the semantic encoder.
[0081] For the application scenario of radar images in this application, the spatial pyramid pooling fast SPPF and the cross-stage partial bottleneck convolution C2f are introduced, making the Backbone module more lightweight and obtaining richer feature information. Compared with the existing SPP, SPPF is more lightweight while improving the model's ability to obtain local and global feature information, greatly improving the target positioning and recognition capabilities. For the current application scenario, the Neck module adopts the path aggregation network feature pyramid structure (PAN-FPN). PAN-FPN is a network structure that combines top-down and bottom-up. Compared with the existing bottom-up (PAN) and top-down (FPN), PAN-FPN can achieve the complementarity of front position information and deep semantic information, thereby realizing the diversity and integrity of feature information and further enhancing the feature expression ability.
[0082] Step S3 is specifically as follows:
[0083] Step S31: A radar image multi-scale fusion module is established to fuse the multi-scale feature information ( , , ) transmitted by the neck network (Neck) module to generate new feature information . By designing this module, the fusion and information transmission between the semantic encoding module and the channel encoding module can be effectively realized. The specific fusion process is as follows:
[0084] (2)
[0085] where is the upsampling operation; represents the concatenation operation; , , are the outputs of three different scales of the semantic encoder Neck network, that is, the outputs of the sixth cross-stage partial bottleneck convolution C2f, the seventh cross-stage partial bottleneck convolution C2f, and the eighth cross-stage partial bottleneck convolution C2f.
[0086] Step S32: Normalize the feature information generated in step S31 to ensure that the signal maintains its quality and integrity as much as possible during transmission. The specific process is as follows:
[0087] , (3)
[0088] In the formula, (•) represents the redundant layer operation, and its parameter is ; BN(•) is the normalization operation.
[0089] Specifically, the overall model diagram of step S3 is as Figure 8 shown.
[0090] Step S4 is specifically as follows:
[0091] Step S41: Establish a physical channel, which can be divided into a Gaussian white noise channel and a Rayleigh fading channel.
[0092] Step S42: Transmit the feature information X generated in step S32 through the physical transmission channel established in step S41. Assume that X is transmitted through the physical channel, then the received signal at the receiving end can be given by Y, and Y can be expressed as
[0093] (4)
[0094] In the formula, h represents the Rayleigh fading channel gain, η represents Gaussian noise with a mean of 0 and a variance of σ 2 , that is, η~(0, σ 2 ); for AWGN channel is the channel model representation under the Gaussian channel, for fading channel is the channel model representation under the Rayleigh fading channel.
[0095] The experimental environment of the embodiments of this application is configured with a Windows 10 operating system, an Intel 13th generation Core i9-13900K processor, an NVIDIA GeForce RTX 4090 graphics card, 24GB of RAM, a Pytorch version of 2.1.1, and a CUDA version of 12.1. The radar dataset used is the Sonar Common Target Detection Dataset (SCTD).
[0096] The specific data transmission process of this design invention can be simplified into the following steps. Assume that the input image of the STCD dataset is an RGB image , and the result after the input data passes through the semantic encoding model is S. The result after the semantic encoding output S is encoded by the neural network-based channel encoder is X.
[0097] The result data is obtained at signal-to-noise ratios of 0, 5, 10, 15, and 25 dB respectively. Ten groups of X data are taken for each signal-to-noise ratio case, and the results are shown in Table 1.
[0098] Table 1 Ten groups of X data under different SNRs
[0099]
[0100] It can be seen from the data in Table 1 that as the signal-to-noise ratio increases, the data becomes more and more stable. From Figure 10 , it can be more intuitively seen that as the signal-to-noise ratio increases, the data has better robustness to noise interference. This demonstrates the rationality of the encoding method of a communication and sensing fusion system for radar images proposed by us, and further proves the effectiveness of the algorithm we proposed.
[0101] Currently, the designs of radar image communication and sensing technologies based on object detection are all designed separately, and there is no integrated design yet. This patent designs an encoding method for a communication and sensing fusion system for radar images.
[0102] The semantic encoder of this encoding method combines two modules, the Backbone network and the Neck network, and can effectively realize the fusion of object detection technology and semantic communication technology.
[0103] The neural network-based channel encoder designed by this encoding method consists of a multi-scale fusion module and a redundancy increase module. The designed multi-scale fusion module can realize the organic coupling of the semantic encoder and the channel encoder.
[0104] It should be emphasized that the term "including / containing" when used in this article refers to the presence of features, elements, steps, or components, but does not exclude the presence or addition of one or more other features, elements, steps, or components.
[0105] Although the present invention has been disclosed above by the description of specific embodiments of the present invention, it should be understood that all the above embodiments and examples are exemplary and not restrictive. Those skilled in the art can design various modifications, improvements or equivalents to the present invention within the spirit and scope of the appended claims. These modifications, improvements or equivalents should also be considered to be included within the protection scope of the present invention.
Claims
1. A coding method for a radar image synaesthesia fusion system, characterized by: include: Step S1: Acquire a data set, which includes a synthetic aperture radar image data set, a test environment parameter, and a signal-to-noise ratio parameter; Step S2: establishing a semantic encoder model, the semantic encoder model comprising a backbone network module and a neck network module, for extracting multi-layer deep feature information of a synthetic aperture radar image dataset; Step S3: establishing a neural network channel encoder model, the neural network channel encoder model including a multi-scale fusion module and a redundancy increase module, for processing the multi-layer deep feature information output by the semantic encoder model and outputting the feature information; Step S4: Load the feature information output by the neural network channel encoder model onto the physical channel and send it; Wherein, in the step S2, the semantic encoder model is established based on the backbone network module and the neck network module of the single target detection framework; The backbone network module includes a first convolutional layer CBS, a second convolutional layer CBS, a first cross-stage partial bottleneck convolution C2f, a third convolutional layer CBS, a second cross-stage partial bottleneck convolution C2f, a fourth convolutional layer CBS, a third cross-stage partial bottleneck convolution C2f, a fifth convolutional layer CBS, a fourth cross-stage partial bottleneck convolution C2f and a spatial pyramid fast pooling SPPF connected in sequence; The neck network module includes a first upsampling operation Upsample, a first scale splicing operation Concat, a fifth cross-stage partial bottleneck convolution C2f, a second upsampling operation Upsample, a second scale splicing operation Concat, a sixth cross-stage partial bottleneck convolution C2f, a sixth convolution layer CBS, a third scale splicing operation Concat, a seventh cross-stage partial bottleneck convolution C2f, a seventh convolution layer CBS, a fourth Concat and an eighth cross-stage partial bottleneck convolution C2f, which are connected in sequence; In step S2, the multi-layer depth feature information of the synthetic aperture radar image data set is extracted by the semantic encoder model and expressed as follows: Assume that the input of the synthetic aperture radar image dataset is an RGB image , where M is the input image, R is the RGB image in the synthetic aperture radar image dataset, H represents the height of the RGB image, W represents the width of the RGB image, and 3 represents the number of channels of the RGB image is 3; the result of the input data after passing through the semantic encoding model is S, which is expressed as: ; Where S is the output of the semantic encoder, represents the semantic encoding network, whose parameter is β, β includes the first convolution layer CBS, the second convolution layer CBS, the first cross-stage partial bottleneck convolution C2f, the third convolution layer CBS, the second cross-stage partial bottleneck convolution C2f, the fourth convolution layer CBS, the third cross-stage partial bottleneck convolution C2f, the fifth convolution layer CBS, the fourth cross-stage partial bottleneck convolution C2f and the spatial pyramid fast pooling SPPF of the backbone network module, and the first upsampling operation Upsample, the first scale splicing operation Concat, the fifth cross-stage partial bottleneck convolution C2f, the second upsampling operation Upsample, the second scale splicing operation Concat, the sixth cross-stage partial bottleneck convolution C2f, the sixth convolution layer CBS, the third scale splicing operation Concat, the seventh cross-stage partial bottleneck convolution C2f, the seventh convolution layer CBS, the fourth Concat and the eighth cross-stage partial bottleneck convolution C2f of the neck network module; Step S3 specifically includes: Step S31: Establish a multi-scale fusion module to fuse the multi-scale feature information transmitted by the neck network module to generate new feature information X1; the specific fusion process is as follows: ; in is the upsampling operation; represents the concatenation operation; M1, M2, and M3 are the outputs of the sixth cross-stage partial bottleneck convolution C2f, the seventh cross-stage partial bottleneck convolution C2f, and the eighth cross-stage partial bottleneck convolution C2f of the semantic encoder, respectively; Step S32: The characteristic information X1 is first normalized by the redundancy adding module to obtain the normalized characteristic information X, so as to ensure that the signal maintains its quality and integrity as much as possible during the transmission process; the specific process is as follows: ; In the formula, X is the normalized feature information, (•) indicates redundant layer operation, whose parameters are ; BN (•) is the normalization operation.
2. The encoding method of the synaesthesia fusion system according to claim 1, characterized in that: In the step S1, the synthetic aperture radar image dataset includes a size and an input batch.
3. The encoding method of the synaesthesia fusion system according to claim 1, characterized in that: The step S4 specifically includes: Step S41: establishing a physical channel, where the physical channel adopts a Gaussian white noise channel or a Rayleigh fading channel; Step S42: The normalized characteristic information X is transmitted through the physical channel established in step S41, and the signal Y received at the receiving end is expressed as: (4) Where Y is the signal received at the receiving end; h is the Rayleigh fading channel gain, η is the mean value of 0 and the variance is σ 2 Gaussian noise, that is, η~(0,σ 2 ) ; for AWGN channel is the channel model representation under Gaussian channel, for fading channel is the channel model representation under Rayleigh fading channel.
4. A coding device for a radar image synaesthesia fusion system, characterized in that: The device executes the steps of the encoding method of the synaesthesia fusion system as described in any one of claims 1-3.
Citation Information
Patent Citations
Target identification method and device for SAR image
CN117036975A
Semantic communication coding and decoding method and device, equipment and storage medium
CN118540024A
Multi-source remote sensing image semantic segmentation method and device based on noise reduction diffusion probability model
CN118587439A
Image processing method, system and device and storage medium
CN114022496A
Semantic segmentation method based on improved ASPP and fusion module in complex scene
CN116342877A