Image super-resolution based signal enhancement algorithm design and system

By converting the received signal into an RGB image and enhancing it using an image super-resolution model, the problem of difficult signal feature extraction in the synaesthesia integrated system under low SNR conditions is solved, and high-quality signal recovery and accurate recognition are achieved, which is suitable for scenarios such as intelligent transportation and security monitoring.

CN120510041BActive Publication Date: 2025-10-10YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510993228.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-10
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

In complex environments, the low signal-to-noise ratio (SNR) in existing synaesthesia integrated systems makes it difficult to extract signal features, resulting in reduced communication quality and perception accuracy. Traditional methods are unable to effectively recover the key features of the signal.

Method used

The received time domain signal is converted into image form, enhanced using the image super-resolution model, and an RGB image is constructed through short-time Fourier transform. A denoising neural network based on a diffusion model is used to enhance the signal and extract key signal features.

Benefits of technology

Significantly improves signal demodulation and perception parameter recognition accuracy, and enhances the system's stable operation in low SNR scenarios. It is suitable for scenarios such as intelligent transportation, security monitoring, and counter-drone operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510041B_ABST
    Figure CN120510041B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of wireless communication, and particularly discloses a signal enhancement algorithm based on image super-resolution and an application system thereof, aiming to solve the problems of a signal quality difference, low spectral image resolution, and degraded identification and demodulation performance of a sensing-integrated system under low signal-to-noise ratio (SNR). The method converts a received signal into an RGB image through short-time Fourier transform, inputs the image into an artificial intelligence image super-resolution model for enhancement, improves the image resolution and details, extracts more accurate communication and sensing features such as amplitude, phase and frequency, and realizes signal denoising and enhancement. The algorithm takes into account the image structure and spectral characteristics, significantly improves the recoverability of sensing information in a complex environment, and is suitable for various intelligent systems such as unmanned aerial vehicles, vehicle-mounted radars, security monitoring and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of unmanned aerial vehicle (UAV) communication and perception integration, and particularly relates to a signal enhancement algorithm based on image super-resolution. Background Art

[0002] With the rapid development of wireless communications and intelligent sensing technologies, integrated communication and sensing systems are becoming an important development direction for next-generation wireless networks. These systems, by reusing wireless resources to integrate communication and sensing functions, are widely applicable in scenarios such as intelligent transportation, urban security, and aerial target monitoring. However, in complex environments (such as urban blocks, underground passages, or mountainous areas), the signal-to-noise ratio (SNR) of the signal received by the receiver is significantly reduced due to factors such as multipath propagation, non-line-of-sight propagation, and electromagnetic interference. This not only affects communication quality but also poses significant challenges to target perception accuracy. In particular, in conditions such as weak echoes and high obstruction, traditional time-domain or frequency-domain signal processing methods struggle to effectively recover the key signal features.

[0003] In recent years, artificial intelligence technologies, particularly super-resolution reconstruction algorithms in image processing, have demonstrated remarkable capabilities in restoring image detail. Converting communication or sensory signals into images and processing them with image super-resolution models promises to achieve higher-quality signal enhancement and information extraction in weak signal conditions. Currently, studies have attempted to use convolutional neural networks or generative adversarial networks to enhance short-time Fourier transform (STFT) spectra and spectrograms to restore target detail. However, most studies focus solely on communication or perception tasks. There is a lack of signal image enhancement frameworks for integrated sensory systems, and these approaches do not fully consider the synergistic information representation between multimodal signal features (such as amplitude, frequency, and phase). Therefore, there is an urgent need to design a signal enhancement method that combines the characteristics of integrated sensory systems, image representation techniques, and deep neural networks to improve signal demodulation and sensory parameter recognition accuracy in low-SNR scenarios, ensuring stable system operation in complex environments. Summary of the Invention

[0004] The present invention aims to solve the problem in existing synaesthesia integrated systems that signal features are difficult to extract in low signal-to-noise ratio (SNR) environments, resulting in reduced communication quality and perception accuracy. This invention provides a signal enhancement algorithm design based on image super-resolution and its application system. By converting the received time domain signal into an image form and enhancing it using an image super-resolution model, the communication and perception performance of the system in weak signal scenarios is improved.

[0005] The application is achieved by a signal enhancement algorithm design based on image super-resolution and an application system thereof, which comprises the following steps:

[0006] Step one, integrated deployment of communication and sensing: the system is composed of a transmitting end and a receiving end, which are located at different positions to form a bistatic architecture. The transmitting end sends a communication and sensing multiplexing signal with a specific encoding or modulation method, which can complete the communication task and provide sensing information for the receiving end.

[0007] Step two, signal reception and preprocessing: the receiving end receives the echo signal reflected by the target (such as unmanned aerial vehicle, vehicle, personnel, etc.) or the environment, and extracts the corresponding time domain complex signal sequence.

[0008] Step three, signal image conversion: the received time domain complex signal is converted into a two-dimensional time-frequency spectrum through short-time Fourier transform, and then an RGB image form is constructed.

[0009] Step four, image super-resolution and signal enhancement: the constructed image is input into an image super-resolution reconstruction model based on deep neural network (such as a diffusion model) to reconstruct the image details and denoise, and enhance the image clarity and texture information.

[0010] Step five, information feature reconstruction: key signal features (such as amplitude, phase, frequency) are extracted from the enhanced image to reconstruct the original signal information, and the sensing parameters such as target position, angle, distance and speed or the modulation content of the communication signal are recovered.

[0011] Further, the transmitting end forms multiple beams through a uniform planar array antenna, and adopts an integrated communication and sensing strategy for beam configuration, in which part of the fixed beams are used for communication tasks, and part of the scanning beams are used for target sensing. represents the transmitting beamforming vector, represents the phase offset vector, represents the power allocation factor, and respectively represent the beamforming vectors of the sensing and communication sub-beams. The transmitted signal carries communication information and sensing signal at the same time through modulation and encoding, so that the system realizes the collaborative and parallel processing of communication and sensing.

[0012] Further, the system comprises at least one transmitting end and one receiving end, and the transmitting end and the receiving end are located at different positions. The transmitting end is used for transmitting a signal with specific modulation or encoding that integrates communication and sensing functions; and the receiving end receives a signal containing target reflection or environmental echo.

[0013] Further, the transmitting antenna adopts a uniform planar array to realize multi-beam design, and the transmitted signal can be represented as wherein Represents the Orthogonal Frequency Division Multiplexing (OFDM) baseband signal, and Respectively represent the number of OFDM symbols and the number of subcarriers, represents the baseband symbol of the mth OFDM symbol on the nth subcarrier, represents the nth subcarrier frequency, represents the subcarrier spacing, represents a matrix function, Denotes the symbol duration. According to the signal transmission process, the channel model can be expressed as ,in and Represent the LoS component and NLoS component of the channel respectively, represents the Ricean channel K factor, Indicates the number of multipaths, , , Respectively represent The channel coefficient, Doppler frequency and delay of each path. and denote the arrival angle and departure angle respectively. In addition, and They represent the fading coefficients of the LoS path and the NLoS path respectively, represents the wavelength of the signal, Indicates the distance between the transmitter and the receiver of the LoS path. and They represent the distance between the transmitter and the reflector, and the distance between the reflector and the receiver, respectively. Represents the reflection coefficient.

[0014] The technical solution adopted by the present invention is as follows: providing a dual-base synaesthesia integrated system architecture; performing signal reception and preprocessing within this architecture; designing a signal image conversion module; proposing an image super-resolution and signal enhancement algorithm; reconstructing information features; performing data simulation, and comparing and analyzing the results. Specifically, the following steps are included:

[0015] S1. First, a model of a dual-base synaesthesia integrated system is presented. The system includes at least one transmitter and one receiver, located at different locations. The transmitter is configured to transmit signals with specific modulation or coding that integrate communication and perception functions; the receiver receives signals containing target reflections or environmental echoes.

[0016] S2. The receiving end receives the echo signal reflected by the target and extracts the corresponding time domain complex signal sequence.

[0017] S3. Convert the received time-domain communication / perception signal into a two-dimensional time-frequency image through short-time Fourier transform, and construct a three-channel RGB image corresponding to the signal strength, frequency distribution, and phase structure. The specific steps are as follows:

[0018] A1. The receiving end performs a short-time Fourier transform on the acquired time-domain communication / perception mixed signal, and obtains a complex-valued spectrum in the time-frequency domain through a sliding window function and a fast Fourier transform. .

[0019] A2. Change the modulus length of the complex spectrum The amplitude spectrum of the signal is extracted and mapped to the red channel (R channel) of the image through linear normalization. The frequency axis calculated by STFT is taken, its value is normalized, and it is copied along the time axis to generate a two-dimensional frequency distribution matrix to form the green channel (G channel) of the image. The phase component of the image is normalized to form the blue channel (B channel) of the image.

[0020] A3. Merge the above three channels into a three-channel RGB image. The resulting image provides an intuitive visualization of the spectral characteristics of the received signal, making it suitable for subsequent deep learning or pattern recognition tasks.

[0021] S4. Since the received signal SNR is low, after converting it into an RGB image, a denoising neural network based on a diffusion model is used to enhance the image, thereby improving the recoverability and perception accuracy of the original signal. The specific steps are as follows:

[0022] A1. Construct the forward diffusion process: transform the RGB image corresponding to the signal As the initial input, at each diffusion time step In the process, Gaussian noise is gradually added to the image to build an image sequence, forming a forward diffusion chain from the clear image to the noisy image; this process is modeled by the following formula: .in, Indicates the mean , the variance is The normal distribution of represents the identity matrix, Represents the noise attenuation factor for the current time step.

[0023] A2. Build a conditional diffusion model to learn at a given time step and noisy images Predicting the original noise A neural network, where the network structure may include UNet, attention module and residual module; by minimizing the mean square error loss: .

[0024] A3. In the test phase, from pure noise images First, use the trained denoising network to iteratively predict the residual noise and inversely obtain the image The estimated value of ; each step of the reverse diffusion process is: ,in . Predict the noise for the current time step, .

[0025] A4, denoising is performed through the trained diffusion model, and finally Get a clear image The image has a higher signal-to-noise ratio and clearer spectral details, thus providing a more accurate image basis for subsequent frequency, amplitude, and phase information extraction.

[0026] S5. In order to obtain the communication and perception information contained in the signal, the super-resolution image is restored and inverse STFT is performed to convert it into an enhanced RF signal. The specific steps are as follows:

[0027] A1. Perform inverse normalization and inverse transformation on the reconstructed high-resolution image, and apply algorithms such as inverse STFT to restore the image to the enhanced RF signal, ensuring that the restored signal in the time domain has a higher signal-to-noise ratio and stronger structural features.

[0028] A2. Perform traditional communication signal processing on the reconstructed time domain signal to demodulate and recover the symbol sequence.

[0029] A3. Target detection and parameter estimation are performed using the enhanced signal. The target angle is extracted using the high-resolution azimuth estimation algorithm Multiple Signal Classification (MUSIC). The two-dimensional Discrete Fourier Transform (2D-DFT) algorithm is then used to perform the inverse DFT and DFT on each row and column of the received signal matrix, respectively. This yields a two-dimensional range-velocity spectrum of the target, with the horizontal and vertical axes representing the target's range, position, and velocity, respectively.

[0030] S6. Data simulation is performed based on the proposed algorithm. The communication bit error rate and perception estimation accuracy are analyzed under different signal-to-noise ratios to verify the effectiveness of the proposed algorithm. At the same time, it is compared with other basic algorithms to verify the robustness and practicality of the proposed method under complex conditions such as low SNR and multipath.

[0031] Beneficial effects of the present invention:

[0032] First, the present invention's signal enhancement algorithm based on image super-resolution can effectively address the problems of low communication quality and low target perception accuracy caused by low SNR signals. Compared with traditional synchronization methods, the present invention can significantly improve signal demodulation and perception parameter recognition accuracy, ensuring stable operation of the system in complex environments. The method of the present invention has the following advantages:

[0033] 1. The present invention converts the received low signal-to-noise ratio communication / perception signals into image form and introduces an artificial intelligence image super-resolution model to effectively improve the clarity and detail retention ability of the spectrum image, and significantly enhance the recognizability and processability of weak signals.

[0034] 2. The proposed algorithm has good generalization ability and can maintain stable performance in various noise types (such as Gaussian white noise, multipath interference, etc.) and complex environments, especially showing stronger robustness in low signal-to-noise ratio scenarios below -10 dB.

[0035] 3. This invention leverages the multi-channel fusion mechanism of RGB images to separately encode information such as amplitude, frequency, and phase, enhancing the model's ability to extract physical quantities such as modulation mode, target position, speed, and distance. By constructing a signal image through STFT and employing a diffusion model or other advanced denoising networks for high-fidelity image restoration, this method effectively reduces bit error rates and estimation errors, thereby enhancing the system's overall recognition accuracy.

[0036] 4. The solution of the present invention can be smoothly integrated with existing orthogonal frequency division multiplexing (OFDM) communication systems and radar perception systems, has low dependence on hardware, and is suitable for rapid deployment and application in scenarios such as intelligent transportation, security monitoring, and anti-drone (unmanned aerial vehicle) countermeasures.

[0037] Second, by visualizing received signals and employing image super-resolution technology, this invention significantly improves the performance of communication and perception tasks in complex scenarios. In particular, it demonstrates excellent recognition accuracy and robustness in low signal-to-noise ratio and non-line-of-sight environments. This approach can be effectively applied to high-value scenarios such as intelligent transportation, security monitoring, and countering illegal drones. Furthermore, the proposed fusion solution of image signal processing and super-resolution models exhibits excellent versatility, making it widely applicable not only to traditional communication systems (such as 5G / 6G networks) but also to emerging integrated synaesthesia systems, drone perception platforms, and radar communication fusion equipment, providing fundamental support for the upgrading of related industries. Furthermore, this algorithm has low hardware deployment requirements and is suitable for rapid deployment via embedded systems, FPGAs, or GPUs. It is particularly well-suited for implementation in terminal-side scenarios such as smart cities, edge devices, and industrial perception, demonstrating excellent engineering translation potential and cost-effectiveness.

[0038] Most existing communication and perception systems rely on frequency-domain filtering, spatial equalization, or classical decision algorithms for signal enhancement. These methods struggle to effectively preserve subtle target features, particularly in low signal-to-noise ratio (SNR) environments, resulting in high bit error rates and short perception distances. This paper, however, proposes, for the first time, converting the received signal into an RGB image and restoring its high-resolution structure using an image super-resolution neural network. This provides a novel modeling and processing approach for signal enhancement in communication and perception systems, filling a gap in the integration of traditional signal and image processing. The RGB image constructed by this invention encodes amplitude, frequency, and phase information in its three color channels, visualizing and structuredly representing the three core physical properties of the original signal. This provides a unified input format that can be directly processed by deep learning models for signal enhancement and perception parameter extraction. This representation method is the first proposed in existing literature and patents. In urban multipath, severe occlusion, or strong interference scenarios, traditional algorithms suffer from a significant decline in perception performance due to their lack of context modeling capabilities. Leveraging the unique properties of image modeling, this paper enables deep models to identify contextual features within the time-frequency structure, thereby achieving robust perception prediction and effectively improving the system's practicality in non-line-of-sight (NLOS) and extremely low SNR conditions.

[0039] In complex urban environments, severe electromagnetic interference or long-distance weak echo scenarios, high communication reception bit error rate and low perception accuracy are long-standing core problems that have not yet been effectively solved. Traditional enhancement methods such as filtering, modulation reconstruction, space-time processing, etc. are often powerless under low SNR conditions, especially unable to take into account both communication and perception functions. The present invention significantly improves signal quality and target perception capabilities in low SNR environments by introducing an end-to-end signal enhancement method based on image super-resolution and diffusion model, successfully breaking through this technical bottleneck. For a long time, how to simultaneously achieve high-precision target detection and feature recognition in a communication system, and to build a unified processing model under a deep learning framework, has always been a difficult problem that the industry urgently needs to overcome. The present invention proposes for the first time to image-encode the received signal into an RGB image, and then process it through image super-resolution + diffusion model, which is further used for communication modulation symbol restoration and target physical parameter extraction, realizing a unified modeling and joint processing technology route in an integrated communication and perception system.

[0040] In traditional communication and perception systems, communication signal processing and perception information extraction are often viewed as two unrelated processes. The engineering community has long tended to adopt independent optimization strategies: the communication side focuses on waveform design and modulation recognition, while the perception side focuses on target echo modeling and parameter estimation. Consequently, there is a widespread technical bias toward isolating the two functions into modules and processing them separately, lacking a technical path for unified modeling and synergistic enhancement. This invention breaks down the barriers between traditional signal processing and image processing, proposing for the first time that communication and perception signals be image-encoded into RGB images after time-frequency transformation, and then introducing advanced image super-resolution algorithms and diffusion neural network structures for enhancement. Key physical information, such as frequency, amplitude, and phase, is then uniformly extracted from the image to recover communication symbols and target parameters. This approach significantly overcomes the conventional wisdom that communication and perception cannot be processed in a unified manner. Furthermore, the industry generally believes that image processing methods struggle to adapt to the non-stationary and noise interference characteristics of complex communication signals, and therefore image super-resolution or diffusion models are rarely introduced into the communication physical layer. This invention verifies for the first time that the image modeling method can effectively enhance the quality of communication signals and significantly improve perception accuracy under low SNR and strong interference conditions. It breaks through the industry's cognitive limitation that "image methods are not suitable for signal enhancement" and has significant technological leadership and disruptiveness.

[0041] Third, the signal enhancement algorithm based on image super-resolution proposed in this paper addresses core technical issues faced by integrated synaesthesia systems in low signal-to-noise ratio environments, including signal quality degradation, spectral image blurring, and a significant decrease in perception accuracy and communication demodulation capabilities. Traditional methods, which primarily rely on frequency domain filtering or feature engineering, struggle to extract accurate information from weak signals, limiting the system's adaptability in complex environments.

[0042] To address these issues, the present invention establishes a mathematical model that maps a mixed communication / perception signal into a three-channel RGB image after a short-time Fourier transform. These three channels correspond to the signal's amplitude, frequency, and phase information, enabling visualization of the signal's spectrum. This model not only preserves the signal's physical characteristics but also enhances its structural representation, facilitating optimization using image processing and deep learning methods.

[0043] This paper introduces the diffusion model as the core algorithm of image super-resolution reconstruction. By constructing a diffusion-reconstruction process of forward denoising and reverse denoising, it significantly improves the clarity and detail expression of spectrum images under low SNR conditions, enhances the recoverability of signals in complex environments such as multipath interference, occlusion and reflection, and provides a higher precision foundation for subsequent information extraction.

[0044] Finally, by inverse STFT operation and frequency domain restoration of the enhanced image, the application realizes high-accuracy demodulation of communication symbols and high-resolution estimation of key perception parameters such as target distance, speed and azimuth, significantly improves the communication stability and environmental perception ability of the system in complex scenes such as urban environment and unmanned platform, and has good engineering application prospect and technical popularization value. BRIEF DESCRIPTION OF DRAWINGS

[0045] Figure 1 A scene diagram of the integrated sensing and communication system of the application.

[0046] Figure 2 An algorithm flowchart of the application.

[0047] Figure 3 A key signal processing step flowchart in the signal enhancement algorithm based on image super-resolution.

[0048] Figure 4 An improved UNet neural network architecture introduced in the reverse diffusion process.

[0049] Figure 5 And Figure 6 A performance comparison simulation diagram of the proposed algorithm and other basic algorithms in communication bit error rate and perception accuracy under different signal-to-interference-and-noise ratios. DETAILED DESCRIPTION

[0050] In order to facilitate those skilled in the art to understand the technical content of the application, the content of the application is further explained below in combination with the drawings.

[0051] The present invention relates to the field of wireless communication technology, in particular to the fusion technology of wireless telepathy systems and artificial intelligence. Specifically, a signal enhancement algorithm design based on image super-resolution and its application system are disclosed. The invention aims to solve technical problems such as poor received signal quality, low spectrum image resolution, and reduced target recognition and demodulation performance in telepathy systems under low signal-to-noise ratio (SNR) environments. In current telepathy systems, the receiving end usually obtains a type of multimodal signal that integrates communication and perception information, including target reflection signals, environmental noise, and communication information. Due to the complexity of actual deployment scenarios, such as urban building obstruction and electromagnetic interference, the signal-to-noise ratio of the received signal is often reduced. Traditional filtering or demodulation methods are difficult to accurately restore communication information and perception features, limiting the improvement of the overall performance of the system. In response to the above problems, the present invention proposes a signal enhancement algorithm that integrates image super-resolution reconstruction. The core idea is to transform the received signal into an RGB image with semantic representation capabilities through short-time Fourier transform and other means, and then input this image into an AI-based image super-resolution model for enhancement, thereby improving the image's resolution and level of detail. Key communication and perception features such as amplitude, phase, and frequency can then be extracted from the higher-quality image to achieve signal denoising, enhancement, and accurate recognition. This method takes into account both the spatial structure of the image and the spectral characteristics of the signal. By introducing an image super-resolution reconstruction mechanism, it significantly improves the recoverability of synaesthesia information under low SNR conditions. It has high versatility and deployment flexibility, and is suitable for a variety of intelligent system scenarios that integrate communication and perception, such as drone detection, vehicle-mounted radar communications, and security monitoring.

[0052] The technical problems of the prior art solved by the present invention and the significant technical progress achieved are mainly reflected in the following aspects:

[0053] Technical issues:

[0054] 1. Severe signal quality degradation under low signal-to-noise ratio conditions: When facing complex targets, the received reflected signals are often severely attenuated, blocked, or subject to multipath interference. This blurs the amplitude, phase, and frequency characteristics of the original communication and perception signals, making high-precision target recognition and parameter estimation difficult.

[0055] 2. Traditional processing methods have limited ability to recover weak signals: Currently commonly used signal enhancement methods based on filtering, compressed sensing, or time-frequency analysis have difficulty in accurately recovering signal structure at the image level. This is especially true in scenarios where weak echoes and strong background noise coexist, making it difficult to guarantee perception performance.

[0056] 3. Lack of diffusion image denoising model for integrated sensing and communication system: Especially under non-ideal conditions such as low SNR and high interference, there is still a lack of a specially constructed image-level joint denoising enhancement algorithm with robustness and precision.

[0057] Technical progress:

[0058] 1. Map communication / sensing signals to three-channel RGB images, and fuse multi-modal feature information with physical meaning: Use short-time Fourier transform to convert the original one-dimensional complex signal to a two-dimensional time-frequency image, and use the red channel to represent the amplitude information, the green channel to represent the frequency distribution, and the blue channel to represent the phase information, to realize the unified modeling of communication and sensing, and provide high-dimensional input for subsequent image enhancement.

[0059] 2. Establish a mapping mechanism between the spectrum image and the communication / sensing parameters to realize physical consistency reconstruction: The image super-resolution model based on the diffusion model enhances the RGB image, significantly improves the image clarity and signal-to-noise ratio while preserving the original signal structure, further extracts key parameters such as target distance, angle, speed and communication modulation symbols, and ensures that the information after image enhancement has engineering usability and physical interpretability through 2D-FFT, angle spectrum analysis, demodulation recovery and other means.

[0060] 3. Propose a diffusion model enhancement for robustness under extreme low SNR: Compared with traditional filtering or deep learning methods, the diffusion model can gradually reconstruct the original signal spectrum from high-noise images, effectively improving the sensing accuracy and communication reliability in complex environments such as occlusion, multipath, and interference.

[0061] 4. Data simulation verifies the effectiveness of the algorithm: The present application verifies the effectiveness of the proposed algorithm through data simulation, proving the feasibility and superiority of the algorithm in practical applications.

[0062] In summary, the present application successfully solves the key problems of severe signal quality degradation under low SNR conditions and limited weak signal recovery capability of traditional processing methods in the prior art, and achieves significant technical progress and application value.

[0063] Embodiment: Low-altitude unmanned aerial vehicle monitoring scene in urban environment

[0064] 1. With the strengthening of urban airspace management, the monitoring of low-altitude drones faces challenges due to dense buildings and complex signal processing. This embodiment designs a dual-base synaesthesia-integrated low-altitude drone monitoring system that adapts to complex urban environments. The system can be deployed on building rooftops, base stations, or monitoring poles. The transmitter uses multi-beam synaesthesia signals for communication and perception tasks, while the receiver, located at a ground monitoring point or vehicle-mounted terminal, receives reflected signals and combines them with image super-resolution enhancement technology to extract high-precision perception information from obscured or low-signal-to-noise ratio channels, enabling the positioning, velocity estimation, and identification of low-altitude aircraft (such as drones).

[0065] 2. Data Collection: Transmitters and receivers are deployed in urban environments at varying heights and densities to collect target reflection signals and ambient background signals. The system operates under various weather conditions, including daytime and nighttime, sunny and rainy, to collect multipath signal data from low-altitude drones at different flight attitudes, paths, and speeds. The collected multipath signals are paired with the actual drone positions to generate a labeled signal dataset for subsequent algorithm training and validation.

[0066] 3. Signal preprocessing: The collected multipath signals are bandpass filtered and down-converted. The complex time domain signals are converted into a two-dimensional spectrogram (time-frequency matrix) using short-time Fourier transform. The spectrogram is normalized and mapped into an RGB three-channel image, where the R channel represents the normalized amplitude information, the G channel represents the normalized frequency axis, and the B channel represents the normalized phase information. This provides high-dimensional input for subsequent image enhancement.

[0067] 4. Using the improved diffusion model: The low-quality RGB image generated above is fed into a super-resolution network based on the diffusion model. This super-resolution model employs an improved residual attention network, with a loss function optimized using a combination of pixel, edge, and perceptual losses. By extracting multi-scale features and using an attention mechanism to enhance spectral detail, the model outputs a high-resolution, well-restored enhanced image. The enhanced, high-quality image is decoded into amplitude, frequency, and phase components, allowing key parameters (such as target distance, angle, velocity, and communication modulation symbols) to be extracted.

[0068] 5. System deployment: Deploy transmitters and receivers within the monitoring area, and set up monitoring terminals to identify, track, and share information about illegal drones.

[0069] This embodiment combines signal visualization with denoising processing based on a diffusion model, significantly improving the perception accuracy and communication capabilities of the low-altitude UAV monitoring system, and providing strong technical support for airspace safety in urban environments.

[0070] Figure 1This is a system scenario diagram of the multi-beam dual-base synaesthesia integration of the present invention. Figure 1 As shown in Figure 1, the algorithm of the present invention is based on the following system: the system consists of a transmitter and a receiver, located at different locations, forming a dual-base architecture. The transmitter sends a synaesthesia multiplexed signal with a specific coding or modulation method. This signal not only completes the communication task but also provides perceptual information to the receiver. The transmitting antennas are all uniform linear arrays to achieve multi-beam design. The transmitted signal can be expressed as ,in represents the baseband signal, represents the transmit beamforming vector, represents the phase offset vector, represents the power allocation factor, and Represent the beamforming vectors of the sensing and communication sub-beams respectively. By flexibly adjusting and , which can simultaneously meet the needs of communication and perception. The beamforming vector can be directly generated using the least squares method based on the mission requirements, taking into account the stability of the communication link and the resolution of target perception. It can achieve spatial multiplexing and low sidelobe radiation, and adapt to the needs of synaesthesia fusion in complex multi-target scenarios.

[0071] Figure 2 The algorithm flow chart of the present invention is as follows:

[0072] A1. Build a dual-base interaceptive integrated system. The system consists of a transmitter and a receiver, located in different locations, forming a dual-base architecture. The transmitter sends an interaceptive multiplexed signal with a specific coding or modulation scheme. This signal not only completes the communication task but also provides perceptual information to the receiver.

[0073] A2. The receiving end receives the echo signal reflected by the target and extracts the corresponding time domain complex signal sequence.

[0074] A3. Convert the received time-domain communication / perception signal into a two-dimensional time-frequency image through short-time Fourier transform, and construct a three-channel RGB image corresponding to the signal strength, frequency distribution, and phase structure. The specific steps are as follows:

[0075] A31. The receiving end performs a short-time Fourier transform on the acquired time-domain communication / perception mixed signal, and obtains a complex-valued spectrum in the time-frequency domain through a sliding window function and a fast Fourier transform.

[0076] A32, the modulus length of the complex spectrum is extracted as the amplitude spectrum of the signal, and is mapped to the red channel (R channel) of the image through linear normalization. The frequency axis obtained by STFT calculation is normalized and copied along the time axis to generate a two-dimensional frequency distribution matrix, which constitutes the green channel (G channel) of the image. The phase component of is standardized to constitute the blue channel (B channel) of the image.

[0077] A33, the above three channels are combined into a three-channel RGB image, and the image thus generated provides intuitive visualization of the spectral characteristics of the received signal, making it suitable for subsequent deep learning or pattern recognition tasks.

[0078] A4, due to the low SNR of the received signal, after conversion to an RGB image, a denoising neural network based on a diffusion model is used to enhance the image, thereby improving the recoverability and perceptual accuracy of the original signal. The specific steps are as follows:

[0079] A41, construct the forward diffusion process: the RGB image corresponding to the signal is taken as the initial input, and at each diffusion time step , Gaussian noise is gradually added to the image to construct an image sequence and form a forward diffusion chain from a clear image to a noisy image; this process is modeled by the following formula: . Wherein, represents a normal distribution with mean and variance , I represents the identity matrix, and represents the noise attenuation factor at the current time step.

[0080] A42, construct a conditional diffusion model to learn a neural network that predicts the original noise at a given time step and noisy image , where the network structure can include UNet, attention modules and residual modules; by minimizing the mean square error loss: .

[0081] A43, in the test phase, starting from a pure noise image , the trained denoising network is used to iteratively predict the residual noise and obtain the estimated value of the image by back propagation; each step of the inverse diffusion process is: , where . is the predicted noise at the current time step, .

[0082] A44, denoising is performed through the trained diffusion model, and finally a clear image is obtained at .​ The image has higher signal-to-noise ratio and clearer spectral details, thereby providing a more accurate image basis for subsequent frequency, amplitude and phase information extraction.

[0083] A5, in order to obtain the communication and perception information contained in the signal, the super-resolution image is subjected to image restoration and inverse STFT operation, and is converted into an enhanced radio frequency signal. The specific steps are as follows:

[0084] A1, inverse normalization and inverse transformation are performed on the reconstructed high-resolution image, and an enhanced radio frequency signal is recovered by using inverse STFT algorithm, so that the recovered signal in time domain has higher signal-to-noise ratio and stronger structural characteristics.

[0085] A2, the reconstructed time domain signal is subjected to traditional communication signal processing, and the symbol sequence is demodulated and recovered.

[0086] A3, the enhanced signal is used for target detection and parameter estimation. The target angle is extracted by using high-resolution azimuth angle estimation algorithm multiple signal classification (MUSIC), and two-dimensional discrete Fourier transform (2D-DFT) algorithm is used to perform inverse DFT and DFT on each row and each column of the received signal matrix, respectively, so that the two-dimensional distance-velocity spectrum of the target can be obtained, and the horizontal coordinate and the vertical coordinate represent the distance position and the velocity information of the target, respectively.

[0087] Figure 3 The key signal processing steps in the signal enhancement algorithm based on image super-resolution proposed in the present application are forward diffusion process and reverse diffusion process. The process is based on the theoretical framework of diffusion model, combines image generation and denoising technology, and significantly improves the quality of communication / perception image under low signal-to-noise ratio condition. The forward diffusion process refers to adding Gaussian noise to the original RGB image Step by step, a series of gradually degraded images are generated until the moment of pure noise image The process simulates the process of gradual deterioration of the signal from the clear state in the complex physical environment, and the spectral details, phase structure and amplitude characteristics in the image are gradually covered. Subsequently, an improved UNet neural network is introduced as a denoising generation module, and in the reverse diffusion process, the network starts from the last state image , and gradually reconstructs the image close to the original noise-free signal . The denoising process consists of predictions of multiple time steps, and each step estimates the previous state based on the current state, thereby restoring the true structure of the signal. The advantages of this method are: not only does it introduce time step embedding to control the intensity of noise addition in the diffusion stage, but it also uses the modeling capabilities of deep networks to restore the frequency domain structure and image texture in the denoising stage, while also maintaining the decomposability and interpretability of communication and perception information. Through the collaborative optimization of the positive and negative stages, the system can retain key features such as target frequency, phase, amplitude, etc. under extremely low signal-to-noise ratio conditions, providing a high-quality data foundation for subsequent modulation recognition and target detection.

[0088] Figure 4 This is the improved UNet neural network architecture introduced in the reverse diffusion process proposed in this paper. This architecture consists of an encoder, a decoder, a skip connection module, a residual connection module, and an attention module, which together achieve multi-scale feature extraction and efficient information reconstruction. The encoder extracts deep features of the image layer by layer and includes the following modules:

[0089] Single convolution block: includes convolution layer, activation function and normalization operation, used to extract basic image texture.

[0090] Residual block: Introduces cross-layer residual connections to improve the stability of deep network training.

[0091] Attention module: used to focus on perception-related feature areas and guide feature extraction to focus more on target areas, such as signal texture edges and frequency jump points.

[0092] Downsampling operation: Use convolution with a stride of 2 instead of traditional pooling to achieve feature dimensionality reduction and space compression.

[0093] The decoder restores spatial resolution layer by layer to reconstruct a high-quality image. Each decoder level contains multiple residual blocks and attention modules to fuse multi-scale features. Upsampling is also used for image restoration, using nearest neighbor interpolation and convolution or transposed convolution. The encoder features at each level are transferred to the decoder at the corresponding layer via skip connections, preserving lower-level spatial texture information and improving detail recovery. Finally, a 1x1 convolution is performed to output the enhanced image channels.

[0094] This network introduces a residual structure to prevent gradient vanishing and improve training convergence speed. It also incorporates an attention module to enhance feature selectivity and suppress redundant interference. It effectively supports high-quality reconstruction of complex perceptual images (such as spectrograms and phase maps) and is suitable for super-resolution and denoising tasks under low SNR conditions.

[0095] In the simulation process of this invention, it is assumed that a transmitter and a receiver are located at different locations, forming a dual-base architecture to perform the task of synergy integration, realizing communication cooperation while monitoring low-altitude drones, and the detection range covers the airspace within a radius of 500 meters on the receiver side. The specific process is as follows:

[0096] Figure 5 and Figure 6 The simulation diagram shows the performance comparison between the proposed algorithm and other basic algorithms in terms of communication bit error rate and perception distance estimation accuracy under different signal-to-noise ratios. Figure 3 and Figure 4 It can be seen that the traditional algorithm shows the worst error performance under all SNR conditions because its signal processing process does not introduce advanced filtering or denoising strategies. The method based on adaptive filtering can gradually suppress noise during iteration, but the convergence speed is significantly reduced in extremely low SNR scenarios, and the performance is unstable. Although the neural network-based method is good at learning noise patterns in training data, its generalization ability under extreme noise is limited, and its feature extraction shows a certain lack of robustness in complex environments. The algorithm proposed in the present invention integrates the denoising strategy guided by domain knowledge and the multi-scale feature analysis technology through an innovative network structure design, overcoming the shortcomings of traditional adaptive algorithms and pure data-driven methods. The method shows excellent robustness throughout the entire SNR range, and is particularly suitable for communication and perception tasks in low SNR environments. It has significant practical value and engineering feasibility.

[0097] Those skilled in the art will appreciate that the embodiments described herein are intended to aid the reader in understanding the principles of the present invention, and it should be understood that the scope of the present invention is not limited to such specific descriptions and embodiments. Various modifications and variations are readily apparent to those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are intended to be included within the scope of the claims.

Claims

1. A signal enhancement method based on image super-resolution, characterized in that: The steps include: Step 1: The transmitter sends a synaesthesia integrated signal with a coding or modulation structure; Step 2: The receiving end receives the mixed signal from the target or the environment and converts it into a two-dimensional time-frequency image through short-time Fourier transform; Step 3: Construct the image into an RGB image with three elements: amplitude, phase, and frequency. The specific steps are as follows: A1. The receiving end performs a short-time Fourier transform on the acquired time-domain communication / perception mixed signal, and obtains a complex-valued spectrum X(f,t) in the time-frequency domain through a sliding window function and a fast Fourier transform. A2. Extract the modulus length |X(f,t)| of the complex spectrum as the signal's amplitude spectrum and map it to the image's red channel (R channel) through linear normalization. Take the frequency axis calculated by STFT, normalize its values, and replicate it along the time axis to generate a two-dimensional frequency distribution matrix to form the image's green channel (G channel). Normalize the phase component of X(f,t) to form the image's blue channel (B channel). A3. Merge the above three channels into a three-channel RGB image. The resulting image provides an intuitive visualization of the spectral characteristics of the received signal, making it suitable for subsequent deep learning or pattern recognition tasks; Step 4: Input the RGB image into an image super-resolution neural network for enhancement to improve image resolution and details; Step 5: Extract communication and perception parameters from the enhanced image to achieve signal denoising and enhancement. The specific steps are as follows: A1. Perform inverse normalization and inverse transformation on the reconstructed high-resolution image, and apply algorithms such as inverse STFT to restore the image to the enhanced RF signal, ensuring that the restored signal in the time domain has a higher signal-to-noise ratio and stronger structural features; A2. Perform traditional communication signal processing on the reconstructed time domain signal to demodulate and recover the symbol sequence; A3. Use the enhanced signal for target detection and parameter estimation. The target angle is extracted using the high-resolution azimuth estimation algorithm Multiple Signal Classification (MUSIC). The two-dimensional Discrete Fourier Transform (2D-DFT) algorithm is used to perform inverse DFT and DFT on each row and column of the received signal matrix, respectively. This yields a two-dimensional range-velocity spectrum of the target, where the horizontal and vertical axes represent the target's range, position, and velocity, respectively.

2. The method according to claim 1, characterized in that The image super-resolution neural network is a diffusion model network, which includes a forward diffusion process and a reverse denoising and reconstruction process. The forward process adds noise, and the reverse process estimates the noise residual based on the neural network and gradually restores a clear image.

3. The method according to claim 1, characterized in that The red channel of the image is the normalized signal amplitude, the green channel is the frequency distribution, and the blue channel is the signal phase.

4. The method according to claim 1, wherein The transmitting end and the receiving end are physically separated to form a dual-base cooperative detection structure, and the detection robustness is improved through spatial diversity.

5. A signal enhancement system based on image super-resolution using the method according to any one of claims 1 to 4, characterized in that: include: Transmitter module, used to generate and send signals with communication and perception functions; A receiving module, used for receiving the mixed echo signal and performing short-time Fourier transform; An image construction module, used to construct the spectrum information into a three-channel RGB image; Image enhancement module, used to improve image resolution through diffusion model; The parameter extraction module is used to obtain communication and perception parameters from the enhanced image.

6. The system according to claim 5, characterized in that The transmitting module uses a uniform planar array antenna to perform multi-beam transmission and allocates power and phase of a sensing beam and a communication beam respectively, thereby achieving spatial multiplexing and interference suppression.

7. The system according to claim 5, characterized in that The receiving module is suitable for an orthogonal frequency division multiplexing structure and completes frequency domain receiving modeling in a multipath and Doppler environment.

8. The system according to claim 5, wherein: The image enhancement module includes a conditional diffusion network, and the network structure includes a UNet backbone structure, an attention module and a residual connection.

9. The system according to claim 5, characterized in that It also includes a restoration module for restoring the enhanced image to an enhanced radio frequency signal through inverse short-time Fourier transform, and performing demodulation and perception processing.

10. The restoration module according to claim 9, characterized in that: The target's position information is extracted using a multiple signal classification algorithm, and the target's range-velocity spectrum is obtained using a two-dimensional discrete Fourier transform.

Citation Information

Patent Citations

  • Multi-antenna blind modulation identification method based on short-time Fourier transform time-frequency analysis

    CN111901267A

  • Audio generation method and system

    CN115700883A