Underwater Chirp spread spectrum communication demodulation method and system based on deep learning
By employing a deep learning-based underwater chirp spread spectrum demodulation method, the problems of high bit error rate and high energy consumption in underwater communication systems under complex channels are solved, achieving low power consumption and high robustness in communication.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FUZHOU UNIV
- Filing Date
- 2026-01-28
- Publication Date
- 2026-05-01
AI Technical Summary
Existing underwater communication systems suffer from high bit error rates and high energy consumption when facing complex underwater acoustic channels. Traditional demodulation techniques are unable to effectively capture the subtle time-frequency features of signals under low signal-to-noise ratio conditions, making it difficult to balance communication reliability and energy consumption.
A deep learning-based underwater Chirp spread spectrum communication demodulation method is adopted. Multi-scale feature extraction is performed through signal preprocessing and deep learning models, and time-frequency feature recognition is performed using SwingTransformer, thereby reducing hardware dependence and transmission power requirements.
It significantly reduces transmission power consumption, lowers the average symbol error rate by 14.6%, maintains stable communication under low signal-to-noise ratio conditions, and is suitable for energy-constrained underwater sensor networks.
Smart Images

Figure CN121967132A_ABST
Abstract
Description
A Deep Learning-Based Demodulation Method and System for Underwater Chirp Spread Spectrum Communication Technical Field
[0001] This invention relates to the field of underwater wireless communication technology, and in particular to an underwater chirp spread spectrum communication demodulation method and system based on deep learning. Background Technology
[0002] Underwater wireless communication, as a core technology supporting critical tasks such as marine environmental monitoring, deep-sea resource exploration, and disaster early warning, is gaining increasing strategic value and experiencing continuous market demand expansion. Against this backdrop, underwater wireless communication technology, due to its deployment flexibility, network scalability, and ability to adapt to complex underwater environments, has become a key enabling technology for building underwater wireless sensor networks (UWSNs), playing an irreplaceable role in marine scientific exploration, resource development, and environmental protection.
[0003] Among the various existing underwater wireless communication technologies, each has its own applicable scenarios and limitations. Radio frequency (RF) communication suffers severe attenuation in water, resulting in extremely limited propagation distances, typically suitable only for very short-range communication. While optical communication possesses the potential for high transmission rates, it heavily relies on stringent line-of-sight transmission conditions and is susceptible to interference from suspended particles in the water, limiting its applicability. Magnetic induction communication faces technical bottlenecks such as complex hardware implementation, high deployment costs, and short operating distances. In contrast, acoustic communication, with its excellent propagation characteristics in water, lower environmental adaptability requirements, and relatively mature system architecture, is widely considered the most practical and reliable underwater communication method, particularly suitable for large-scale, long-distance underwater applications.
[0004] However, underwater acoustic channels are among the most complex communication channels on Earth, their characteristics equivalent to a multidimensional filter exhibiting time-varying, space-varying, and frequency-varying properties. Sound waves attenuate rapidly with increasing frequency as they propagate in water, severely limiting usable bandwidth and resulting in low channel capacity. Simultaneously, due to surface and seabed reflections and the inhomogeneity of the ocean medium, signals experience significant multipath effects, leading to waveform distortion and time spread, severely impacting communication timing synchronization and data integrity. Furthermore, the noise composition in real underwater environments is highly complex, including not only non-uniformly distributed Gaussian colored noise in the frequency domain but also ubiquitous burst impulse interference in the time domain, such as transient noise introduced by ship navigation and biological activity. These factors collectively lead to extremely high bit error rates for underwater communication systems.
[0005] To address these challenges, traditional underwater acoustic communication systems typically employ two main strategies to improve communication reliability. One strategy relies on high-performance hardware, such as commercial modems, which integrate complex signal processing algorithms and high-precision hardware components to maintain stable communication under harsh channel conditions. However, these devices are often complex, expensive, and bulky, limiting their application in large-scale deployments, resource-constrained systems, or miniaturized platforms (such as micro-submarines and low-cost sensor nodes). The other strategy increases transmitter power to enhance the signal-to-noise ratio (SNR), thereby combating channel attenuation and noise interference. However, this approach leads to a significant increase in system power consumption; in some typical applications, the transmit power can be more than 100 times the receive power. For energy-constrained underwater sensor networks or long-term monitoring systems, excessively increasing transmit power is not feasible, severely limiting the system's endurance and lifespan.
[0006] In recent years, chirp spread spectrum (CSS) modulation techniques have attracted widespread attention due to their excellent resistance to multipath fading, Doppler shift, and environmental noise in long-range wireless communications such as LoRa. This technology achieves energy-distributed transmission through a wideband linearly swept frequency signal, maintaining high demodulation reliability even under low signal-to-noise ratio conditions, providing a new technical path for resource-constrained underwater communications. Existing inventions have demonstrated that CSS technology has good application potential in underwater scenarios, effectively balancing communication robustness and system complexity.
[0007] While CSS technology has demonstrated certain advantages in underwater communication, existing solutions still have significant limitations. On the one hand, most CSS-based underwater communication systems still rely on repetitive transmissions or complex modulation mechanisms to combat channel noise, leading to a significant increase in power consumption at the transmitter. This "trading power consumption for reliability" strategy is unsustainable in energy-constrained underwater applications. On the other hand, most existing demodulation techniques follow the traditional "dechirping + peak detection" process, which can only extract coarse-grained features of the signal and is difficult to capture subtle but crucial structural information in the time and frequency domains. Especially in complex and variable underwater environments, weak signals are often masked by background noise, and traditional methods are prone to misjudgment due to insufficient feature utilization.
[0008] Therefore, there is an urgent need for a new demodulation method that can reduce dependence on transmission power and fully exploit the inherent time-frequency characteristics of the received signal to support the practical deployment of low-power, highly robust underwater communication systems. Summary of the Invention
[0009] In view of this, the purpose of this invention is to provide an underwater chirp spread spectrum communication demodulation method and system based on deep learning. This method improves demodulation performance through intelligent signal processing technology, overcoming the limitations of traditional technologies in terms of performance, cost, and energy consumption, and providing a new technical path for the widespread application of underwater communication technology.
[0010] To achieve the above objectives, the present invention adopts the following technical solution: a deep learning-based underwater chirp spread spectrum communication demodulation method, comprising the following steps:
[0011] Step 1: Signal preprocessing; Receive underwater Chirp spread spectrum communication signals, and sequentially perform packet detection, time-frequency offset correction, and symbol conversion operations to obtain the input features of the deep learning model;
[0012] Step 2: Demodulation based on deep learning; Input the input features obtained in Step 1 into a preset deep learning model, and use the deep learning model to extract and identify time-frequency features at multiple scales, outputting the symbol category probability distribution to complete signal demodulation.
[0013] In a preferred embodiment, the chirp signal is generated by superimposing a discrete frequency offset onto a base up-chirp signal. To achieve signal modulation, different The values correspond to different signs, and the modulated signal is represented as follows:
[0014] (1)
[0015] in, Based on the up-frequency chirped signal, This represents the frequency offset corresponding to the data symbol. As an offset component, the starting frequency of the fundamental up-frequency chirp is adjusted to... The sweep rate in the formula Due to bandwidth Duration of the Chirp signal Joint decision;
[0016] The spreading factor SF, as a core parameter, defines the size of the coded symbol set. The signal bandwidth BW and the chirp duration T satisfy the following relationship:
[0017] (2)
[0018] In signal bandwidth Under fixed conditions, spreading factor The increase will cause the chirp symbol duration to increase. Corresponding extension; each Chirp symbol is characterized by a differentiated configuration of its initial frequency. A variety of distinct symbols.
[0019] In a preferred embodiment, the packet structure in packet detection includes a preamble, a synchronization word SFD, and a data payload; the preamble consists of 6 basic up-chirps for detecting the presence of packets and coarse-grained signal synchronization; the synchronization word SFD consists of 2.25 basic down-chirps for fine-grained signal synchronization and eliminating carrier frequency offset (CFO); and the data payload is used to transmit encoded valid data.
[0020] In a preferred embodiment, step 1, packet detection specifically includes: the receiver dividing the received signal into sliding windows with a symbol duration T as the window length, and the signal in each window is called a "window signal"; performing dechirping and peak detection operations on each window signal; when the energy peaks of multiple window signals appear at the same frequency position, it indicates that a data packet has been detected.
[0021] In a preferred embodiment, the time-frequency offset correction in step 1 specifically includes: when the sliding window is aligned with a single Chirp symbol, after performing dechirping and peak detection on the basic up-chirp in the preamble, its FFT peak should appear at the position of frequency 0; by detecting the deviation between the peak position and the zero-frequency position in the window signal, the time offset is calculated, and the signal is time-domain shifted to achieve alignment of the symbol boundary.
[0022] In a preferred embodiment, the symbol transformation in step 1 specifically includes: after aligning the window signal with the Chirp symbol boundary, dividing each Chirp symbol sequentially using the length of a single symbol as the length of the sliding window, performing a short-time Fourier transform (STFT) to generate a two-dimensional time-frequency spectrum, and extracting its real part as the input feature of the neural network.
[0023] In a preferred embodiment, step 2 specifically includes: the deep learning model receives the real part of the time-frequency map generated in the preprocessing stage as input, and first performs frequency domain feature enhancement through the time-frequency transformation module FTB; subsequently, the input signal is divided into multiple local token sequences after block segmentation and linear embedding processing, and enters a four-stage hierarchical processing flow; wherein, the first, third and fourth stages each contain two basic processing units, and the second stage is expanded to six repeating units to enhance the core feature extraction capability; each processing unit is composed of the FTB module and the SwinTransformer block in collaboration; the FTB module is responsible for enhancing local time-frequency characteristics, and the SwinTransformer uses a shift window mechanism to model long-distance dependencies and achieve global context awareness; a patch merging layer is introduced between each stage to gradually downsample the spatial resolution of the feature map, while doubling the number of channels to construct a pyramid-like structure.
[0024] This invention also provides a deep learning-based underwater Chirp spread spectrum communication demodulation system, comprising a transmitter, a receiver, and a signal processing module; the transmitter is used to generate and transmit Chirp spread spectrum communication signals, and includes a main control unit, an audio power amplifier, and a vibration motor; the receiver is used to receive underwater acoustic Chirp signals and convert them into electrical signals, and includes a piezoelectric ceramic sensor, an audio decoder, and a power supply device; the signal processing module is used to execute the deep learning-based underwater Chirp spread spectrum communication demodulation method according to any one of claims 1-7.
[0025] In a preferred embodiment, the main control unit is a Raspberry Pi used to generate a Chirp spread spectrum modulation signal containing a preamble, a synchronization word SFD, and a data load. The audio power amplifier is used to amplify the power of the modulation signal, and the vibration motor is used to convert the electrical signal into an underwater acoustic Chirp signal.
[0026] In a preferred embodiment, the piezoelectric ceramic sensor is used to pick up underwater vibration signals, the audio decoder is used to digitize the signals, and the power supply device supplies power to the receiver.
[0027] Compared with existing technologies, the present invention has the following beneficial effects: As an underwater chirp spread spectrum communication method and system based on deep learning, the present invention has the following significant advantages: First, the hardware cost is greatly reduced. The transmitter only requires a common commercial vibration motor, and the receiver uses an inexpensive piezoelectric ceramic ring as a sensor, eliminating the need for dedicated acoustic transducers, filtering circuits, or active noise reduction modules; Second, the transmission power consumption is significantly reduced. Under the same decoding accuracy requirements, up to 40% of the transmission power can be saved, effectively extending the service life of underwater nodes, which is particularly suitable for long-term monitoring tasks with limited energy; Finally, the demodulation robustness is improved. Compared with the traditional "de-chirp + peak detection" method, the symbol error rate is reduced by an average of 14.6%, and stable communication can still be maintained in harsh underwater environments with a signal-to-noise ratio as low as -10dB. Attached Figure Description
[0028] Figure 1 is a schematic diagram of the overall framework of a preferred embodiment of the present invention;
[0029] Figure 2 is a schematic diagram of the signal processing flow design of a preferred embodiment of the present invention;
[0030] Figure 3 is a schematic diagram of the data packet structure of a preferred embodiment of the present invention;
[0031] Figure 4 is a schematic diagram of the demodulation model structure of a preferred embodiment of the present invention;
[0032] Figure 5 is a schematic diagram of hardware deployment according to a preferred embodiment of the present invention, wherein (a) is the deployment of the sender and (b) is the deployment of the receiver;
[0033] Figure 6 is a top view of an underwater experimental scene according to a preferred embodiment of the present invention, wherein (a) is a small pond and (b) is a slow-moving stream;
[0034] Figure 7 shows the relationship between SER and different bandwidths (BWs) and diffusion factors (SFs) in a preferred embodiment of the present invention, where (a) represents different SFs and (b) represents different BWs;
[0035] Figure 8 shows the SER performance variation curves under different SNRs in the preferred embodiment of the present invention;
[0036] Figure 9 is a schematic diagram comparing the SER performance of a preferred embodiment of the present invention under different water bodies;
[0037] Figure 10 is a schematic diagram of the performance variation curves under different transmission powers in a preferred embodiment of the present invention. Detailed Implementation
[0038] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0039] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0040] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations according to this application; as used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise; furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0041] The underwater chirp spread spectrum communication demodulation method based on deep learning, as shown in Figure 1-10, includes the following steps:
[0042] Step 1: Signal Preprocessing; The underwater Chirp spread spectrum communication signal is received, and packet detection, time-frequency offset correction, and symbol conversion are performed sequentially to obtain the input features for the deep learning model. The signal preprocessing stage, as the system's front-end processing unit, is responsible for converting the raw received signal into a high-quality time-frequency representation. This stage first employs a sliding window mechanism for packet detection, utilizing the redundant upsampling chirp structure in the preamble to achieve reliable signal detection under low signal-to-noise ratio conditions. Subsequently, time-frequency alignment is completed through a dual correction mechanism—precise timing of symbol boundaries is achieved based on the preamble, and the downsampling chirp in the synchronization word (SFD) is used to accurately estimate and compensate for carrier frequency offset. Finally, a short-time Fourier transform (STFT) is performed on the corrected symbols, retaining the real part as the neural network input.
[0043] Step 2: Demodulation based on deep learning; The input features obtained in Step 1 are input into a preset deep learning model. The deep learning model performs multi-scale extraction and recognition of time-frequency features, outputting the symbol category probability distribution to complete signal demodulation. The demodulation stage uses SwinTransformer as the backbone network, constructing a four-stage processing flow. Through hierarchical window partitioning and shift window attention mechanisms, multi-scale feature capture from local details to global structure is achieved while reducing computational complexity. A time-frequency transform module (FTB) is embedded in each processing unit to dynamically highlight discriminative time-frequency patterns, effectively addressing the problem of ambiguous signal trajectories in underwater environments. This design achieves high-precision symbol recognition even under low signal-to-noise ratio conditions, successfully realizing a key shift from "hardware dependence" to "algorithm intelligence," providing an innovative solution for resource-constrained underwater communication scenarios.
[0044] Specifically, a chirp signal is a spread spectrum signal that performs a linear frequency scan of a preset bandwidth BW within a fixed time T. Due to its excellent spreading gain, strong resistance to multipath interference, and high demodulation reliability even in low signal-to-noise ratio (SNR) environments, it is widely used in long-distance wireless communication systems. Specifically, it is achieved by superimposing a discrete frequency offset onto a basic up-chirp signal. To achieve signal modulation, different The modulated signal can be represented as follows, with different values corresponding to different symbols:
[0045] (1)
[0046] in, Based on the up-frequency chirped signal, This represents the frequency offset corresponding to the data symbol. As an offset component, the starting frequency of the fundamental up-frequency chirp is adjusted to... The sweep rate in the formula Due to bandwidth Duration of the Chirp signal A joint decision.
[0047] The spreading factor (SF) is a core parameter that defines the size of the set of coded symbols. The signal bandwidth (BW) and the chirp duration (T) satisfy the following relationship:
[0048] (2)
[0049] In signal bandwidth ( With a fixed spread factor, the spread factor ( An increase in ) will lead to an increase in the duration of the chirp symbol ( The corresponding extension occurs. At this point, each Chirp symbol can be characterized by a differentiated configuration of its initial frequency. A variety of distinct symbols.
[0050] The traditional dechirp method, currently the most widely used demodulation technique, will serve as the performance benchmark for this invention. Its core principle involves mixing a locally generated reference chirp signal with the received signal to convert the linearly swept frequency signal into a time-domain pulse, which is then used for symbol recognition through peak detection. The mathematical expression for this method is:
[0051] (3)
[0052] In the formula For basic upward chirping The conjugate of the two, multiplied together, yields... This represents the encoded shift frequency offset. By performing a Fourier transform (FFT) operation on it, in... The frequency shows a significant peak.
[0053] Regarding the design of the data packet structure: Similar to most communication systems using CSS modulation, the data packet structure designed in this scheme consists of three parts: a preamble, a synchronization word (SFD), and a data payload, as shown in Figure 3. The preamble consists of 6 basic up-chirps, used to detect the presence of the data packet and for coarse-grained signal synchronization; the synchronization word (SFD) consists of 2.25 basic down-chirps, used for fine-grained signal synchronization and to eliminate carrier frequency offset (CFO); the data payload is used to transmit the encoded valid data.
[0054] Step 1 specifically includes: Signal preprocessing, a crucial pre-processing step in the underwater Chirp spread spectrum communication system, is responsible for converting the raw received signal into a high-quality time-frequency representation suitable for subsequent deep learning demodulation. This step sequentially completes packet detection, time-frequency offset correction, and symbol feature conversion, ensuring that structurally clear and time-aligned input features can still be obtained in complex underwater environments. The specific processing flow is as follows:
[0055] 1) Packet Detection: The receiver divides the received signal into sliding windows with a symbol duration T as the window length. The signal within each window is called the "window signal". A "dechirping + peak detection" operation is performed on each window signal. Since the preamble consists of multiple known redundant basic up-chirps, when the energy peaks of multiple window signals appear at the same frequency position, it indicates that a data packet has been detected.
[0056] 2) Time-Frequency Offset Correction: When the sliding window is precisely aligned with a single chirp symbol, after performing "dechirping + peak detection" on the basic up-chirp in the preamble, its FFT peak should appear at the position of frequency 0. In practical applications, the time offset is calculated by detecting the deviation between the peak position and the zero-frequency position in the window signal, and the signal is time-domain shifted to achieve symbol boundary alignment. The synchronization word (SFD) consists of 2.25 basic down-chirps and is used to estimate and compensate for the carrier frequency offset (CFO).
[0057] 3) Symbol Transformation: After aligning the window signal with the Chirp symbols, the length of a single symbol is used as the length of the sliding window. Each Chirp symbol is then divided sequentially, and a Short-Time Fourier Transform (STFT) is performed to generate a two-dimensional time-frequency spectrum. The real part is then extracted as the input feature for the neural network. This design preserves the core linear frequency sweep structure of the Chirp signal while avoiding complex number operations, making it suitable for lightweight real-number neural network architectures.
[0058] Step 2 specifically includes: The deep learning-based demodulation stage is the core component of this system for achieving high-precision symbol recognition. Its key objective is to construct an efficient deep learning neural network for Chirp signal demodulation. As shown in Figure 4, this stage employs a hierarchical neural network architecture, achieving end-to-end mapping from time-frequency features to symbol categories. The model uses the real part of the two-dimensional time-spectrum graph of a single Chirp symbol as the network input, employs a SwinTransformer as the backbone feature extractor, and integrates a time-frequency transform module (FTB). Finally, after layer normalization and flattening, the features are fed into a fully connected layer, and the probability distribution of each symbol category is output through the Softmax function.
[0059] The model receives the real part of the time-frequency map generated in the preprocessing stage as input. First, it undergoes frequency domain feature enhancement through a Time-Frequency Transform (FTB) module. Subsequently, the input signal is divided into multiple local token sequences through block segmentation and linear embedding, entering a four-stage hierarchical processing flow. The first, third, and fourth stages each contain two basic processing units, while the second stage expands to six repeating units to enhance the extraction of core features. Each processing unit is composed of an FTB module and a SwinTransformer block: the FTB enhances local time-frequency characteristics, while the SwinTransformer models long-distance dependencies through a shift window mechanism, achieving global context awareness. A patch merging layer is introduced between stages to progressively downsample the spatial resolution of the feature map while doubling the number of channels, constructing a pyramid-like structure. This facilitates the layer-by-layer abstraction of semantic information, forming a hierarchical feature representation from details to the global picture.
[0060] SwinTransformer: As the core component of this architecture, SwinTransformer achieves a balance between local perception and global modeling while reducing computational complexity through hierarchical window partitioning and a shift-window attention mechanism. Its layer-by-layer downsampling structure supports multi-scale feature extraction, making it particularly suitable for processing two-dimensional time-frequency inputs with strong spatiotemporal correlation. This architecture not only effectively captures the linear frequency sweep pattern of chirp symbols but also maintains high sensitivity to small frequency shifts and local distortions, which is crucial for distinguishing multi-level symbols with similar frequency shifts.
[0061] Time-Frequency Transform (FTB) Module: As a lightweight feature enhancement unit, the FTB is designed based on a simplified self-attention mechanism. By modeling the intrinsic correlation between frequency bands and time periods, it dynamically highlights discriminative time-frequency patterns while suppressing interference responses in irrelevant regions. In complex underwater environments, noise masking and channel distortion often lead to unclear sweep trajectories of chirp symbols. This module, by focusing on the temporal consistency within local frequency bands, can more stably capture the essential structural features of signals, significantly improving the model's ability to perceive weak signals.
[0062] Finally, the high-level features are flattened after layer normalization and fed into a fully connected classification head. The probability distribution of each symbol class is then output through the Softmax function. This design achieves in-depth mining and accurate identification of time-frequency features while maintaining a lightweight model.
[0063] This invention constructs an underwater chirp spread spectrum communication demodulation system based on deep learning, which fully realizes the end-to-end communication link from signal transmission and underwater transmission to reception and demodulation. The overall hardware deployment is shown in Figure 5. This system breaks through the dependence of traditional underwater communication on dedicated acoustic transducers and high-cost hardware. By using a combination of commercial components, it significantly reduces system complexity and deployment costs while ensuring communication performance, making it suitable for long-term monitoring applications of large-scale underwater sensor networks.
[0064] The transmitter consists of a 10mm diameter commercially available miniature vibration motor, precisely encapsulated in a custom-designed waterproof silicone sleeve to ensure safe and stable operation in underwater environments. The vibration motor is driven by a high-efficiency PAM8406 audio power amplifier, and system power consumption is monitored in real-time by a Juwei A3 digital power meter for easy evaluation of energy efficiency. The signal source is generated by a Raspberry Pi ZeroW as the main control unit, and the precise generation of the CSS modulated signal, including the complete data packet structure of preamble, synchronization word (SFD), and data payload, is achieved through Matlab.
[0065] The receiver employs a low-cost piezoelectric ceramic ring (70mm outer diameter, 60mm inner diameter, 50mm height) as an acoustic sensor to pick up vibration signals in the water. This element features a wide frequency response range and high sensitivity, effectively capturing weak mechanical vibration signals in the water. The vibration signals are sampled with high precision (up to 96kHz sampling rate) and digitized using a Raspberry Pi Zero equipped with an IQaudioCodec audio expansion board. The entire receiving system is powered by a portable power bank, requiring no external power grid and offering excellent flexibility for field deployment.
[0066] As shown in Figure 1, the overall architecture is divided into three collaborative layers: physical layer communication mechanism, lightweight hardware implementation architecture, and intelligent signal processing flow.
[0067] At the physical layer communication mechanism level, this invention deeply analyzes the essential characteristics and challenges of underwater acoustic channels: when chirp signals propagate in water through mechanical vibrations, they must overcome multiple interferences such as multipath effects, frequency-selective fading, strong background noise, and time-varying propagation delays. This framework is designed based on these non-ideal channel conditions to ensure the system has practical application capabilities.
[0068] At the lightweight hardware implementation architecture level, this invention adopts a lightweight design concept, significantly reducing the system deployment threshold and cost. Specifically, it uses a Raspberry Pi ZERO equipped with an audio decoder as the core processing platform. The transmitter uses only a commercially available vibration motor to excite the acoustic chirp signal, and the receiver uses a low-cost piezoelectric ceramic ring to collect the weak vibration response. The entire system requires no dedicated modulation / demodulation chips, filtering circuits, or active noise reduction modules. This design not only significantly reduces system power consumption but is also suitable for long-term unattended underwater sensor networks, providing a feasible solution for large-scale underwater IoT deployments.
[0069] At the intelligent signal processing level, this invention constructs a complete end-to-end processing chain from signal perception to symbol recovery. The received signal first undergoes high-precision packet detection and adaptive preprocessing, and then is converted to a time-frequency representation via Short-Time Fourier Transform (STFT). Unlike traditional "de-chirping + peak detection" methods, this project innovatively introduces a deep learning demodulation model. Through fine-grained modeling and analysis of time-frequency features, it effectively distinguishes noise interference from useful signal components, achieving high-precision symbol recognition even under signal-to-noise ratio conditions as low as -10dB. This processing flow represents a crucial shift from "hardware dependence" to "algorithm intelligence," opening a new path for improving underwater communication performance.
[0070] Experimental evaluation indicators
[0071] To objectively measure the performance advantages of the proposed communication framework, traditional dechirp is used as the benchmark method for comparison. Dechirp, as the most widely used CSS demodulation technique in current underwater acoustic communication, has advantages such as simple implementation and low computational overhead, making it suitable as a performance reference. The evaluation mainly uses the symbol error rate (SER) as the core indicator.
[0072] Symbol Error Rate (SER): SER is defined as the ratio of erroneously demodulated chirped symbols to the total number of transmitted symbols. It is used to evaluate the robustness of a demodulation model to noise.
[0073] I. Experimental Setup
[0074] To comprehensively evaluate the performance and adaptability of the proposed communication framework in real underwater environments, a systematic outdoor test experiment was designed and implemented.
[0075] As shown in Figure 6, the experimental scenarios selected two representative natural aquatic environments—an artificial shallow pond and a natural slow-moving stream. These two scenarios simulated the still water and slow-flowing water conditions commonly encountered in practical applications, covering the diverse challenges that underwater communication systems may face. During the experiment, a variety of key parameter combinations were systematically tested, including three bandwidths BW (BW=1kHz, 2kHz, 4kHz) and three spreading factors SF (SF=4, 5, 6), forming a complete parameter matrix to explore the impact of parameter configuration on system performance.
[0076] The shallow pond experimental scenario was located in a science and technology park within a university in a certain province, with a water depth of approximately 1.2-1.8 meters and a communication distance set at 55 meters. This scenario exhibits typical characteristics of a still water environment, but also features abundant vegetation (including aquatic plants and duckweed) and significant biological noise interference (mainly from fish activity and insect calls), simulating the complex acoustic environment commonly found in real urban waterways. To ensure the reliability of the test data...
[0077] The slow-flowing stream experimental scenario was located near the library of a university in a certain province, with an average water flow velocity of approximately 0.3 m / s. The communication distance was extended to 80 meters to verify the system's performance over longer distances. This scenario was characterized by significant multipath effects, mainly stemming from the interaction between the water flow and the riverbed rocks, as well as the reflection of water surface ripples, accompanied by low-frequency water flow noise and occasional interference from fallen leaves. Notably, the stream bottom topography was complex and varied, including areas of sand, pebbles, and silt, providing natural conditions for investigating the impact of different bottom materials on signal propagation.
[0078] All experiments used the same hardware configuration, with the transmitter and receiver kept horizontally aligned and at a depth of 0.5 meters underwater to reduce multipath effects in the vertical direction.
[0079] II. Analysis of Experimental Results
[0080] 1) Performance comparison under different SF and BW configurations
[0081] To investigate the performance differences of the system under different SF and BW configurations and to find the optimal experimental configuration, a total of 176,000 Chirp symbols were collected for performance evaluation in the small pond environment shown in Figure 6(a), for various combinations of bandwidths (BW=1kHz, 2kHz, 4kHz) and spreading factors (SF=4, 5, 6) (a total of 5 effective configurations). The experimental results are shown in Figure 7.
[0082] Specifically, under a fixed bandwidth BW = 2kHz (Figure 7(a)), as the spreading factor SF increases from 4 to 6, the SER of both methods decreases, which is consistent with expectations—lower SF values correspond to shorter chirped symbols that are more susceptible to noise interference. When SF = 6, the SER of the traditional Decirp method is as high as 19.5%, while that of the method of this invention is only 1.7%, widening the performance gap to 17.8 percentage points. This indicates that deep learning models can more effectively mine fine-grained features of signals and suppress noise interference under high spreading factors.
[0083] Similarly, under the condition of fixed SF=5 (Figure 7b), as the bandwidth increases from 1kHz to 4kHz, both methods exhibit a trend of first decreasing and then increasing, reaching optimal performance at BW=2kHz. However, the performance of the method of this invention (SER=2.3%) is still significantly lower than that of the traditional Decirp method (SER=21.8%). This phenomenon occurs because, while keeping the spreading factor (SF) constant, increasing the bandwidth (BW) shortens the duration of the chirped symbol. This leads to a decrease in spreading gain, thereby weakening the system's resistance to environmental noise. Conversely, when the bandwidth is reduced below a critical threshold (e.g., 1kHz), although the symbol duration increases, the actual spectral spread is limited, resulting in excessive concentration of signal energy and making it more susceptible to masking effects from narrowband interference and spectral selective noise. This explains why the system performance is optimal at BW=2kHz—this configuration achieves the best balance between symbol duration and spectral spread, ensuring sufficient spreading gain while avoiding excessive concentration of spectral energy, thus achieving optimal noise robustness and demodulation reliability in complex underwater environments.
[0084] Based on the above experimental results, and in order to achieve the best balance between communication reliability and transmission efficiency, the combination of BW=2kHz and SF=5 was finally selected as the system default configuration for subsequent experimental evaluation.
[0085] 2) Performance variation trend under different signal-to-noise ratios (SNR)
[0086] To systematically evaluate the adaptability and robustness of the proposed communication framework under different noise environments, this invention conducted noise immunity tests in a controlled underwater experimental environment shown in Figure 6(a) using the determined optimal parameter configuration (BW=2kHz, SF=5).
[0087] The experiment first acquired a benchmark dataset under high signal-to-noise ratio (SNR) conditions: by adjusting the transmit power to a controllable maximum output power (approximately 0.6W) and continuously transmitting data packets. After rigorous screening and processing, 32,000 high-quality, interference-free Chirp symbol samples were obtained, forming a clean high SNR reference dataset. Based on this, software simulation was used to precisely inject additive white Gaussian noise (AWGN) of varying intensities into the benchmark dataset, simulating the SNR gradient changes from -35dB to 5dB. Reliable performance evaluation results were obtained through statistical analysis.
[0088] Figure 8 illustrates the symbol error rate (SER) performance of the two demodulation methods under different signal-to-noise ratio (SNR) conditions. As the SNR gradually decreases from 5 dB to -35 dB, the SER of both methods shows an increasing trend, reflecting the negative impact of increased noise intensity on communication quality. When the SNR drops to -35 dB, the SER of both methods approaches 100%, indicating that the system has completely lost its decoding capability. Notably, the SER curve of the deep learning demodulation method proposed in this invention remains below that of the traditional Decirp method throughout the entire test range, confirming its performance advantage under various noise conditions. Simultaneously, under the same SNR conditions, the proposed method reduces the symbol error rate by an average of 5.8% compared to the traditional Decirp method. More importantly, as the SNR decreases from 5 dB to -10 dB, the performance gap between the two methods initially widens gradually, reaching a peak at SNR = -10 dB (SER = 18.1%); when the SNR further decreases to -35 dB, the gap shows a converging trend. This phenomenon can be explained as follows: at moderate noise levels (-20dB to 0dB), deep learning models can effectively extract discriminative information from time-frequency features, significantly outperforming traditional methods that rely solely on coarse-grained peak detection; however, in the extremely low signal-to-noise ratio region (<-20dB), even deep learning models struggle to extract effective signal features from strong noise, leading to a convergence in the performance of the two methods.
[0089] 3) Performance comparison under different water bodies
[0090] In all the aquatic environments shown in Figure 6, with a fixed transmitter-receiver spacing and the device placed at a water depth of 40 cm, using a configuration of BW=2kHz and SF=5, a total of 64,000 measured Chirp symbol samples were collected to evaluate the symbol error rate (SER) performance in different water environments. The experimental results are shown in Figure 9.
[0091] In all test scenarios, the proposed method significantly outperformed the traditional Decirp demodulation method in terms of SER, with an average SER reduction of 13.5%, fully validating its robustness and adaptability in complex underwater channels. Specifically, in a slow-flowing river environment (Figure 6(b)), the proposed model achieved a SER of 2.31%, significantly better than the 3.10% in a shallow pond scenario (Figure 6(a)). This performance gap is mainly due to the extremely shallow water depth of the pond (approximately 0.5m), leading to strong multipath effects and bottom reflections, thus introducing more severe signal distortion and environmental noise. In contrast, rivers and lakes have greater water depth and more stable acoustic propagation conditions, which are conducive to the complete preservation of the Chirp signal and further highlight the advantages of the deep learning demodulation mechanism in suppressing interference in complex environments.
[0092] 4) Performance variation trend under different transmit powers
[0093] To further verify the power efficiency advantage of the proposed method in a real underwater environment, the symbol error rate (SER) performance of the system was tested at different transmit power levels. The experimental results are shown in Figure 10. The experiment was conducted in the controlled underwater environment shown in Figure 6(a), using the determined optimal parameter configuration (BW=2kHz, SF=5). By precisely controlling the output power of the transmitter, the system performance in the range of 0.1W to 0.6W was comprehensively evaluated.
[0094] As shown in Figure 10, the deep learning demodulation method proposed in this invention significantly outperforms the traditional Decirp benchmark method at all test power levels. Of particular note is that when the transmit power is 0.25W, the proposed method achieves a SER reduction of up to 39.6%, lowering the bit error rate from 44.7% of the traditional method to 5%. This performance difference has significant practical implications in the field of underwater communication. More importantly, to achieve the target SER threshold of 10%, the proposed method requires only 0.18W of transmit power, while the traditional Decirp method requires 0.53W, meaning the former consumes only about 54% of the transmit power of the latter. This result demonstrates that, while maintaining the same communication reliability, this method can significantly reduce transmitter energy consumption, which is of significant value for energy-constrained underwater sensing nodes.
Claims
1. A deep learning-based underwater chirp spread spectrum communication demodulation method, characterized in that, The process includes the following steps: Step 1: Signal preprocessing; receiving underwater Chirp spread spectrum communication signals, sequentially performing packet detection, time-frequency offset correction, and symbol conversion operations to obtain the input features of the deep learning model; Step 2: Demodulation based on deep learning; inputting the input features obtained in Step 1 into a preset deep learning model, using the deep learning model to extract and identify time-frequency features at multiple scales, outputting the symbol category probability distribution, and completing signal demodulation.
2. The underwater chirp spread spectrum communication demodulation method based on deep learning according to claim 1, characterized in that, Chirp signals are generated by superimposing a discrete frequency offset onto a base up-chirp signal. To achieve signal modulation, different The values correspond to different signs, and the modulated signal is represented as follows: (1) Among them, Based on the up-frequency chirped signal, This represents the frequency offset corresponding to the data symbol. As an offset component, the starting frequency of the fundamental up-frequency chirp is adjusted to... The sweep rate in the formula Due to bandwidth Duration of the Chirp signal Jointly determined; the spreading factor SF, as a core parameter, defines the size of the coded symbol set; among which, The signal bandwidth BW and the chirp duration T satisfy the following relationship: (2) In signal bandwidth Under fixed conditions, spreading factor The increase will cause the chirp symbol duration to increase. Corresponding extension; each Chirp symbol is characterized by a differentiated configuration of its initial frequency. A variety of distinct symbols.
3. The underwater chirp spread spectrum communication demodulation method based on deep learning according to claim 1, characterized in that, The packet structure in packet inspection includes a preamble, a synchronization word (SFD), and a data payload. The preamble consists of six basic up-chirps and is used to detect the presence of packets and perform coarse-grained signal synchronization. The synchronization word (SFD) consists of 2.25 basic down-chirps and is used for fine-grained signal synchronization and to eliminate carrier frequency offset (CFO). The data payload is used to transmit encoded valid data.
4. The underwater chirp spread spectrum communication demodulation method based on deep learning according to claim 1, characterized in that, The packet detection in step 1 specifically includes: the receiver dividing the received signal into sliding windows with the symbol duration T as the window length, and the signal in each window is called the "window signal"; performing dechirping and peak detection operations on each window signal; when the energy peaks of multiple window signals appear at the same frequency position, it indicates that a packet has been detected.
5. The underwater chirp spread spectrum communication demodulation method based on deep learning according to claim 1, characterized in that, The time-frequency offset correction in step 1 specifically includes: when the sliding window is aligned with a single Chirp symbol, after performing dechirping and peak detection on the basic up-frequency chirp in the preamble, its FFT peak should appear at the position of frequency 0; by detecting the deviation between the peak position and the zero-frequency position in the window signal, the time offset is calculated, and the signal is time-domain shifted to achieve alignment of the symbol boundary.
6. The underwater chirp spread spectrum communication demodulation method based on deep learning according to claim 1, characterized in that, The symbol transformation in step 1 specifically includes: after aligning the window signal with the Chirp symbol boundary, using the length of a single symbol as the length of the sliding window, dividing each Chirp symbol sequentially, performing a short-time Fourier transform (STFT), generating a two-dimensional time-frequency spectrum, and extracting its real part as the input feature of the neural network.
7. The underwater chirp spread spectrum communication demodulation method based on deep learning according to claim 1, characterized in that, Step 2 specifically includes: the deep learning model receives the real part of the time-frequency map generated in the preprocessing stage as input, and first performs frequency domain feature enhancement through the time-frequency transformation module FTB; then, the input signal is divided into multiple local token sequences after block segmentation and linear embedding processing, and enters a four-stage hierarchical processing flow; wherein, the first, third and fourth stages each contain two basic processing units, and the second stage is expanded to six repeating units to enhance the core feature extraction capability; each processing unit is composed of the FTB module and the SwinTransformer block in collaboration; the FTB module is responsible for enhancing local time-frequency characteristics, and the SwinTransformer uses a shift window mechanism to model long-distance dependencies and achieve global context awareness; a patch merging layer is introduced between each stage to gradually downsample the spatial resolution of the feature map, while doubling the number of channels to construct a pyramid-like structure.
8. A deep learning-based underwater chirp spread spectrum communication demodulation system, characterized in that, The system includes a transmitter, a receiver, and a signal processing module. The transmitter is used to generate and transmit Chirp spread spectrum communication signals and includes a main control unit, an audio power amplifier, and a vibration motor. The receiver is used to receive underwater acoustic Chirp signals and convert them into electrical signals and includes a piezoelectric ceramic sensor, an audio decoder, and a power supply device. The signal processing module is used to execute the deep learning-based underwater Chirp spread spectrum communication demodulation method according to any one of claims 1-7.
9. The underwater chirp spread spectrum communication demodulation system based on deep learning according to claim 8, characterized in that, The main control unit uses a Raspberry Pi to generate a Chirp spread spectrum modulation signal containing a preamble, a synchronization word SFD, and a data load. The audio power amplifier is used to amplify the power of the modulation signal, and the vibration motor is used to convert the electrical signal into an underwater acoustic Chirp signal.
10. The underwater chirp spread spectrum communication demodulation system based on deep learning according to claim 8, characterized in that, The piezoelectric ceramic sensor is used to pick up underwater vibration signals, the audio decoder is used to digitize the signals, and the power supply device provides power to the receiver.