Bionic signal demodulation method based on DSCAF-Net
Through the DSCAF-Net demodulation method, a dual-stream feature tensor is constructed and feature fusion is performed to optimize the training process, which solves the problems of high bit error rate and insufficient phase correction in the bionic covert communication system in complex underwater acoustic channels, and achieves low latency, high real-time and efficient signal demodulation.
Patent Information
- Application Number
- CN202510938055.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-14
AI Technical Summary
Existing bionic covert communication systems have a high bit error rate under complex underwater acoustic channel conditions and rely on high-performance hardware equipment. In addition, the deep learning model is highly complex and the input data processing is complicated, making it difficult to meet the requirements of low latency and high real-time performance, and the phase correction capability is insufficient.
A bionic signal demodulation method based on DSCAF-Net is adopted. A dual-stream feature tensor is constructed through signal preprocessing. A dual-stream neural network architecture is used for feature extraction and fusion. Training is combined with mini-batch gradient descent, label smoothing and Adam optimizer, and the learning rate is dynamically adjusted to optimize model performance.
The bionic signal demodulation with low bit error rate, low latency and high real-time performance is achieved under complex underwater acoustic channel conditions, which improves the generalization and robustness of the model and reduces dependence on hardware equipment.
Smart Images

Figure CN120785701A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bionic signal demodulation, and in particular to a bionic signal demodulation method based on DSCAF-Net. Background Art
[0002] In current biomimetic covert communication systems, traditional modulation and demodulation technologies exhibit high bit error rates in complex underwater acoustic channel conditions and severe signal attenuation, and are highly dependent on high-performance hardware, hindering further improvements in system communication quality. In recent years, the application of deep learning in modulation recognition and signal demodulation has gradually demonstrated its theoretical advantages. However, existing methods generally suffer from high model complexity and cumbersome input data processing, making it difficult to meet the low-latency and high-real-time requirements of practical systems. Furthermore, existing research has performed poorly in phase correction. Deep neural networks often rely on external hardware systems to provide phase compensation, lacking the ability to autonomously learn and dynamically correct phase mismatches.
[0003] For example, Korean scholar Jongmin Ahn et al. proposed a bionic communication method based on real dolphin whistles, which uses a DAG-Net and LSTM structure to achieve information recovery under the communication link, and uses link state information for auxiliary enhancement. However, this method has problems with long input sequences and complex network structures, which limit its real-time demodulation capabilities. For another example, in a study by He Dongxuan et al. from Beijing Institute of Technology, by introducing a hardware phase compensation module in front of the deep neural network demodulator and analyzing the difference in the maximum deflection ratio under terahertz channels and AWGN channels, the demodulation performance was improved under specific conditions. However, their method still essentially relies on external phase correction, and the neural network itself does not have end-to-end phase compensation capabilities. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to propose a bionic signal demodulation method based on DSCAF-Net, which can solve at least one technical problem mentioned in the background technology.
[0005] According to one aspect of the present invention, a bionic signal demodulation method based on DSCAF-Net is provided, comprising:
[0006] After receiving the bionic signal, the signal is preprocessed, a dual-stream feature tensor is constructed and sample-level random sorting is performed, global normalization is performed, and the category labels are one-hot encoded. The preprocessed dataset is divided into training, validation, and test sets.
[0007] A two-stream neural network architecture is constructed. The two-stream separation module is used to split the data into two independent channels: the training sequence and the data symbol sequence. The two-stream two-layer residual convolution module is used to extract features from the two channels respectively. The cross-attention feature fusion module completes the two-stream feature fusion and outputs the demodulation result through the pooling layer and classification layer.
[0008] The neural network is trained using mini-batch gradient descent, with a label smoothing method based on the number of categories and smoothing strength as the loss function. The network parameters are updated using the Adam optimizer, and the learning rate scheduler is used to dynamically adjust the learning rate.
[0009] The trained neural network model is verified, and the model weight with the lowest bit error rate on the verification set is selected as the optimal model, and its performance is compared and evaluated with the benchmark algorithm.
[0010] In the above technical solution, first, in the signal preprocessing stage, the biomimetic signal is preprocessed to construct a two-stream feature tensor and perform sample-level random sorting. Global normalization is performed and the category labels are one-hot encoded. The dataset is then divided into training, validation, and test sets. This process provides a high-quality, structured data foundation for subsequent neural network training, helping to improve the model's generalization and demodulation performance. Next, in the network architecture design, a two-stream neural network architecture is constructed. The two-stream separation module splits the data into two independent channels: a training sequence and a data symbol sequence. A two-stream two-layer residual convolution module is used to extract features from the two channels separately. The cross-attention feature fusion module then completes the two-stream feature fusion. Finally, the demodulation result is output through the pooling layer and classification layer. This two-stream architecture can fully utilize the information from different channels. The residual convolution module and the cross-attention mechanism help improve the effect of feature extraction and fusion, thereby improving the accuracy of demodulation. For model training, the neural network is trained using mini-batch gradient descent, with a label smoothing method based on the number of categories and smoothing strength as the loss function. The network parameters are updated using the Adam optimizer, while the learning rate scheduler dynamically adjusts the learning rate. This series of optimization strategies effectively improves model training efficiency and convergence speed, ensuring good performance in complex underwater acoustic channel conditions. Finally, during the model verification and evaluation phase, the model weight with the lowest bit error rate on the validation set is selected as the optimal model, and its performance is compared with the baseline algorithm. This process objectively verifies the performance advantages of this solution.
[0011] This application effectively solves the key problems in existing bionic covert communication systems through signal preprocessing, dual-stream neural network architecture design and optimized training strategies. It is expected to achieve bionic signal demodulation with low bit error rate, low latency and high real-time performance under complex underwater acoustic channel conditions, providing new ideas and technical support for the development of bionic covert communication systems.
[0012] In some embodiments, a dual-stream feature tensor is constructed, including: defining a training sequence as a known reference signal of length L for channel estimation and phase correction of data symbols, defining a data symbol as a signal unit that carries data and contains information to be decoded; the transmitting end adopts a periodic frame structure, each frame starts with a training sequence followed by N data symbols in series, and the training sequence in each frame is copied N times to generate a training sequence copy matrix; the data symbol sequence and the training sequence copy matrix are spatially aligned to construct a three-dimensional matrix, a single sample is a two-dimensional matrix, channel 0 is the training sequence corresponding to the data symbol sequence, and channel 1 is the data symbol sequence.
[0013] In the above technical solution, training sequences and data symbols are first defined. A training sequence is a known reference signal of length L, used for channel estimation and phase correction of data symbols, while a data symbol is a signal unit that carries the information to be decoded. At the transmitter, a periodic frame structure is adopted, with each frame beginning with a training sequence followed by N data symbols in series. This frame structure design ensures that the training sequence provides a reference for channel estimation and phase correction for each data symbol while also imbuing data transmission with a certain regularity, facilitating processing at the receiver. To construct a dual-stream feature tensor, the solution further replicates the training sequence in each frame N times to generate a training sequence replica matrix. The data symbol sequence and the training sequence replica matrix are then spatially aligned to form a three-dimensional matrix. In this three-dimensional matrix, a single sample is a two-dimensional matrix, with channel 0 representing the training sequence corresponding to the data symbol sequence, and channel 1 representing the data symbol sequence. This dual-stream feature tensor construction method not only preserves the original information of the training sequence and data symbols but also, through spatial alignment, strengthens the relationship between them, facilitating subsequent feature extraction and fusion.
[0014] The dual-stream feature tensor construction method provides a rich and easy-to-process data format for signal demodulation. This design not only helps improve the accuracy of channel estimation and phase correction, but also provides more effective input for subsequent neural network feature extraction and demodulation, thereby improving the performance of the entire bionic signal demodulation system.
[0015] In some embodiments, a two-stream two-layer residual convolution module is used to extract features from two channel data respectively, including:
[0016] The normalized three-dimensional tensor is split into two independent channels through the dual-stream separation module. The data of the two channels are respectively input into two layers of convolution for feature extraction. The residual connection is introduced to map the input data to the same shape as the output feature through convolution and add them as the input of the subsequent module.
[0017] In the technical solution, the double-flow double-layer residual convolution module is used to extract features from two-channel data respectively. By combining the characteristics of double-flow architecture and residual convolution, the gradient vanishing problem in deep network training is solved, and the performance and stability of the model are improved.
[0018] Specifically, the normalized three-dimensional tensor is first split into two independent channels by the double-flow separation module. This separation method allows the training sequence and data symbol sequence to be processed separately, avoiding mutual interference between them, while fully utilizing their respective characteristics. Subsequently, the data of the two channels are input into two convolution layers for feature extraction. Convolution operation can effectively capture the local features of the signal, and the design of double-layer convolution further enhances the depth and complexity of feature extraction, enabling the model to better understand the internal structure of the signal.
[0019] During feature extraction, residual connection is introduced. Residual connection solves the gradient vanishing problem in deep network training by mapping the input data through convolution to the same shape as the output features and adding it to the output of the convolution layer. This design not only preserves the original information of the input data, but also enables the network to learn more complex feature representations. Through residual connection, the model can more effectively pass gradients, speed up training, and improve the generalization ability of the model.
[0020] The use of a double-flow double-layer residual convolution module not only fully utilizes the information of the training sequence and data symbol sequence, but also solves the problem of deep network training through residual connection, improving the performance and stability of the model. This design provides high-quality feature representations for subsequent feature fusion and demodulation, helping to further improve the accuracy and reliability of the biomimetic signal demodulation system.
[0021] In some embodiments, the double-flow feature fusion is completed by a cross-attention feature fusion module, including: mapping the feature map of channel 0 to the query space, mapping the feature map of channel 1 to the key and value space, calculating the dot product similarity matrix S of Query and Key, performing Softmax normalization along the last dimension of the similarity matrix S to generate weights for each position, and performing weighted summation on the Value using the attention weights to complete the double-flow feature fusion.
[0022] In the above technical solution, the design of the cross-attention feature fusion module for double-flow feature fusion involves using attention mechanisms to dynamically capture and fuse the relevance between the features of the two channels, thereby improving the performance of signal demodulation.
[0023] Specifically, the module first maps the feature maps of channel 0 (training sequence) into the query space and the feature maps of channel 1 (data symbol sequence) into the key and value spaces. This mapping allows the features of the two channels to interact through the attention mechanism. By calculating the dot product similarity matrix S between the query and key, the model quantifies the correlation between the features of the two channels. The dot product operation captures the degree of match between the features, providing a basis for subsequent weight assignment. The similarity matrix S is then softmax normalized along the last dimension to generate weights for each position. Softmax normalization ensures that the weights sum to 1, allowing the weights to accurately reflect the relative importance of the features. Finally, these attention weights are used to perform a weighted summation of the values to complete the fusion of the two-stream features. This weighted summation method dynamically adjusts the contribution of features based on the weights, thereby effectively fusing the features of the two channels.
[0024] The cross-attention feature fusion module not only fully utilizes the information of the two-channel features, but also dynamically captures the correlation between the features through the attention mechanism. This dynamic adjustment method enables the model to better adapt to different signal conditions and channel environments, thereby improving the accuracy and robustness of signal demodulation. In addition, the introduction of the attention mechanism also provides interpretability for the model, allowing researchers to better understand the process and mechanism of feature fusion. The cross-attention feature fusion module not only improves the performance of signal demodulation, but also provides new ideas and methods for the design of bionic signal demodulation systems. This design is expected to achieve more efficient and accurate signal demodulation under complex underwater acoustic channel conditions, and promote the development of bionic covert communication technology.
[0025] In some embodiments, the label smoothing method based on the number of categories and the smoothing strength includes: customizing the smoothing label and replacing the original one-hot code label with the smoothing label, and changing the true category The probability of is reduced to 1-α, and the remaining K-1 categories share the probability of α, each probability is α is the smoothing intensity function, K is the total number of categories, and the loss function is the cross entropy loss function.
[0026] In the above technical solution, a label smoothing method based on the number of categories and smoothing strength is used as the design of the loss function, aiming to improve the generalization ability and robustness of the model.
[0027] Specifically, the core of this method is to smooth the original one-hot code labels. Traditional one-hot code labels set the probability of the true category to 1 and the probability of the other categories to 0. This rigid division may cause the model to be overconfident in the training data, resulting in insufficient generalization ability when facing noise or unseen data. The label smoothing method introduces a smoothing intensity function α to smooth the true category. The probability of is reduced to 1-α, and the remaining α probability is evenly distributed to the other K-1 categories. The probability of each category is This smoothing process makes the labels no longer absolute 0 and 1, but a probability distribution with a certain degree of ambiguity, thereby encouraging the model to learn smoother decision boundaries and reduce the risk of overfitting.
[0028] The smoothing strength α is a key parameter that determines the degree of label smoothing. Smaller values of α make labels closer to the original one-hot encoding, while larger values increase label ambiguity. By choosing α appropriately, a balance can be achieved between the model's confidence and generalization ability. Furthermore, the total number of categories K also affects the smoothed probability distribution. The probability distribution of smoothed labels takes the total number of categories into account, making the smoothing process more adaptable. Regarding the loss function, this scheme uses the cross-entropy loss function. The cross-entropy loss function is a commonly used loss function in classification tasks. It measures the difference between the probability distribution of the model output and the probability distribution of the true labels. By using the smoothed labels as the target probability distribution, the cross-entropy loss function effectively guides model training, ensuring that the model not only focuses on the true category during learning but also maintains a certain degree of sensitivity to other categories. This smoothed loss function encourages the model to output a smoother probability distribution, thereby improving the model's robustness in complex environments.
[0029] A label smoothing method based on the number of categories and smoothing strength is an effective optimization strategy. By smoothing label information, it mitigates the model's overconfidence in the training data and improves its generalization ability. Furthermore, combined with the cross-entropy loss function, this method effectively guides the model's training process, resulting in better demodulation performance under complex channel conditions. This design not only helps improve the accuracy and reliability of the bionic signal demodulation system but also provides new insights and methods for the application of deep learning in the communications field.
[0030] In some embodiments, the network parameters are updated by an Adam optimizer, including
[0031] The first-order momentum is calculated as the exponential moving average of the gradient direction at each moment, and the second-order momentum is calculated to reflect the gradient change over a period of time. The first-order momentum and the second-order momentum are corrected for deviations, and the gradient formula for updating the neural network weights by the Adam optimizer and the weight update formula of the neural network are obtained. With each training of the neural network, the weights are updated according to the above formula.
[0032] In the above technical solution, the Adam optimizer combines the advantages of the momentum method and adaptive learning rate adjustment, can effectively handle non-stationary targets and sparse gradients, and is suitable for training large-scale data sets and complex models.
[0033] Specifically, the Adam optimizer first calculates first-order momentum (i.e., the first-order moment estimate of the gradient), which is the exponential moving average of the gradient direction at each moment. This first-order momentum can smooth gradient fluctuations, helping the model better adapt to gradient changes. Simultaneously, the Adam optimizer also calculates second-order momentum (i.e., the second-order moment estimate of the gradient), which reflects gradient changes over time. Second-order momentum provides information about the gradient distribution, helping to adjust the learning rate so that the model can update at different rates across different parameter dimensions.
[0034] To eliminate initial gradient deviations, the first-order and second-order momentum are corrected. The corrected momentum values more accurately reflect the true gradient, thereby improving the stability of the optimization process. Based on the corrected momentum values, the Adam optimizer derives the gradient formula for updating the neural network weights and the weight update formula for the neural network. These formulas combine gradient smoothing and adaptive learning rate adjustment, enabling the model to more efficiently converge to a global or local optimal solution during training. With each training session, the neural network updates its weights according to the above formulas. This dynamic weight adjustment allows the model to continuously adapt to data changes during training while avoiding local optimality or slow convergence. The Adam optimizer's adaptive learning rate adjustment mechanism allows the model to flexibly adjust the learning rate based on different parameter dimensions and training stages, thereby improving training efficiency and model performance.
[0035] The Adam optimizer updates network parameters, taking into account gradient smoothing and adaptive learning rate adjustment, providing efficient and stable support for model training. This optimizer not only helps improve the model's demodulation performance in complex communication environments, but also provides a universal and effective optimization strategy for training deep learning models.
[0036] In some embodiments, a learning rate scheduler is used to dynamically adjust the learning rate, including:
[0037] The ReduceLROnPlateau learning rate scheduler is used to monitor the changes in the validation set loss. If the indicator does not improve within 20 consecutive training cycles, the current learning rate is reduced by a preset factor. The learning rate update rule is:
[0038] η new =max(factor·η old , min_lr)
[0039] Where factor is the learning rate attenuation coefficient, min_lr is the lower limit of the learning rate, and five-fold cross-validation is used in the model training stage. The model weight file is saved every time the preset condition is met. The bit error rate on the validation set and test set is monitored in real time during the training process.
[0040] In the above technical solution, the ReduceLROnPlateau learning rate scheduler is used to dynamically adjust the learning rate, which can further optimize the model training process and ensure that the model has better convergence and robustness in complex environments.
[0041] The learning rate is a critical hyperparameter in deep learning model training. Excessively high learning rates can lead to instability or even failure to converge, while excessively low learning rates can slow training and increase training time. Dynamic learning rate adjustment automatically adjusts the learning rate based on performance metrics during training (such as validation set loss), enabling rapid convergence early in training and fine-tuning in later stages to avoid getting stuck in local optima.
[0042] The above solution uses the ReduceLROnPlateau learning rate scheduler, whose core idea is to dynamically adjust the learning rate by monitoring changes in validation set loss. Specifically, when validation set loss does not improve for 20 consecutive training cycles, the learning rate scheduler reduces the current learning rate by a preset decay factor. This mechanism effectively prevents the model from prematurely converging to a local optimum during training and promptly adjusts the learning rate to seek a better solution when validation set performance stagnates.
[0043] The learning rate update rule is: η new =max(factor·η old , min_lr), where factor is the learning rate decay coefficient, which controls the reduction rate; min_lr is the lower limit of the learning rate, preventing the learning rate from being too low and causing the training process to stagnate. In this way, the learning rate scheduler can ensure model convergence while avoiding the problems caused by too low a learning rate.
[0044] During model training, five-fold cross-validation is a very effective validation method. By splitting the dataset into five parts, using one part as the validation set and the remaining four parts as the training set, limited data resources can be fully utilized while reducing bias caused by data partitioning. During each validation process, model performance metrics (such as validation set loss and bit error rate) more accurately reflect the model's actual performance. Furthermore, this solution saves the model weight file every time a preset condition (such as validation set loss improvement or learning rate adjustment) is met during training. This mechanism ensures that training results are not lost due to unexpected interruptions during training and facilitates the selection of optimal model weights for evaluation and deployment after training. During training, the bit error rate on the validation and test sets is monitored in real time. The bit error rate is a key indicator of demodulation performance in communication systems. By monitoring the bit error rate in real time, it is possible to promptly identify changes in model performance on different datasets and adjust the training strategy accordingly. For example, if the validation set bit error rate continues to decrease, but the test set bit error rate does not significantly improve, further adjustments to the model structure or training parameters may be necessary to improve the model's generalization ability.
[0045] Using the ReduceLROnPlateau learning rate scheduler to dynamically adjust the learning rate, combined with five-fold cross-validation and real-time monitoring, not only effectively improves the model's convergence speed and stability, but also prevents the model from falling into local optima by dynamically adjusting the learning rate. Furthermore, real-time monitoring of the bit error rate on the validation and test sets ensures the model's performance in real-world applications. This comprehensive optimization strategy provides strong support for efficient training and performance improvement of the bionic signal demodulation system.
[0046] According to another aspect of the present invention, a bionic signal demodulation device based on DSCAF-Net is provided, which is based on the above method and includes the following modules connected in sequence:
[0047] The receiving module is used to perform signal preprocessing after receiving the bionic signal, construct a dual-stream feature tensor and perform sample-level random sorting, perform global normalization and one-hot encoding on the category labels, and divide the preprocessed dataset into training, validation, and test sets;
[0048] The construction module is used to build a two-stream neural network architecture. The two-stream separation module splits the data into two independent channels: the training sequence and the data symbol sequence. The two-stream two-layer residual convolution module is used to extract features from the two channels respectively. The cross-attention feature fusion module completes the two-stream feature fusion and outputs the demodulation result through the pooling layer and classification layer.
[0049] The training module is used to train the neural network using mini-batch gradient descent, using a label smoothing method based on the number of categories and smoothing strength as the loss function, updating the network parameters through the Adam optimizer, and dynamically adjusting the learning rate using the learning rate scheduler;
[0050] The evaluation module is used to verify the trained neural network model, select the model weight with the lowest bit error rate on the verification set as the optimal model, and compare its performance with the benchmark algorithm.
[0051] In the above technical solution, in order to better use the above method, this application proposes a bionic signal demodulation device based on DSCAF-Net, and each module corresponds to each step of the above method. Its specific principles have been described above and will not be repeated here.
[0052] According to another aspect of the present invention, there is provided a bionic signal demodulation device based on DSCAF-Net, comprising:
[0053] at least one processor and a memory communicatively coupled to the at least one processor;
[0054] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the above method.
[0055] In the above technical solution, in order to better run and process the method, the above method is stored in a memory and a processor is used to execute the stored method. It should be noted that the principle and effect of each step have been described above and will not be further explained here.
[0056] According to another aspect of the present invention, a computer-readable storage medium is provided, storing a computer program, wherein the computer program implements the above method when executed by a processor.
[0057] In the above technical solution, in order to better run and use the method, the above method is stored in a computer-readable storage medium and implemented by a processor. It should be noted that the principle and effect of each step have been described above and will not be further explained here. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0059] Figure 1 1 is a flow chart of an embodiment of a bionic signal demodulation method based on DSCAF-Net according to the present invention;
[0060] Figure 2 2 is a data set diagram of an embodiment of a bionic signal demodulation method based on DSCAF-Net according to the present invention;
[0061] Figure 3 This is a network model diagram of DSCAF-Net (Dual-Stream Cross-Attention Fusion Network) in an embodiment of the bionic signal demodulation method based on DSCAF-Net of the present invention.
[0062] Figure 4 It is a structural diagram of an embodiment of a bionic signal demodulation device based on DSCAF-Net of the present invention. DETAILED DESCRIPTION
[0063] The present invention will be described in further detail below with reference to the accompanying drawings and examples. It is particularly noted that the following examples are intended only to illustrate the present invention and are not intended to limit the scope of the present invention. Similarly, the following examples are only some embodiments of the present invention and are not intended to be exhaustive. All other embodiments obtained by those of ordinary skill in the art without creative effort are intended to fall within the scope of protection of the present invention.
[0064] Example 1
[0065] This embodiment proposes a bionic communication demodulation method based on deep learning, and constructs a corresponding bionic covert communication system to verify the effectiveness of the method. This method targets a bionic communication system based on octal phase rotation modulation, and replaces the demodulation module in the original system with a deep learning model, which significantly reduces the bit error rate. In this bionic communication system, information is transmitted by regulating the phase of the dolphin click signal (click). There are 8 phase modes in total, and each mode carries 3 bits of information. At the same time, a training sequence is inserted before every N data symbols for channel equalization and model training. This embodiment first uses the communication signal collected in the pool experiment to train the deep learning model, and then uses the communication signal collected in the sea trial experiment to verify the robustness of the model. As Figure 1 As shown, the whole process includes the following steps:
[0066] S1. After receiving the biomimetic signal, perform signal preprocessing, construct a dual-stream feature tensor and perform sample-level random sorting, perform global normalization and one-hot encoding on the category labels, and divide the preprocessed dataset into training, validation, and test sets; see Figure 2 , the S1 process includes the following steps:
[0067] S11. Construct a dual-stream network structure input sample. This application proposes a dual-stream network structure input method to meet the channel equalization and signal decoding requirements in the communication system. The specific implementation process is as follows:
[0068] Two signal components are defined: the training sequence is a known reference signal of length L, used for channel estimation and phase correction of data symbols. The data symbol is a signal unit that carries data, with a length of L and contains the information to be decoded.
[0069] Furthermore, the transmitter adopts an "N+1" periodic frame structure: each frame starts with a training sequence for establishing a channel response benchmark; N data symbols are subsequently concatenated to form a complete data segment with a total length of (N+1)×L.
[0070] Furthermore, to achieve dynamic channel equalization during decoding, the two-stream network input is generated according to the following rules: the training sequence in each frame is replicated N times to generate an N × L matrix of training sequence replicas. The N × L data symbol sequence is spatially aligned with the training sequence replica matrix to construct an N × 2 × L three-dimensional matrix. A single sample is a 2 × L two-dimensional matrix, where channel 0 is the training sequence corresponding to the data symbol sequence, and channel 1 is the data symbol sequence.
[0071] S12, global normalization processing, the normalization operation formula is:
[0072]
[0073] Where X is a three-dimensional tensor of M×2×L, M is the number of total data symbol sequences in the received signal. μ is the global mean, σ is the global standard deviation, X norm It represents the result of normalization of each channel signal, that is, converting the data of each channel into a normal distribution with zero mean and unit variance. The calculation formulas of μ and σ are as follows:
[0074]
[0075] Among them, ∈ is a very small positive number.
[0076] S13: Label one-hot encoding: To ensure that the neural network can understand and learn the classification task more clearly and accurately, the binary information carried by each code element is converted into a one-hot encoding as the label corresponding to the sample. Because the bionic signal modulation method uses an octal phase rotation method, an eight-bit one-hot encoding method is used. For example, the binary sequence
[101] will be mapped to [00001000].
[0077] S14. Dataset partitioning: The dataset is divided into a training set (70%), a validation set (20%), and a test set (10%). The input data dimension is a three-dimensional tensor of M×2×L, where M is the total number of codeword signals in the received signal, 2 is the number of channels, and L is the sequence length of each channel.
[0078] S2. Build a two-stream neural network architecture. Use the two-stream separation module to split the data into two independent channels: the training sequence and the data symbol sequence. Use the two-stream two-layer residual convolution module to extract features from the two channels respectively. Use the cross-attention feature fusion module to complete the two-stream feature fusion. After the pooling layer and the classification layer, the demodulation result is output. Please refer to Figure 3 , the steps of building a neural network include:
[0079] S21. Dimension reconstruction: The normalized data is converted into a three-dimensional tensor with a dimension of B×2×L, where B is the sample size of the batch input model. The data is first passed through the dual-stream separation module to split it into two groups of three-dimensional tensors of B×1×L, which are the training sequence and the data symbol sequence respectively.
[0080] S22. Dual-Stream Two-Layer Residual Convolution Module (DSTL-ResConv) module design: First, the normalized data is transformed into a three-dimensional tensor with a dimension of B×2×L, where B is the sample size of the batch input model. It is first passed through the dual-stream separation module to split the data into two independent channels (Dual-Stream), namely the training sequence and the data symbol sequence.
[0081] The data of the two channels are input into two layers of convolution for feature extraction. The feature dimension of each convolution layer output is C in , set the convolution kernel width to 3, the feature map padding to 0, the convolution kernel sliding step to 1, to extract the feature map of the input data, and obtain a dimension of B×C in ×L feature map, the two-layer convolution operation formula is as follows:
[0082]
[0083] in represents convolution operation, + represents matrix addition, is the weight of the first and second layer convolution kernels, is the bias term of the first and second layers.
[0084] The LeakyReLU activation function is defined as follows:
[0085]
[0086] Where α is the slope coefficient of the negative semi-axis, which is usually set to 0.01.
[0087] In order to improve the feature transfer capability and solve the gradient vanishing problem, the DSTL-ResConv module introduces a residual connection to convert the input X in Mapped to the same shape as X2 through a 1×1 convolution (to match the number of channels) and added to X2. The final output is:
[0088]
[0089] Among them, W res and B res are the convolution weights and biases of the residual connection, the convolution kernel size is 1×1, and the stride is 1. The final output size of the module is B×C in ×L, as the input of subsequent modules.
[0090] S23, Cross Attention Fusion Block (CAF Block) design: This module uses the cross attention mechanism to complete the feature fusion of the dual-stream network structure, and extracts the B×C dimension of the dual-stream channels. in ×L feature maps are input to the CAF Block for feature fusion. The feature map F1 of channel 1 is mapped to the query space, and the feature map F2 of channel 2 is mapped to the key and value space. The formula is:
[0091]
[0092] in, is the learnable weight matrix, C out is the feature dimension output by the cross attention mechanism. Further, the dot product similarity matrix between Query and Key is calculated:
[0093]
[0094] Furthermore, a weight is generated for each position to reflect its attention intensity to other positions. Softmax normalization is performed on the similarity matrix S along the last dimension (sequence direction):
[0095]
[0096] Among them, the formula of Softmax is:
[0097]
[0098] in, It represents the similarity between the i-th position of channel 1 and the j-th position of channel 2 in the b-th sample.
[0099] Furthermore, the attention weight W is used to perform a weighted summation on the Value, where the output of each position is a weighted combination of the Value, the weight is determined by W, and A is the result of the weighted summation of the Value feature by the attention weight matrix W (reflecting the attention strength between different positions). The formula is:
[0100]
[0101] Furthermore, the output dimension is B×C out ×L, in order to match the input dimension of the pooling layer, the dimension must first be converted to B×L×C out Then output the transpose A′ of A:
[0102]
[0103] S24, pooling layer design; input A′ to the average pooling layer, for each sample b and feature dimension C out , calculate the mean along the sequence length dimension (dim=1) and compress the sequence of each sample into a global feature:
[0104]
[0105] The output dimension of the average pooling layer is B×C out .
[0106] S25, classification layer design: The classification layer consists of two linear layers and a LeakyReLu activation layer. After the first linear layer, the dimension is increased from B×C out Reduced to B×C linear1 ; After the LeakyReLu activation layer, the dimension remains unchanged. The formula of LeakyReLu is:
[0107]
[0108] Among them, α is the slope of the negative half axis, which controls the leakage degree of the negative area and is set to 0.01. Then the second linear layer is used to compress the dimension to B×C linear2 As the output, it is compared with the one-hot encoded label, C linear2 It can be selected according to the modulation and demodulation mode. The experiment uses the octal phase rotation mode, then C linear2 =8.
[0109] Based on the above, in this embodiment, the DSCAF-Net neural network structure is as follows:
[0110] Table 2DSCAF-Net network structure
[0111]
[0112] S3. Use mini-batch gradient descent to train the neural network. Use a label smoothing method based on the number of categories and smoothing strength as the loss function. Update the network parameters through the Adam optimizer. Use the learning rate scheduler to dynamically adjust the learning rate. The neural network model training steps include:
[0113] S31. Use mini-batch gradient descent to divide the training set into blocks of uniform size and input them in batches. The neural network learns one batch of data per iteration.
[0114] S32. This application designs a label smoothing method based on the number of categories and smoothing strength as the loss function of the neural network. The neural network learns in the direction of minimizing the loss function. First, customize the smooth label
[0115]
[0116] Replace the original one-hot code label with a smooth label and change the true category The probability of is reduced to 1-α, and the remaining K-1 categories share the probability of α, each probability is α is a smooth intensity function (α∈[0,1]), K is the total number of categories, and the experiment uses octal phase rotation modulation, so K=8. The final loss function is:
[0117]
[0118] Where N is the number of neurons in the output layer, is the predicted value output by the i-th neuron in the output layer, and Y is the true value corresponding to the i-th neuron in the output layer. The smaller the value of the loss function, the better the neural network learns the data. Especially when facing category imbalance or label noise, it helps to improve the robustness and generalization ability of the model on unseen samples.
[0119] S33. Using the Adam optimizer to update neural network parameters: During training, the neural network calculates the gradient of the error backpropagation based on the loss function and updates the network weight parameters based on this gradient. The Adam optimizer adjusts the neural network gradient and uses the adjusted gradient to update the neural network weights. The Adam optimizer combines first-order momentum and second-order momentum algorithms to correct bias and adjust the neural network gradient.
[0120] The first-order momentum formula is: m t =β1*m t-1 +(1-β1)*g t
[0121] Among them, g t is the gradient calculated by the neural network at time t, m t is the first-order momentum at time t. The first-order momentum is the exponential moving average of the gradient direction at each moment, which is approximately equal to the average of the sum of the gradient vectors at the most recent (1-β1) moments;
[0122] The second-order momentum formula is: Among them, g t is the gradient calculated by the neural network at time t, V t is the second-order momentum at time t, which reflects the gradient change over a period of time.
[0123] The value of β1 is 0.9, and the value of β2 is 0.999. The initial values of m0 and V0 are both 0. Therefore, in the initial stage of neural network training, m t , V t The value of will be close to 0. Based on this, Adam optimizer t , V t Correct the deviation to solve this problem. The formula for correcting the deviation is as follows:
[0124]
[0125] in, and is the modified first-order momentum and second-order momentum, from which the gradient formula for the Adam optimizer to update the neural network weights is derived.
[0126]
[0127] The weight update formula of the neural network is:
[0128]
[0129] Among them, w t is the neural network weight at time t, and lr is the neural network learning rate. A suitable learning rate can make the neural network converge faster. With each training, the neural network updates the weight according to the above formula to achieve the effect of learning data and accurately identifying.
[0130] S34. This application uses the learning rate scheduler ReduceLROnPlateau provided by the PyTorch deep learning framework. This method monitors the changes in the validation set loss. If the indicator does not improve within 20 consecutive training cycles, the current learning rate is reduced by a preset factor to escape the local optimum in the early stages of training. The ReduceLROnPlateau learning rate scheduler is used to adjust the learning rate. The specific learning rate update rule is:
[0131] η new =max(factor·n old , min_lr)
[0132] factor is the learning rate attenuation coefficient (such as 0.5), and min_lr is the lower limit of the learning rate (such as 1e-30).
[0133] S35. To calculate the bit error rate during training and validation, take the argmax (argument of the maximum) of the one-hot encoding of the labels and convert it to binary:
[0134] pred_bits=dec2bin(argmax(P,dim=1))∈{0,1} B×3
[0135] dec2bin is a built-in function in Python that converts decimal numbers to binary numbers. argmax is used to return the index position corresponding to the maximum value of a tensor in a specified dimension. The argmax operation can be used to convert the vector into the final predicted category. The prediction is compared with the actual value bit by bit to calculate the bit error rate (BER):
[0136]
[0137] S36. During the model training phase, to improve the robustness and generalization of the model, a five-fold cross-validation method is used, with each fold trained for 150 epochs. The model weight file (.pth format) is saved every 10 epochs. In addition, the bit error rate on the validation and test sets is monitored in real time during training, and the current best model is automatically saved each time the validation set loss reaches the lowest value. This strategy is intended to ensure that the final model is fully trained and performs optimally on the validation and test sets, thereby avoiding overfitting and improving the model's practical application effect.
[0138] S4. Verify the trained neural network model, select the model weight with the lowest bit error rate on the validation set as the optimal model, and compare its performance with the benchmark algorithm. The neural network model verification steps include:
[0139] S41. To verify the model's performance on unseen samples, all weight files (.pth format) saved during training were used in the model testing phase. To evaluate the performance of each model in a real-world communication environment, this paper uses bit error rate as the primary performance metric. The model weight with the lowest bit error rate on the validation set is selected as the final optimal model, and its demodulation performance is systematically compared with that of an existing benchmark algorithm. The benchmark algorithm estimates and corrects signal phase offset based on the Hilbert transform. The specific process is as follows: First, a Hilbert transform is performed on the received signal to extract its analytical envelope. Then, the vector angle between the training sequence after passing through the channel and the original training sequence is calculated to estimate the phase offset angle introduced by the channel. Finally, this offset angle is used as a rotation compensation parameter and uniformly applied to the data symbol sequence at the receiving end, completing the phase correction process. This method can achieve relatively accurate phase compensation under ideal conditions, but it is highly dependent on the training sequence in complex underwater acoustic channels, resulting in limited robustness.
[0140] This evaluation strategy can not only effectively reflect the robustness of the model under complex channel conditions, but also verify the advantages and applicability of the proposed method in actual application scenarios from the perspective of generalization ability. In order to illustrate the recognition performance proposed by the present invention, this example uses the DualChannel-SE-Net method to compare the transmission bit error rate of the existing multi-level phase rotation benchmark algorithm under different environments. In this study, two sets of real sea trial data are used to evaluate the performance of the model. The structure of sea trial data 1 is: a training sequence is inserted for every 50 data symbol sequences for phase offset estimation and correction; and sea trial data 2 inserts a training sequence after every 125 data symbol sequences. Both sets of data were collected from the same sea area, but the acquisition time periods were different, thus covering the potential impact of environmental changes on communication performance, and have certain representativeness and practicality. Different data structure designs are intended to examine the robustness and adaptability of the model under conditions of changing training sequence density. The results are shown in Table 1:
[0141] Table 1 Bit error rate performance of different algorithms in different environments
[0142]
[0143] As shown in Table 2, when using deep learning for demodulation, the bit error rates of the three data types are significantly lower than those obtained using existing equalization algorithms, demonstrating that deep learning can effectively improve communication performance in this application scenario. This method significantly improves the information transmission reliability of the bionic covert communication system, demonstrating superior performance.
[0144] Based on one of the embodiments, the present application proposes a new low-complexity dual-stream neural network architecture and integrates a cross-attention mechanism to achieve channel equalization and adaptive phase correction. Specifically, the network extracts feature information of the data signal and the training sequence through parallel convolution channels, and then realizes dual-modal feature fusion through a cross-attention mechanism, thereby enhancing the network's robustness to phase drift and improving demodulation accuracy. In addition, to further improve the generalization performance of the network, the present application designs an improved label smoothing cross entropy loss function. In multi-classification tasks, the loss function converts discrete one-hot encoded target labels into soft distributions by introducing a smoothing factor, alleviating the model's overfitting of noisy labels and enhancing the model's robustness in complex environments. Experimental results show that this method achieved a bit error rate of 0.19% in sea trials under 2dB signal-to-noise ratio conditions, which has significant performance advantages over traditional demodulators based on equalization algorithms. More importantly, this method has a lightweight structure and good edge deployment capabilities, which can achieve efficient real-time demodulation in underwater acoustic communication scenarios.
[0145] Example 2
[0146] See also Figure 4 A bionic signal demodulation device based on DSCAF-Net, based on the above method, includes the following modules connected in sequence:
[0147] The receiving module is used to perform signal preprocessing after receiving the bionic signal, construct a dual-stream feature tensor and perform sample-level random sorting, perform global normalization and one-hot encoding on the category labels, and divide the preprocessed dataset into training, validation, and test sets;
[0148] The construction module is used to build a two-stream neural network architecture. The two-stream separation module splits the data into two independent channels: the training sequence and the data symbol sequence. The two-stream two-layer residual convolution module is used to extract features from the two channels respectively. The cross-attention feature fusion module completes the two-stream feature fusion and outputs the demodulation result through the pooling layer and classification layer.
[0149] The training module is used to train the neural network using mini-batch gradient descent, using a label smoothing method based on the number of categories and smoothing strength as the loss function, updating the network parameters through the Adam optimizer, and dynamically adjusting the learning rate using the learning rate scheduler;
[0150] The evaluation module is used to verify the trained neural network model, select the model weight with the lowest bit error rate on the verification set as the optimal model, and compare its performance with the benchmark algorithm.
[0151] In the above technical solution, in order to better use the method described in one of the embodiments, the present application proposes a bionic signal demodulation device based on DSCAF-Net, each module corresponding to each step of the above method, and its specific principles have been described above and will not be repeated here.
[0152] Embodiment 3
[0153] A bionic signal demodulation device based on DSCAF-Net, comprising: at least one processor and a memory communicatively connected to the at least one processor;
[0154] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the above method.
[0155] In the above technical solution, in order to better run and process the method described in one of the embodiments, the above method is stored in a memory and executed by a processor. It should be noted that the principle and effect of each step have been described above and will not be further explained here.
[0156] Example 4
[0157] A computer-readable storage medium stores a computer program, which implements the above method when executed by a processor.
[0158] In the above technical solution, in order to better run and use the method described in one of the embodiments, the above method is stored in a computer-readable storage medium and implemented using a processor. It should be noted that the principles and effects of each step have been described above and will not be further explained here.
[0159] The above descriptions are only some embodiments of the present invention and do not limit the scope of protection of the present invention. Any equivalent device or equivalent process transformation made by using the contents of the description and drawings of the present invention, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A bionic signal demodulation method based on DSCAF-Net, characterized in that: The method comprises: After receiving the bionic signal, the signal is preprocessed, a dual-stream feature tensor is constructed and sample-level random sorting is performed, global normalization is performed, and the category labels are one-hot encoded. The preprocessed dataset is divided into training, validation, and test sets. A two-stream neural network architecture is constructed. The two-stream separation module is used to split the data into two independent channels: the training sequence and the data symbol sequence. The two-stream two-layer residual convolution module is used to extract features from the two channels respectively. The cross-attention feature fusion module completes the two-stream feature fusion and outputs the demodulation result through the pooling layer and classification layer. The neural network is trained using mini-batch gradient descent, with a label smoothing method based on the number of categories and smoothing strength as the loss function. The network parameters are updated using the Adam optimizer, and the learning rate scheduler is used to dynamically adjust the learning rate. The trained neural network model is verified, and the model weight with the lowest bit error rate on the verification set is selected as the optimal model, and its performance is compared and evaluated with the benchmark algorithm.
2. A bionic signal demodulation method based on DSCAF-Net according to claim 1, characterized in that: Constructing a dual-stream feature tensor includes: defining a training sequence as a known reference signal of length L for channel estimation and phase correction of data symbols, defining a data symbol as a signal unit that carries data and contains information to be decoded; the transmitting end adopts a periodic frame structure, each frame starts with a training sequence followed by N data symbols in series, and the training sequence in each frame is copied N times to generate a training sequence copy matrix; the data symbol sequence and the training sequence copy matrix are spatially aligned to construct a three-dimensional matrix, a single sample is a two-dimensional matrix, channel 0 is the training sequence corresponding to the data symbol sequence, and channel 1 is the data symbol sequence.
3. A bionic signal demodulation method based on DSCAF-Net according to claim 1, characterized in that: A two-stream two-layer residual convolution module is used to extract features from the two channel data separately, including: The normalized three-dimensional tensor is split into two independent channels through the dual-stream separation module. The data of the two channels are respectively input into two layers of convolution for feature extraction. The residual connection is introduced to map the input data to the same shape as the output feature through convolution and add them as the input of the subsequent module.
4. A bionic signal demodulation method based on DSCAF-Net according to claim 1, characterized in that: The dual-stream feature fusion is completed through the cross-attention feature fusion module, including: Map the feature map of channel 0 to the query space, and map the feature map of channel 1 to the key and value space. Calculate the dot product similarity matrix S between the query and the key. Perform Softmax normalization on the similarity matrix S along the last dimension to generate weights for each position. Use the attention weight to perform weighted summation on the value to complete the dual-stream feature fusion.
5. The bionic signal demodulation method based on DSCAF-Net according to claim 1, characterized in that: Label smoothing methods based on the number of categories and smoothing strength include: Customize the smooth label and replace the original one-hot code label with the smooth label, and change the true category The probability of is reduced to 1-α, and the remaining K-1 categories share the probability of α, each probability is α is the smoothing intensity function, K is the total number of categories, and the loss function is the cross entropy loss function.
6. The bionic signal demodulation method based on DSCAF-Net according to claim 1, characterized in that: Update the network parameters through the Adam optimizer, including The first-order momentum is calculated as the exponential moving average of the gradient direction at each moment, and the second-order momentum is calculated to reflect the gradient change over a period of time. The first-order momentum and the second-order momentum are corrected for deviations, and the gradient formula for updating the neural network weights by the Adam optimizer and the weight update formula of the neural network are obtained. With each training of the neural network, the weights are updated according to the above formula.
7. The bionic signal demodulation method based on DSCAF-Net according to claim 1, characterized in that: Use the learning rate scheduler to dynamically adjust the learning rate, including: The ReduceLROnPlateau learning rate scheduler is used to monitor the changes in the validation set loss. If the indicator does not improve within 20 consecutive training cycles, the current learning rate is reduced by a preset factor. The learning rate update rule is: or new =max(factor·η old ,min_lr) Where factor is the learning rate attenuation coefficient, min_lr is the lower limit of the learning rate, and five-fold cross-validation is used in the model training stage. The model weight file is saved every time the preset condition is met. The bit error rate on the validation set and test set is monitored in real time during the training process.
8. A bionic signal demodulation device based on DSCAF-Net, characterized in that: The method according to any one of claims 1 to 7, comprising the following modules connected in sequence: The receiving module is used to perform signal preprocessing after receiving the bionic signal, construct a dual-stream feature tensor and perform sample-level random sorting, perform global normalization and one-hot encoding on the category labels, and divide the preprocessed dataset into training, validation, and test sets; The construction module is used to build a two-stream neural network architecture. The two-stream separation module splits the data into two independent channels: the training sequence and the data symbol sequence. The two-stream two-layer residual convolution module is used to extract features from the two channels respectively. The cross-attention feature fusion module completes the two-stream feature fusion and outputs the demodulation result through the pooling layer and classification layer. The training module is used to train the neural network using mini-batch gradient descent, using a label smoothing method based on the number of categories and smoothing strength as the loss function, updating the network parameters through the Adam optimizer, and dynamically adjusting the learning rate using the learning rate scheduler; The evaluation module is used to verify the trained neural network model, select the model weight with the lowest bit error rate on the verification set as the optimal model, and compare its performance with the benchmark algorithm.
9. A bionic signal demodulation device based on DSCAF-Net, characterized in that: include: at least one processor and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Deep learning communication system optimization method based on channel perception
CN121603136A