Lightweight Yellow River underwater signal classification method and system based on knowledge distillation

Through a lightweight underwater signal classification method based on knowledge distillation, the ResNet and GhostNet networks combined with self-attention and gradual distillation loss are used to solve the robustness and computational cost of underwater signal classification in the Yellow River Basin, and efficient and low-cost signal classification is achieved.

CN120277481APending Publication Date: 2025-07-08HENAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510313318.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Traditional underwater signal classification methods are not robust enough in the Yellow River Basin dam environment, making it difficult to adapt to multi-signal interference scenarios, and the high computational cost of convolutional neural networks limits its application on embedded devices.

Method used

The lightweight Yellow River underwater signal classification method based on knowledge distillation is adopted, and through ResNet as the teacher network, the sample relationship progressive knowledge distillation loss is used to guide GhostNet student network training, combining self-attention and progressive distillation loss, reducing the calculation amount and maintaining high accuracy.

Benefits of technology

While achieving high-precision signal classification in the Yellow River Basin, it significantly reduces the calculation cost and parameter quantity, improves the efficiency and applicability of signal classification, and is suitable for environments with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277481A_ABST
    Figure CN120277481A_ABST
Patent Text Reader

Abstract

The invention provides a lightweight Yellow River underwater signal classification method and system based on knowledge distillation, and relates to the technical field of Yellow River basin dam safety monitoring. The method comprises the following steps: acquiring original data of an underwater acoustic signal of the Yellow River at a signal acquisition end, and performing normalization processing; at the preprocessing end, a training set, a verification set and a test set are segmented through a sliding window; at a training end, ResNet is used as a teacher network for pre-training, and knowledge is migrated to a lightweight GhostNet student network in combination with SRPKD (Sample Relationship Progressive Knowledge Distillation) loss; and at the classification end, signal features are extracted through a GhostModule unit, the calculation amount is compressed, and finally real-time classification is realized. According to the lightweight signal classification method, the challenge of implementing deep learning on a resource-constrained system is successfully handled, the calculation cost can be effectively reduced, the excellent signal classification performance can be kept, and the communication requirement of the underwater complex environment of the Yellow River basin is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Yellow River dam safety monitoring, and in particular to a lightweight Yellow River underwater signal classification method and system based on knowledge distillation. Background Art

[0002] The Yellow River basin is vast, with complex currents and turbid water. The underwater environment is highly complex and variable. This makes the collection and classification of underwater signals in the Yellow River a very challenging task. Underwater signals, including acoustic signals, vibration signals, and electromagnetic signals, play a vital role in many applications, such as underwater communications, water quality monitoring, and underwater detection. However, due to the particularity of the Yellow River waters (such as high turbidity, large amounts of sediment, complex currents and climatic conditions, etc.), traditional underwater signal classification methods face many difficulties.

[0003] Traditional signal classification methods rely on artificial feature extraction and traditional signal processing algorithms, such as spectrum analysis and wavelet transform. These methods can identify different types of signals to a certain extent, but they are not robust enough in the special underwater environment of the Yellow River dam and are difficult to adapt to multi-signal interference scenarios such as the Yellow River Basin. Convolutional neural networks (CNNs) perform well in underwater signal classification, but their high computational cost limits their application in embedded devices. Therefore, there is an urgent need for an underwater signal classification method with both high precision and low computational complexity to meet the actual needs of underwater communications in the Yellow River Basin. Summary of the invention

[0004] In response to the above problems, the present invention proposes a lightweight Yellow River underwater signal classification method and system based on knowledge distillation. This lightweight classification method successfully copes with the challenge of implementing deep learning on resource-constrained systems. It can effectively reduce computing costs while maintaining excellent signal classification performance, thus meeting the actual needs of underwater communications in the Yellow River basin.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] On the one hand, the present invention proposes a lightweight Yellow River underwater signal classification method based on knowledge distillation, comprising:

[0007] Step 1: Collect the original data of underwater acoustic signals from the Yellow River Basin and perform normalization processing;

[0008] Step 2: Use a sliding window to divide the normalized underwater acoustic signal data into training, verification, and test samples;

[0009] Step 3: Input the training and validation samples into ResNet for pre-training; use the pre-trained ResNet as the teacher network in knowledge distillation, and use the sample relationship progressive knowledge distillation loss to guide the training of the GhostNet student network;

[0010] Step 4: Input the test samples into the trained GhostNet student network for underwater acoustic signal classification in the Yellow River Basin.

[0011] Further, in the above Step 1, the normalization process is carried out in the following manner:

[0012]

[0013] where x org is the original data of the underwater acoustic signal, μ is the mean of the original data of the underwater acoustic signal, and σ is the standard deviation.

[0014] Further, the above Step 2 includes:

[0015] Set the length and step size of the sliding window, and perform overlapping sampling on the collected underwater acoustic signals in the Yellow River Basin; adopt a stratified random division strategy, and allocate the training set, validation set and test set proportionally to ensure that the sample distributions of all categories are consistent; perform data augmentation on the training set samples.

[0016] Further, in the above Step 3, extract the feature map of the last convolutional layer of the teacher network as the soft target of knowledge distillation.

[0017] Further, the sample relationship progressive knowledge distillation loss is defined as follows:

[0018] L SRPKD = L SRKD + L PKD + L CE

[0019] where L SRPKD represents the sample relationship progressive knowledge distillation loss, L SRKD represents the sample correlation distillation loss; L PKD represents the progressive distillation loss; L CE represents the cross-entropy loss of the student model itself.

[0020] Further, the sample correlation distillation loss is calculated in the following manner:

[0021]

[0022] where

[0023]

[0024]

[0025] Among them, L is the number of layers of the teacher network and the student network. represents the SRKD loss of the corresponding layer l. represents the sample correlation based on self-attention between any two samples X i and X j in the l-th layer of the teacher model. represents the sample correlation based on self-attention between any two samples X i and X j in the l-th layer of the student model. and respectively represent the self-attention-based feature vectors of samples X i and X j in the l-th layer of the teacher model. · represents the dot product of two feature vectors, and || ||2 represents the L2 norm of the feature vector. and respectively represent the self-attention-based feature vectors of samples X i and X j in the l-th layer of the student model.

[0026] Furthermore, the progressive distillation loss is calculated as follows:

[0027] L PKD = L KL (P′ T , P S )

[0028] where

[0029] P′ T = P T · M + P T · (1 - M) · λ

[0030]

[0031] where L KL is the KL divergence function, P′ T is the output probability of the adjusted teacher model, P S is the output probability of the student model, P T represents the softened label; M is a binary mask, with elements corresponding to the correct class being 1 and others being 0; λ is a dynamic decay factor, L T represents the output of the teacher model, T is the temperature during softening, e represents the current training epoch of the student model, and E is the total number of epochs of the student model.

[0032] On the other hand, the present invention proposes a lightweight Yellow River underwater signal classification system based on knowledge distillation, including:

[0033] A signal acquisition module, which is used to collect the original underwater acoustic signal data from the Yellow River Basin and perform normalization processing;

[0034] A preprocessing module, which is used to divide the normalized underwater acoustic signal data into training, validation, and test samples using a sliding window;

[0035] A training module, which is used to input the training and validation samples into ResNet for pre-training; use the pre-trained ResNet as the teacher network in knowledge distillation, and use the progressive knowledge distillation loss of sample relationship to guide the training of the GhostNet student network;

[0036] A classification module, which is used to input the test samples into the trained GhostNet student network for classifying the underwater acoustic signals in the Yellow River Basin.

[0037] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the above-mentioned method is implemented.

[0038] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned method is implemented.

[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0040] The present invention first collects the original underwater acoustic signal data from the Yellow River Basin and performs normalization processing; then uses a sliding window to divide the normalized underwater acoustic signal data into training, validation, and test samples; inputs the training and validation samples into ResNet for pre-training; uses the pre-trained ResNet as the teacher network in knowledge distillation, and uses the progressive knowledge distillation loss of sample relationship to guide the training of the GhostNet student network; finally, inputs the test samples into the trained GhostNet student network for classifying the underwater acoustic signals in the Yellow River Basin. Through the present invention, while maintaining high signal classification accuracy, the number of parameters and the amount of calculation can be significantly reduced, and the efficiency of classifying the underwater signals in the Yellow River can be improved. In addition, due to the reduction of the amount of calculation, the present invention also reduces the requirements for hardware devices, enabling it to operate stably in an environment with relatively limited resources, and further improving the applicability and universality of the present invention in classifying the underwater signals in the Yellow River. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a schematic flowchart of a lightweight method for classifying underwater signals in the Yellow River based on knowledge distillation provided by an embodiment of the present invention;

[0042] Figure 2 Schematic diagram of the network training process provided by an embodiment of the present invention;

[0043] Figure 3 Schematic diagram of the architecture of a lightweight Yellow River underwater signal classification system based on knowledge distillation provided by an embodiment of the present invention;

[0044] Figure 4 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0045] The following further explains the present invention in conjunction with the accompanying drawings and specific embodiments:

[0046] As Figure 1 shown, a lightweight Yellow River underwater signal classification method based on knowledge distillation includes:

[0047] S101: At the signal acquisition end, collect the original underwater acoustic signal data from the Yellow River Basin and perform normalization processing.

[0048] Specifically, collect the original underwater acoustic signal data from underwater in the Yellow River Basin. To enable better subsequent analysis and classification, perform mean value and standard deviation adjustment on the data. The specific method is to use the normalization formula to eliminate the dimension difference and perform normalization processing:

[0049]

[0050] In the formula, x org is the original underwater acoustic signal data, μ is the mean value of the original underwater acoustic signal data, and σ is the standard deviation.

[0051] S102: At the preprocessing end, use a sliding window to divide the data set (composed of the normalized underwater acoustic signal data) into training, validation, and test samples;

[0052] Specifically, set the sliding window length to 1024 sampling points and the step size to 512 points, and perform overlapping sampling on the collected acoustic signals; adopt a stratified random division strategy, and allocate the training set, validation set, and test set in a ratio of 6:2:2 to ensure that the sample distributions of each category are consistent; perform data augmentation on the training set samples: add Gaussian white noise with a signal-to-noise ratio of 5 dB and perform random scaling with an amplitude of ±10%;

[0053] S103: At the training end, input the training and validation samples into ResNet for pre-training; use the pre-trained ResNet as the teacher network in knowledge distillation, and use the progressive knowledge distillation loss of the sample relationship to guide the training of the GhostNet student network;

[0054] Among them, at the training end, the training and validation samples are input into ResNet for pre-training. The main steps are as follows:

[0055] (1) Use the Adam optimizer for training. Set the initial learning rate to 3e-4 and apply the cosine annealing strategy. Set the batch size to 64;

[0056] (2) Design the cross-entropy loss as the loss function. Implement the early stopping strategy: terminate the training when the accuracy of the validation set has not improved for 5 consecutive epochs, and save the model parameters with the highest accuracy of the validation set;

[0057] (3) Extract the feature map of the last convolutional layer of the teacher network as the soft target for knowledge distillation.

[0058] Among them, at the training end, as Figure 2 shown, use the pre-trained ResNet as the teacher network in knowledge distillation, and use the Sample Relationship Progressive Knowledge Distillation (SRPKD) loss to guide the training of the GhostNet student network. The main steps are as follows:

[0059] (1) Calculate the SRKD loss. For the features output by each layer of the network, calculate its sample features based on the attention mechanism. Considering that the self-attention mechanism in Transformer can effectively capture the information of the key parts in the time series signal, especially in identifying the feature information of different signal types, the present invention introduces the self-attention mechanism to further enhance the feature information of the samples. Specifically, the input matrix vector Z is linearly transformed to generate the query matrix Q, the key matrix K, and the value matrix V. QK T is used to calculate the similarity matrix between the inputs, and a scaling factor is introduced to map the result of QK T to a reasonable range. This design effectively avoids the numerical instability problem that may be caused by directly using SoftMax, ensures the stability of gradient update, and thus improves the feature extraction ability of the model. Multiply the scaled result by V to obtain the weighted average output. Therefore, for a specific layer l of the teacher model, T l represents the features output by this layer, and the features enhanced by the self-attention mechanism can be defined as:

[0060]

[0061] Secondly, construct the correlation between any two samples based on the self-attention features. Considering that the cosine similarity is an effective method to measure the similarity of the directions of two vectors, especially suitable for calculating the correlation between samples in a noise interference environment, the present invention introduces the cosine similarity to calculate the correlation between sample features. For a specific layer l of the teacher model, any two samples Xi and X j Sample correlation based on self-attention can be defined as:

[0062]

[0063] Similarly, for a specific layer l of the student model, for any two samples X i and X j Sample correlation based on self-attention can be defined as:

[0064]

[0065] where, and respectively represent the self-attention-based feature vectors of samples X i and X j in the l-th layer of the teacher model, and respectively represent the self-attention-based feature vectors of samples X i and X j in the l-th layer of the student model, · represents the dot product of the two feature vectors, and || ||2 represents the L2 norm of the feature vector.

[0066] Next, the sample correlation of a certain layer in the teacher model is passed to the student model and defined as the SRKD loss of this layer. The present invention draws on the idea of the L1 loss to construct the distillation loss. For a given specific layer l, its SRKD loss can be defined as:

[0067]

[0068] Finally, the SRKD losses of all single layers are added up and defined as the overall network SRKD loss, specifically as follows:

[0069]

[0070] where L is the number of layers of the teacher network and the student network. SRKD captures the correlation between samples at different levels of the network through the self-attention mechanism and provides more useful knowledge for the GhostNet student network.

[0071] (2) Calculate the PKD loss. To transfer more comprehensive knowledge, the present invention introduces a traditional Logits-based knowledge distillation method. However, the Logits-based knowledge contains negative information of incorrect predictions for each instance, and as training progresses, this negative information may hinder the further improvement of the student model's performance. To solve this problem, the present invention proposes a Progressive Knowledge Distillation (PKD) method, which makes the Logits-based knowledge more effective and valuable by gradually reducing the output of incorrect Logits during the distillation process. The specific implementation process of PKD can be described as follows:

[0072] In the first step, process the Logits output by the teacher model to generate softened labels through the Softmax function. For the output L of the teacher model T , the softened label P T can be defined as:

[0073]

[0074] where T is the temperature during softening.

[0075] In the second step, adjust the softened probabilities so that the probabilities of correct Logits remain unchanged and the probabilities of incorrect Logits gradually decrease. To achieve this process, the present invention adopts a dynamic adjustment strategy. Let λ be the decay factor, which changes with the training stage. The adjusted probability P' T can be expressed as:

[0076] P' T = P T ·M + P T ·(1 - M)·λ (8)

[0077] where M is a binary mask, the elements corresponding to the correct class are 1, and others are 0; λ is the dynamic decay factor, defined as follows:

[0078]

[0079] where e represents the current training epoch of the student model, and E is the total number of epochs of the student model.

[0080] In the third step, calculate the KL loss between the adjusted teacher model output and the student model output. The progressive distillation loss L PKD can be defined as:

[0081] L PKD = L KL (P' T , P S ) (10)

[0082] where LKL is the KL divergence function, and P' T is the output probability of the adjusted teacher model, and P S is the output probability of the student model. PKD dynamically adjusts the knowledge transfer strategy to enable the student model to obtain the most suitable knowledge at different training stages. In the initial stage of distillation, the teacher model provides comprehensive knowledge to the student model, thus quickly improving the diagnostic ability of the student; as the training progresses, the teacher model gradually reduces the impact of mispredictions to enhance the generalization ability of the student model.

[0083] (3) Calculate the overall distillation loss. The present invention simultaneously considers the sample - related knowledge (SRKD), progressive knowledge (PKD), and the cross - entropy knowledge (CE) of the student model itself. The teacher model transfers rich knowledge to the student model, enabling high - quality knowledge distillation and thus improving the overall performance of the diagnostic model. The overall distillation loss of SRPKD is defined as follows:

[0084] L SRPKD = L SRKD + L PKD + L CE (11)

[0085] where, L SRKD represents the sample - related distillation loss; L PKD represents the progressive distillation loss; L CE represents the cross - entropy loss of the student model itself.

[0086] S104: Input the test sample into the trained GhostNet network for signal classification.

[0087] Specifically, with the test sample as the model input, it first undergoes a 16 * 16 convolution for preliminary feature extraction. The BN layer is used to accelerate the network convergence speed and prevent overfitting, and the ReLU activation function is used to enhance the model's expressive ability; then the features obtained after the max - pooling operation are used as the input of the GhostModule. The model extracts features from the input through 4 GhostModules with the same structure, and the output feature data is output as the final result through the average - pooling operation and the fully - connected layer.

[0088] Specifically, the network structure of the GhostNet network is as follows:

[0089] Table 1 GhostNet network structure

[0090]

[0091]

[0092] Corresponding to the above method, such asFigure 3 As shown in the figure, an embodiment of the present invention further provides a lightweight underwater Yellow River signal classification system based on knowledge distillation, including: a signal acquisition module, a preprocessing module, a training module, and a classification module.

[0093] Among them, the signal acquisition module is set at the acquisition end, and is used to collect the original underwater acoustic signal data from the Yellow River Basin and perform normalization processing; the preprocessing module is set at the preprocessing end, and is used to divide the normalized underwater acoustic signal data into training, verification, and test samples by using a sliding window; the training module is set at the training end, and is used to input the training and verification samples into ResNet for pre-training; the pre-trained ResNet is used as the teacher network in knowledge distillation, and the progressive knowledge distillation loss of sample relationship is used to guide the training of the GhostNet student network; the classification module is set at the classification end, and is used to input the test samples into the trained GhostNet student network for classifying the underwater acoustic signals in the Yellow River Basin.

[0094] Figure 4 An example of the physical structure diagram of an electronic device is shown in Figure 4 As shown in the figure, the electronic device may include: a processor 401, a communication interface 402, a memory 403, and a communication bus 404. Among them, the processor 401, the communication interface 402, and the memory 403 complete mutual communication through the communication bus 404. The processor 401 can call the logical instructions in the memory 403 to execute the Yellow River underwater signal classification method, which includes: at the signal acquisition end, collecting the original underwater acoustic signal data of the Yellow River and performing normalization processing; at the preprocessing end, dividing it into training, verification, and test sets through a sliding window; at the training end, using ResNet as the teacher network for pre-training, combining the progressive knowledge distillation (SRPKD) loss of sample relationship, and migrating the knowledge to the lightweight GhostNet student network; at the classification end, extracting signal features and compressing the computational amount through the GhostModule unit, and finally realizing real-time classification.

[0095] In addition, when the logical instructions in the above-mentioned memory 403 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0096] The embodiments of the present invention also provide a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided in the above-mentioned method embodiments. For example, it includes: at the signal acquisition end, acquiring the original data of the underwater acoustic signals of the Yellow River and performing normalization processing; at the preprocessing end, dividing into training, validation, and test sets through a sliding window; at the training end, using ResNet as a teacher network for pre-training, combining the sample relationship progressive knowledge distillation (SRPKD) loss, and migrating the knowledge to the lightweight GhostNet student network; at the classification end, extracting signal features through the GhostModule unit and compressing the computational amount to finally achieve real-time classification.

[0097] The embodiments of the present invention also provide a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the methods provided in the above-mentioned method embodiments. For example, it includes: at the signal acquisition end, acquiring the original data of the underwater acoustic signals of the Yellow River and performing normalization processing; at the preprocessing end, dividing into training, validation, and test sets through a sliding window; at the training end, using ResNet as a teacher network for pre-training, combining the sample relationship progressive knowledge distillation (SRPKD) loss, and migrating the knowledge to the lightweight GhostNet student network; at the classification end, extracting signal features through the GhostModule unit and compressing the computational amount to finally achieve real-time classification.

[0098] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0099] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A lightweight underwater Yellow River signal classification method based on knowledge distillation, characterized in that, Including: Step 1: Collect the original underwater acoustic signal data from the Yellow River Basin and perform normalization processing; Step 2: Use a sliding window to divide the normalized underwater acoustic signal data into training, validation, and test samples; Step 3: Input the training and validation samples into ResNet for pre-training; use the pre-trained ResNet as the teacher network in knowledge distillation, and use the sample relationship progressive knowledge distillation loss to guide the training of the GhostNet student network; Step 4: Input the test samples into the trained GhostNet student network for classifying the underwater acoustic signals in the Yellow River Basin.

2. The lightweight Yellow River underwater signal classification method based on knowledge distillation according to claim 1, wherein In the said Step 1, the normalization processing is carried out in the following manner: where x org is the original data of the underwater acoustic signal, μ is the mean value of the original data of the underwater acoustic signal, and σ is the standard deviation.

3. The lightweight Yellow River underwater signal classification method based on knowledge distillation according to claim 1, characterized in that, The said Step 2 includes: Set the length and step size of the sliding window, and perform overlapping sampling on the collected underwater acoustic signals in the Yellow River Basin; adopt a stratified random division strategy, and allocate the training set, validation set, and test set proportionally to ensure the consistent distribution of samples in each category; perform data augmentation on the training set samples.

4. The lightweight Yellow River underwater signal classification method based on knowledge distillation according to claim 1, characterized in that, In the said Step 3, extract the feature map of the last convolutional layer of the teacher network as the soft target of knowledge distillation.

5. The lightweight Yellow River underwater signal classification method based on knowledge distillation according to claim 1, wherein The said sample relationship progressive knowledge distillation loss is defined as follows: L SRPKD = L SRKD + L PKD + L CE Among them, L SRPKD represents the sample relationship progressive knowledge distillation loss, and L SRKD represents the sample correlation distillation loss; L PKD represents the progressive distillation loss; L CE represents the cross-entropy loss of the student model itself.

6. The lightweight Yellow River underwater signal classification method based on knowledge distillation according to claim 5, characterized in that The said sample correlation distillation loss is calculated in the following manner: wherein Among them, L is the number of layers of the teacher network and the student network. denotes the SRKD loss of the corresponding layer l. denotes the sample correlation based on self-attention of any two samples X i and X j in the teacher model at the corresponding layer l. denotes the sample correlation based on self-attention of any two samples X i and X j in the student model at the corresponding layer l. and respectively denote the self-attention-based feature vectors of samples X i and X j in the l-th layer of the teacher model. and respectively denote the self-attention-based feature vectors of samples X i and X j in the l-th layer of the student model. · denotes the dot product of two feature vectors, and || ||2 denotes the L2 norm of the feature vector.

7. The lightweight Yellow River underwater signal classification method based on knowledge distillation according to claim 5, characterized in that The said progressive distillation loss is calculated in the following manner: L PKD = L KL (P T ′, P S ) wherein P′ T = P T ·M + P T ·(1 - M)·λ Among them, L KL is the KL divergence function, P′ T is the output probability of the adjusted teacher model, P S is the output probability of the student model, P T represents the label after softening; M is a binary mask, the elements corresponding to the correct category are 1, and the others are 0; λ is a dynamic decay factor, L T represents the output of the teacher model, T is the temperature during softening, e represents the current training epoch of the student model, and E is the total number of epochs of the student model.

8. A lightweight underwater Yellow River signal classification system based on knowledge distillation, characterized in that, Including: A signal acquisition module, which is used to collect the original underwater acoustic signal data from the Yellow River Basin and perform normalization processing; A preprocessing module, which is used to use a sliding window to divide the normalized underwater acoustic signal data into training, validation, and test samples; A training module, which is used to input the training and validation samples into ResNet for pre-training; use the pre-trained ResNet as the teacher network in knowledge distillation, and use the sample relationship progressive knowledge distillation loss to guide the training of the GhostNet student network; A classification module, which is used to input the test samples into the trained GhostNet student network for classifying the underwater acoustic signals in the Yellow River Basin.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the said processor executes the said program, it implements the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing a computer program thereon, characterized in that, When the said computer program is executed by the processor, it implements the method according to any one of claims 1 to 7.