Dynamic unmanned underwater vehicle echo feature-oriented deep learning identification method and system and application
By constructing an improved feature fusion-recognition network, fusing multi-domain features and utilizing deep learning methods, the problem of insufficient recognition accuracy of dynamic UUVs in complex marine environments was solved, achieving higher recognition accuracy and robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies struggle to effectively identify the acoustic characteristics of dynamic unmanned underwater vehicles. Traditional methods are susceptible to contamination or distortion in feature extraction within complex marine environments, and lack the ability to identify UUV radiated noise, resulting in insufficient identification accuracy and weak generalization ability.
An improved feature fusion-recognition network is constructed, which integrates beam domain, time-frequency domain, and spatial-frequency domain features. Through cross-modal feature interaction and computation, and the collaborative work of multilayer perceptron branches and KAN branches, accurate UUV recognition is achieved.
It improves the accuracy and robustness of UUV identification, enabling accurate identification of UUVs in complex environments and enhancing feature representation capabilities and matching effects.
Smart Images

Figure CN121634064A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of calculation, estimation or counting, in particular to a deep learning identification method, system and application for dynamic unmanned underwater vehicle echo characteristics. BACKGROUND
[0002] With the continuous rise of the global population and the continuous progress of human society, the contradiction between the accelerated consumption of land resources and the increasingly tense surface space has become increasingly prominent. In this context, fully developing the vast space and rich resources contained in the ocean has become a key strategy to promote the steady progress of human society. To translate this strategic blueprint into reality, the core intelligent equipment, unmanned underwater vehicle, cannot be ignored.
[0003] Unmanned Underwater Vehicle (UUV) as the core intelligent equipment of ocean development, has unique advantages such as safety, reliability, autonomy and efficiency, and is a valuable tool to ensure the safety and efficiency of marine activities. It can be used for various underwater activities, including resource exploration, facility construction, rescue and seafood fishing.
[0004] However, as a new type of intelligent underwater platform, the acoustic stealth performance and maneuvering characteristics of UUV pose a serious challenge to traditional underwater acoustic target recognition technology. UUV has unique low signal-to-noise ratio acoustic characteristics. When encountering complex marine environment, interference of other marine equipment, etc., the extracted features will be easily contaminated or distorted. The underwater target recognition model based on traditional feature extraction faces the core bottlenecks of insufficient feature representation and weak cross-domain generalization. More importantly, most sonar equipment is designed for anti-submarine needs and almost has no recognition ability for UUV radiated noise, which has a serious shortcoming. Therefore, it is urgent to innovate the recognition method suitable for the acoustic characteristics of UUV to provide technical support for the development of anti-UUV equipment.
[0005] In order to cope with these challenges, it is necessary to redesign the deep learning recognition method suitable for UUV, especially for the special needs of UUV in dynamic motion scene, and to propose a more robust solution. SUMMARY
[0006] In view of the problems existing in the prior art, the present application provides a deep learning identification method, system and application for dynamic unmanned underwater vehicle echo characteristics.
[0007] The technical scheme adopted by the present application is a deep learning identification method for dynamic unmanned underwater vehicle echo characteristics. The method constructs an improved feature fusion-recognition network for fusing features of different domains and realizing target recognition based on the fused features.
[0008] training the improved feature fusion-recognition network based on the basic features of the echo signals;
[0009] acquiring echo signals, inputting the basic features into the trained improved feature fusion-recognition network to obtain the recognition result of the unmanned underwater vehicle.
[0010] Preferably, the basic features of the echo signals include beam domain features, time-frequency domain features and space-frequency domain features.
[0011] Preferably, a plurality of real numbers are randomly generated by a sonar array, and corresponding angles are defined, active sonar echo signals are acquired based on the angles, taken as beam domain data, and the beam domain features are obtained after preprocessing, and the space-frequency domain features and the time-frequency domain features are extracted based on the beam domain data.
[0012] Preferably, a data extraction window is dynamically delimited with the target coordinates of the target unmanned underwater vehicle to be identified as the center, the space-frequency domain features are cropped, and the beam domain features and the time-frequency domain features are correspondingly compressed.
[0013] Preferably, the improved feature fusion-recognition network comprises an improved feature fusion unit and an improved unmanned underwater vehicle recognition unit arranged in sequence.
[0014] Preferably, the improved feature fusion unit comprises a multi-domain feature input and primary feature extraction module, a cross-modal feature interaction and operation module, a high-order fusion and feature reconstruction module, and a fusion feature output module.
[0015] The cross-modal feature interaction and operation module performs a transpose operation on a specific feature matrix in the primary feature matrix group output by the multi-domain feature input and primary feature extraction module, and obtains an interaction feature after performing a matrix multiplication operation on the feature matrix adjusted by the transpose operation and other feature matrices; the transpose operation here involves three groups of feature matrices, which are performed in pairs, and there is a specific feature matrix that needs to be transposed in each group of feature matrices.
[0016] The high-order fusion and feature reconstruction module further fuses the interaction feature into a transit high-order fusion feature, and fuses the transit high-order fusion feature and the primary feature to obtain a high-dimensional fusion feature.
[0017] The fusion feature output module fuses the high-dimensional fusion feature to generate a high-order fusion feature.
[0018] Preferably, the improved unmanned underwater vehicle recognition unit comprises a feature extraction module and a recognition module connected in sequence.
[0019] The feature extraction module comprises two first extraction blocks and four second extraction blocks arranged in sequence.
[0020] The identification module comprises a plurality of layers of a perceptron branch and a KAN branch arranged side by side, and the outputs of the perceptron branch and the KAN branch are outputted after weighted summation.
[0021] Preferably, the first-level extraction block comprises two convolution layers, a self-attention module and a max-pooling layer arranged in sequence.
[0022] The second-level extraction block comprises three convolution layers, a self-attention module and a max-pooling layer arranged in sequence.
[0023] A deep learning identification system for dynamic unmanned underwater vehicle echo features comprises:
[0024] At least one processor; and
[0025] A memory in communication connection with the at least one processor; wherein,
[0026] The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to realize the deep learning identification method for dynamic unmanned underwater vehicle echo features.
[0027] An application of the deep learning identification method for dynamic unmanned underwater vehicle echo features is applied to extracting multi-domain features in original echo signals and performing complementary and enhancement before executing the identification of unmanned underwater vehicles.
[0028] The present application relates to a deep learning identification method, system and application for dynamic unmanned underwater vehicle echo features, an improved feature fusion-identification network is constructed, which is used to fuse features of different domains and realize target identification based on the fused features; the improved feature fusion-identification network is trained based on the basic features of echo signals; the echo signals are acquired, the basic features are extracted and input into the trained improved feature fusion-identification network, and the identification result of the unmanned underwater vehicle is obtained; the method realizes the system, and is applied to extracting multi-domain features in original echo signals and performing complementary and enhancement before executing the identification of unmanned underwater vehicles.
[0029] The present application has the following advantages:
[0030] (1) Multi-domain features are extracted in original echo signals, composite beam domain features, time-frequency domain features and space-frequency domain features are extracted, and the problem of single UUV feature type is solved;
[0031] (2) A feature fusion scheme is proposed to solve the problem of insufficient fusion between features, which not only considers the related information in a single data, but also considers the complementary and enhancement effect between data, so as to improve the matching effect and enhance the representation ability of features;
[0032] (3) In view of the performance deficiency of the existing general underwater acoustic recognition network in identifying UUVs, a UUV-specific recognition unit is proposed, which can more accurately capture the acoustic features of UUVs and perform recognition and classification, thereby improving the accuracy of the model in identifying UUVs. BRIEF DESCRIPTION OF DRAWINGS
[0033] Figure 1 is a flowchart of the present application;
[0034] Figure 2 is a structural diagram of the improved feature fusion-recognition network in the present application;
[0035] Figure 3 is a structural diagram of the improved feature fusion unit in the present application;
[0036] Figure 4 is a structural diagram of the improved unmanned underwater vehicle recognition unit in the present application. DETAILED DESCRIPTION
[0037] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the embodiments. It should be understood that the specific implementation described herein is only used to explain the present application and does not limit the present application.
[0038] As shown in Figure 1 The present application relates to a deep learning recognition method for dynamic unmanned underwater vehicle echo features, which comprises the following steps:
[0039] (1) An improved feature fusion-recognition network is constructed, which is used to fuse features of different domains and realize target recognition based on the fused features;
[0040] (2) The improved feature fusion-recognition network is trained based on the basic features of the echo signal;
[0041] (3) The echo signal is obtained, the basic features are extracted, and then the trained improved feature fusion-recognition network is inputted to obtain the recognition result of the unmanned underwater vehicle.
[0042] The steps are described below in combination with the specific implementation method.
[0043] (1) An improved feature fusion-recognition network is constructed, which is used to fuse features of different domains and realize target recognition based on the fused features;
[0044] As shown in Figure 2As shown, the improved feature fusion-recognition network comprises an improved feature fusion unit and an improved unmanned underwater vehicle recognition unit arranged in sequence; the former can fully fuse features of different domains, not only considering relevant information within a single data, but also considering the complementary and enhancing effect between data, thereby improving the matching effect, and the latter introduces attention selective emphasis on key regions and channels while maintaining multi-scale spatial semantics, significantly improving the accuracy and robustness of UUV recognition.
[0045] (1-1) As shown in the improved feature fusion-recognition network, the improved feature fusion unit comprises a multi-domain feature input and primary feature extraction module, a cross-modal feature interaction and operation module, a high-order fusion and feature reconstruction module, and a fusion feature output module. Figure 3
[0046] The modules are described one by one as follows.
[0047] (1-1-1) Multi-domain feature input and primary feature extraction module
[0048] The multi-domain feature input and primary feature extraction module is used to convert the original features into primary abstract feature representations more suitable for the recognition task.
[0049] First, feature inputs from three different signal processing domains are received, including:
[0050] Beam domain feature (B): derived from the beamforming output of the sensor array and subjected to simple data preprocessing (downsampling);
[0051] Time-frequency domain feature (T): derived from the time-frequency spectrum of the short-time Fourier transform;
[0052] Space-frequency domain feature (S): derived from the joint analysis of space and frequency, representing the joint distribution of signals in space and frequency;
[0053] Subsequently, primary feature extraction is performed, and each domain signal is processed by an independent feature extraction network to generate a set of high-dimensional abstract feature vectors, including:
[0054] Beam domain primary feature: , , ;
[0055] Time-frequency domain primary feature: , , ;
[0056] Space-frequency domain primary feature: , , ;
[0057] The independent feature extraction network here mainly uses 1*1 convolution with a step of 1 and padding of 0, so that the number of channels is doubled.
[0058] (1-1-2) Cross-modal feature interaction and operation module
[0059] To realize the depth information complementation and enhancement between different domain features, a series of structured tensor operations are performed to construct cross-modal interaction features. Specifically, the cross-modal feature interaction and operation module performs a transpose operation on a specific feature matrix in the primary feature matrix group output by the multi-domain feature input and the primary feature extraction module, and obtains an interaction feature after matrix multiplication of the feature matrix adjusted by the transpose operation and other feature matrices.
[0060] The transpose operation is to perform a transpose operation on a specific feature matrix, aiming to adjust the feature dimension structure and prepare for dimension alignment for subsequent matrix multiplication. This operation changes the view of the feature tensor without changing its essential information. The calculation process satisfies:
[0061]
[0062] Wherein, A is an m n matrix, and its transpose is an n m matrix;
[0063] The feature matrix adjusted by the transpose operation and other feature matrices are subjected to matrix multiplication. This operation, as an affine transformation and feature combination, can generate a new feature representation that contains the complex correlation between the two feature modes. When two matrices come from different sources or different transformed feature matrices, their product will generate a covariance interaction matrix. Each element of the interaction matrix captures the strength of the correlation between specific feature dimensions from the two matrices. For example, in the figure, the feature T is multiplied by the transposed feature S to generate the interaction feature ST, which satisfies:
[0064]
[0065] Wherein, A is an m n matrix, B is an n p matrix, and the matrix product C is an m p matrix;
[0066] Through the above operation, the system generates a series of intermediate fusion features with rich cross-modal information, such as ST, BS, and BT.
[0067] (1-1-3) High-order fusion and feature reconstruction module
[0068] The high-order fusion and feature reconstruction module further fuses the interaction features into transit high-order fusion features through a second round of tensor operation, and fuses the transit high-order fusion features and the primary features to obtain high-dimensional fusion features.
[0069] To this end, element-wise multiplication operations (Hadamard product) need to be performed on feature matrices from the same or different domains to realize feature modulation or gating mechanisms, highlight shared significant information, and suppress irrelevant features; this operation is a kind of soft mask, and the values in one matrix can determine whether the elements at the corresponding positions in another matrix should be amplified, retained, or suppressed, which is a dynamic feature selection mechanism; for example, the Hadamard product of the features ST, BS, and BT generates the interaction feature BST, which satisfies:
[0070]
[0071] where A is an m n matrix, B is an m n matrix, and the calculation result C is also an m n matrix.
[0072] Specifically, the Hadamard product of the interaction features ST, BS, and BT generates the transit high-order fusion feature BST, and the Hadamard product of the remaining three primary features and the transit high-order fusion feature BST generates three high-dimensional fusion features (features 1, 2, and 3), realizing the evolution from bimodal interaction to trimodal and higher-order collaboration, and the generated high-order fusion features maximally integrate the complementary information of the beam, time-frequency, and space-frequency domains, forming a global and context-aware feature representation.
[0073] (1-1-4) Fusion feature output module
[0074] The high-dimensional fusion features are fused to generate high-order fusion features.
[0075] The generated three high-dimensional fusion features are Hadamarded to finally generate high-order fusion features, which will serve as a series of final fusion features.
[0076] The high-order fusion features represent the essence of the input signal after multi-level and multi-angle fusion, and have high discriminability, strong robustness, and compactness.
[0077] In the present application, the improved feature fusion unit discards the traditional splicing and simple weighting strategy, and instead adopts a multi-level fusion path based on structured tensor operation, receives multiple primary features from the beam domain, time-frequency domain and space-frequency domain, and realizes shallow interaction to deep fusion through the cooperative combination of transpose operation, matrix multiplication and Hadamard product.
[0078] (1-2) As shown in Figure 4 The improved unmanned underwater vehicle recognition unit comprises a feature extraction module and a recognition module connected in sequence.
[0079] (1-2-1) The feature extraction module comprises two first extraction blocks and four second extraction blocks arranged in sequence.
[0080] The first extraction block comprises two convolution layers, a self-attention module and a max-pooling layer arranged in sequence.
[0081] The second extraction block comprises three convolution layers, a self-attention module and a max-pooling layer arranged in sequence.
[0082] In the present application, the first first extraction block applies 3x3 convolution twice in sequence, with a channel number of 64, and a multi-head self-attention mechanism module is introduced after convolution to explicitly model long-range dependencies and suppress background interference, and finally max-pooling is used for down-sampling to obtain multi-scale representation and enhance feature robustness; the second first extraction block repeats the structure of the previous one, and the channel number is increased to 128; then the first second extraction block, which increases the number of convolutions to 3 and the channel number to 256; the second second extraction block repeats the structure of the previous one, and the channel number is increased to 512; the third and fourth second extraction blocks both repeat the structure of the previous one, and finally the channel number is increased to 1024, and the receptive field is gradually expanded in the process.
[0083] After multiple convolutions and pooling, the feature map is compressed into a low spatial resolution, high channel number tensor.
[0084] (1-2-2) The recognition module comprises a multi-layer perceptron branch and a KAN branch arranged side by side, and the outputs of the multi-layer perceptron branch and the KAN branch are output after weighted summation.
[0085] Among them, the multi-layer perceptron (MLP) branch carries out nonlinear feature transformation through fully connected layer and fixed activation function, and its core is the linear combination of weight matrix and input feature followed by activation function, which provides strong fitting ability as a standard feature processor in deep learning; the KAN branch adopts a learnable spline function to replace the traditional activation function, and its weight parameter is replaced by an adaptive adjustment unary function, thereby significantly reducing the parameter amount while maintaining high expressiveness, and due to its inherent interpretability, it is more suitable for modeling complex nonlinear relationships;
[0086] In fact, the effect of KAN is better than that of MLP, and the actual output of KAN is consistent with that of MLP, that is, the probability of whether the target is recognized, considering that KAN is slower in processing speed due to a large number of parameters, therefore, weights a and b are attached to KAN and MLP respectively, and the weighted sum of the output results [A, B] and [C, D] of the two is taken in [a * A + b * C, a * B + b * D], that is, the one with higher probability is taken, and a and b are automatically adjusted in the model training process; in the process, MLP is responsible for regular feature transformation and compression, and KAN is used as a more efficient and flexible classifier or feature enhancer, and the two work together to improve the model performance and efficiency. The weighted sum of the outputs of the two modules is taken as the final logarithmic probability, and the one with higher probability is taken as the recognition result.
[0087] (2) training the improved feature fusion-recognition network based on the basic features of the echo signal;
[0088] The basic features of the echo signal include beam domain features, time-frequency domain features and space-frequency domain features.
[0089] The original data comes from the active sonar echo signal obtained by the sonar array from 180 angles and 8000 frequency data, and then the basic features are extracted therefrom to construct a data set;
[0090] Specifically, a plurality of real numbers are randomly generated by the sonar array, such as 180, and the corresponding angles are defined, and the 180 angles are determined by the inverse trigonometric function here; in the signal acquisition process, the range parameter is set to [1°, 180°] to collect signals from different angles to the array, and to ensure that the intervals between the angles are not completely the same; the angle selection satisfies wherein, y represents a real number;
[0091] The active sonar echo signal is obtained based on the angle, and 8000 frequency data are collected at each angle as beam domain data, and the beam domain features are obtained after preprocessing, and the space-frequency domain features and the time-frequency domain features are extracted based on the beam domain data;
[0092] (2-1) extracting the space-frequency domain features according to the collected signals
[0093] Fast Fourier transform (FFT) is used to implement, specifically:
[0094] The two-dimensional Fourier transform is performed on the beam domain data to obtain complex spectrum data,
[0095]
[0096] Wherein, F(u, v) is a two-dimensional frequency domain function, u and v are frequency variables, j is a virtual unit, x and y are space variables, after FFT transformation, the formed space spectrum diagram is of complex number type;
[0097] Then the amplitude and phase of the spectrum data are calculated, the amplitude represents the intensity of each frequency component, and the phase represents the displacement of each frequency component. The spectrum diagram is centralized so that the low frequency component is located at the center of the spectrum diagram, facilitating observation and analysis;
[0098] (2-2) Extracting time-frequency domain features according to beam domain data
[0099] Short-time Fourier transform (STFT) is used for implementation. STFT can provide local information of signals in time and frequency, thereby effectively analyzing the frequency characteristics of time-varying signals. Specifically:
[0100] The beam domain data is analyzed and target positioning is analyzed to confirm the spatial coordinates of the target to be identified. According to the abscissa of the spatial coordinates, i.e. the beam number, 8000 data contained in the beam number are extracted from the beam domain data.
[0101] The continuous 8000 data are divided into shorter time segments, usually by multiplying a sliding window function, such as the Hanning window, the Hamming window, etc., as follows:
[0102]
[0103] Wherein, x(t) is the original signal, represents the time offset, is a window function, used to intercept the local time domain region of the signal, and determines the resolution of the signal in the time domain, is the kernel function of Fourier transform;
[0104] Fourier transform, Fourier transform is applied to each time segment to calculate the frequency components of the signal in the time segment; the Fourier transform results of all time segments are combined to obtain the time-frequency representation of the signal, and the complex number result is taken as a modulus to convert it to a real number type for subsequent operation and application.
[0105] It should be noted that directly processing the full amount of beam domain data can easily lead to huge consumption of computing resources and low processing efficiency. Therefore, the target coordinates of the unmanned underwater vehicle to be identified are taken as the center to dynamically define the data extraction window, such as 64 units x 64 units, to crop the space-frequency domain features, and the beam domain features and time-frequency domain features are correspondingly compressed, which can achieve an optimal balance between retaining key feature information of the target and maximizing the suppression of redundant data.
[0106] The data set is divided into 7:3 as a training set and a test set.
[0107] In the training, the loss function of the improved feature fusion-recognition network is constructed, which is associated with the classification prediction result and the true result of the training sample; in the embodiment, a cross-entropy loss function is adopted, and the prediction probability of the UUV recognition label is evaluated to ensure that the model can improve the recognition accuracy of the dynamic UUV.
[0108] After the training is completed, the improved feature fusion-recognition network is tested by using the test set.
[0109] (3) After the echo signal is acquired and the basic features are extracted, the improved feature fusion-recognition network trained is inputted, and the recognition result of the unmanned underwater vehicle is obtained.
[0110] Through experiments, the performance of the method is obviously better than that of the classical deep learning algorithm and the underwater acoustic recognition algorithm from the overall performance, which reflects the superiority of the method in recognizing the dynamic UUV, and the superiority is due to the following points:
[0111] Firstly, the mechanism based on the feature fusion network can effectively extract the visual features of the UUV under different domains and receptive fields, and the robustness to scale changes and occlusions is enhanced through cross-domain feature fusion, and the integrity of the feature representation is improved;
[0112] Secondly, the robust extraction of multi-scale features is realized through the deep convolution and the multi-head self-attention mechanism, and the representation ability and performance of the model are further improved, and the discrimination ability of the model in the variable underwater environment is enhanced;
[0113] In addition, through the collaborative training of the MLP and the KAN, the nonlinear fitting ability of the MLP and the knowledge guiding characteristics of the KAN are fully utilized, the weights are dynamically adjusted, the model can better balance the data-driven learning and the knowledge-guided reasoning in different scenes, and the overall performance is improved;
[0114] Finally, the classification network is combined with the cross-entropy loss, and the feature extraction network is cooperatively optimized, and the robustness and accuracy of the overall framework under various complex conditions are ensured.
[0115] These designs jointly lay the significant performance advantage of the method in the complex environment.
[0116] The application also relates to a deep learning recognition system for echo features of dynamic unmanned underwater vehicles, which comprises:
[0117] at least one processor; and
[0118] a memory in communication connection with the at least one processor; wherein
[0119] The memory stores instructions executable by the processor, the instructions being for execution by the processor to implement the deep learning identification method for dynamic unmanned underwater vehicle echo characteristics.
[0120] The application further relates to an application of the deep learning identification method for dynamic unmanned underwater vehicle echo characteristics, which is applied to the identification of unmanned underwater vehicles after multi-domain features are extracted from original echo signals and are complemented and enhanced.
[0121] Those skilled in the art will understand that embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0122] The present application is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus generate an apparatus that implements the flow Figure 1 one or more flows and / or blocks Figure 1 means for performing the function specified by one or more of the flows or blocks.
[0123] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including instruction means, which implements the flow Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0124] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable data processing apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable data processing apparatus provide a process for implementing the flow Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0125] While the preferred embodiments of the application have been described, additional variations and modifications can be made to these embodiments by those skilled in the art once they have the benefit of the present disclosure without departing from the spirit and scope of the application. Accordingly, it is intended that the appended claims include all such modifications and variations as fall within the scope of the present application.
[0126] It is apparent that those skilled in the art can make various changes and modifications to the application without departing from the spirit and scope of the application. It is therefore intended that the present application cover all such changes and modifications that are within its scope.
Claims
1. A deep learning identification method for dynamic unmanned underwater vehicle echo characteristics, characterized in that: The method constructs an improved feature fusion-recognition network for fusing features of different domains and realizing target recognition based on the fused features; The improved feature fusion-recognition network is trained based on basic features of echo signals; After obtaining the echo signals and extracting the basic features, the improved feature fusion-recognition network is inputted to obtain the recognition result of the unmanned underwater vehicle.
2. The method of claim 1, wherein: The basic features of the echo signals include beam domain features, time-frequency domain features and space-frequency domain features.
3. The method of claim 2, wherein: Through a sonar array, a plurality of real numbers are randomly generated and corresponding angles are defined, the active sonar echo signals are obtained based on the angles as beam domain data, the beam domain features are obtained after preprocessing, and the space-frequency domain features and the time-frequency domain features are extracted based on the beam domain data.
4. The method of claim 3, wherein: A data extraction window is dynamically delimited with the target coordinates of the target unmanned underwater vehicle to be recognized as the center, the space-frequency domain features are cropped, and the beam domain features and the time-frequency domain features are correspondingly compressed.
5. The method of claim 1, wherein: The improved feature fusion-recognition network includes an improved feature fusion unit and an improved unmanned underwater vehicle recognition unit arranged in sequence.
6. The method of claim 5, wherein: The improved feature fusion unit includes a multi-domain feature input and primary feature extraction module, a cross-modal feature interaction and operation module, a high-order fusion and feature reconstruction module, and a fusion feature output module. The cross-modal feature interaction and operation module performs a transpose operation on a specific feature matrix in a primary feature matrix group output by the multi-domain feature input and primary feature extraction module, and obtains interaction features after matrix multiplication of the feature matrix adjusted by the transpose operation and other feature matrices. The high-order fusion and feature reconstruction module further fuses the interaction features into a transit high-order fusion feature, and fuses the transit high-order fusion feature and the primary feature to obtain a high-dimensional fusion feature. The fusion feature output module fuses the high-dimensional fusion feature to generate a high-order fusion feature.
7. The method of claim 5, wherein: The improved unmanned underwater vehicle recognition unit includes a feature extraction module and a recognition module connected in sequence. The feature extraction module includes two primary extraction blocks and four secondary extraction blocks arranged in sequence. The recognition module includes a multi-layer perceptron branch and a KAN branch arranged side by side, and the outputs of the multi-layer perceptron branch and the KAN branch are outputted after weighted summation.
8. The method of claim 7, wherein: The primary extraction block includes two convolution layers, a self-attention module and a max-pooling layer arranged in sequence. The secondary extraction block includes three convolution layers, a self-attention module and a max-pooling layer arranged in sequence.
9. A deep learning recognition system oriented to dynamic unmanned underwater vehicle echo features, characterized in that: It includes: At least one processor; And The memory is connected in communication with the at least one processor; wherein The memory stores instructions executable by the processor, and the instructions are executed by the processor to implement the deep learning recognition method for dynamic unmanned underwater vehicle echo features in any one of claims 1-8.
10. Use of a deep learning recognition method according to one of claims 1 to 8 for dynamic unmanned underwater vehicle echo features. It is applied to extract multi-domain features in original echo signals and perform complementary and enhancement, and then perform unmanned underwater vehicle recognition.