A Continuous Tracking Method for Underwater Targets Based on Deep Convolutional Neural Network and Adaptive Expert Inference Rules
By combining deep convolutional neural networks and adaptive expert inference rules, the target feature variation problem of water target tracking technology in complex environments is solved, and continuous tracking capabilities are improved under conditions such as multi-objective, strong interference, and low signal-to-noise ratio.
Patent Information
- Application Number
- CN202111336321.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-11
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-11-11
AI Technical Summary
Existing in-water target tracking technologies are difficult to adapt to dynamic changes in the field target-interference-environment under low signal-to-noise ratio, multi-objective and strong interference, resulting in target characteristics variation and loss.
Combining deep convolutional neural networks and adaptive expert inference rules, by constructing a deep convolutional neural network model and expert inference rule library for beam-time spectrogram feature extraction, historical data and field data are used for adaptive updates to achieve continuous tracking of target features.
It improves the adaptability and continuity of water target tracking, enhances the tracking ability under conditions such as multi-objective, strong interference, low signal-to-noise ratio, and can better adapt to on-site environmental changes.
Smart Images

Figure CN114186580B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of underwater target tracking and artificial intelligence, and particularly relates to a method for continuously tracking underwater targets based on a deep convolutional neural network and an adaptive expert inference rule. Background Art
[0002] Passive tracking of underwater targets is an important link in underwater acoustic detection, and can provide important information for target recognition and the formation of underwater attack and defense postures.
[0003] At present, the research foundation of domestic passive underwater target tracking technology is still very weak, mainly relying on energy and azimuth information, and it is easy to lose targets in the case of low signal-to-noise ratio, multiple targets, etc. Although the target tracking method based on target feature identification has received attention, the current feature acquisition method mainly establishes a model based on the known target data / characteristic knowledge. Affected by factors such as multi-target interference, spatio-temporal changes in the ocean channel, platform and environmental noise, it is difficult to obtain clean, clear and practical scene-adaptive target features, which has become a bottleneck restricting the development of target feature tracking technology.
[0004] In recent years, deep learning technology has developed rapidly and has been widely used in the fields of speech, image, etc. It has also received attention from many domestic and foreign scholars in the field of underwater acoustics, but mainly focuses on target recognition and positioning, etc., and there are relatively few research applications in target tracking. Given the complexity of the underwater acoustic environment and the difficulty of obtaining underwater acoustic data, the known underwater acoustic target data / features often have certain differences from the actual application environment. A simple intelligent processing system driven by historical data may be difficult to adapt to the target feature variation caused by the time-varying of the on-site target-interference-environment.
[0005] Therefore, the present invention proposes a method for continuously tracking underwater targets based on a deep convolutional neural network and an adaptive expert inference rule. By comprehensively applying a deep learning model driven by historical data and an expert inference rule driven by on-site data, the entire processing process can better adapt to the on-site situation, which helps to improve the ability of continuously tracking underwater targets. Summary of the Invention
[0006] One of the purposes of the present invention is to provide a method for continuously tracking underwater targets based on a deep convolutional neural network and an adaptive expert inference rule to solve the problems such as the existing intelligent processing system in the background art being difficult to adapt to the target feature variation caused by the time-varying of the on-site target-interference-environment.
[0007] To achieve the above purpose, the present invention provides the following technical solutions:
[0008] A method for continuously tracking underwater targets based on a deep convolutional neural network and an adaptive expert inference rule, the method comprising the following steps:
[0009] Step 1: Construct and train a deep convolutional neural network model for beam time-frequency spectrogram feature extraction;
[0010] Step 2: Construct an expert inference rule base for target feature identification of unknown features, and initialize and update the parameters of the expert inference rule base;
[0011] Step 3: Obtain an initial time-frequency spectrogram sample set of the target of interest;
[0012] Step 4: For each newly acquired batch of multi-beam time-domain data, generate an unknown time-frequency spectrogram sample set based on the short-time Fourier transform;
[0013] Step 5: Use the trained deep convolutional neural network model to extract the features of the target of interest and unknown features from the initial time-frequency spectrogram sample set and the unknown time-frequency spectrogram sample set respectively;
[0014] Step 6: Adaptively update the parameters of the expert inference rule base based on the target feature set of the target of interest;
[0015] Step 7: Based on the target feature set of the target of interest, use the expert inference rule base with adaptively updated parameters to identify the target features of the extracted unknown features.
[0016] Preferably, the method for constructing the deep convolutional neural network model is as follows: successively add an input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a pooling layer, a fourth convolutional layer, and a fifth convolutional layer. The input data size of the input layer is 256*256*1, the convolutional kernel of the first convolutional layer is 5*5*32, and the stride is 2; the convolutional kernel of the second convolutional layer is 3*3*64, and the stride is 1; the convolutional kernels of the third convolutional layer and the fourth convolutional layer are both 3*3*128, and the strides are both 1; the convolutional kernel of the fifth convolutional layer is 1*1*256, and the stride is 2; the pooling kernel of the pooling layer is 3*3*2, and the stride is 1; add 3 successively connected first ResNet-Inception modules; add 1 fourth ResNet-Inception module; add 3 successively connected first ResNet-Inception modules; add 8 successively connected second ResNet-lnception modules; add a sixth convolutional layer, the convolutional kernel of the sixth convolutional layer is 3*3*1024, and the stride is 1; add a seventh convolutional layer, the convolutional kernel of the seventh convolutional layer is 1*1*2048, and the stride is 2; add 5 successively connected third ResNet-Inception modules; add a global average pooling layer; add a fully connected layer, and the output feature dimension is 512; construct the loss function of the deep convolutional neural network model and set the training parameters.
[0017] Preferably, the first ResNet-Inception module is constructed as follows: After the data input layer, four parallel branches are added: Branch 1 is a direct branch, Branch 2 includes two convolutional layers, the parameters of convolutional layer 1 are (1×1, n, 1), and the parameters of convolutional layer 2 are (3×3, n, 1). Branch 3 includes four convolutional layers, and the parameters of each convolutional layer are (1×1, n, 1), (1×5, n, 1), (5×1, n, 1), and (1×1, n, 1) respectively. Branch 4 includes four convolutional layers, and the parameters of each convolutional layer are (1×1, n, 1), (1×7, n, 1), (7×1, n, 1), and (1×1, n, 1) respectively. After Branches 2, 3, and 4, a feature dimension expansion operation is added. By integrating the output features of Branches 2, 3, and 4, a comprehensive feature set output is obtained, and the total number of features is 3n. After the output feature of Branch 1 and the comprehensive feature set, a direct addition operation is added to obtain the top-level feature of the first ResNet-Inception module, and by adding a ReLU activation function, the final convolutional feature is output. The second ResNet-Inception module is constructed as follows: After the data input layer, three parallel branches are added. Branch 1 is a direct branch; Branch 2 includes three convolutional layers, and the parameters of the three convolutional layers are (1×1, n, 1), (1×5, n, 1), and (5×1, n, 1) respectively; Branch 3 includes two parallel sub-branches at the input end. Sub-branch 1 includes three convolutional layers, and the parameters are (1×1, n, 1), (3×3, n, 1), and (3×3, n, 1) respectively. Sub-branch 2 includes three convolutional layers, and the parameters are (1×1, n, 1), (3×1, n, 1), and (1×3, n, 1) respectively. The output ends of Sub-branch 1 and Sub-branch 2 are jointly connected to a convolutional layer with parameters (1×1, n, 1).After branches 2 and 3, a feature dimension expansion operation is added. By integrating the output features of branches 2 and 3, a comprehensive feature set output is obtained, and the total number of features is 2n. After the output features of branch 1 and the comprehensive feature set, a direct addition and summation operation is added to obtain the top-level features of the second ResNet-Inception module, and by adding a ReLU activation function, the final convolutional features are output. The construction of the third ResNet-Inception module is as follows: After the data input layer, 2 parallel branches are added. Branch 1 is a direct branch, and branch 2 contains 2 parallel sub-branches at the input end. Sub-branch 1 includes 1 convolutional layer with parameters (1×1, n, 1), and sub-branch 2 includes 3 convolutional layers with parameters (1×1, n, 1), (1×3, n, 1), and (3×1, n, 1) respectively. The output ends of sub-branch 1 and sub-branch 2 are jointly connected to 1 convolutional layer with parameters (1×1, 2n, 1). After branches 1 and 2, a direct addition and summation operation is added to obtain the top-level features of the third ResNet-Inception module, and by adding a ReLU activation function, the final convolutional features are output. The construction of the fourth ResNet-Inception module is as follows: After the data input layer, 3 parallel branches are added. Branch 1 includes 2 convolutional layers with parameters (1×1, n, 1) and (1×1, n, 2) respectively, branch 2 includes 2 convolutional layers with parameters (1×3, n, 1) and (3×1, n, 2) respectively, and branch 3 includes 3 convolutional layers with parameters (1×1, n, 1), (3×3, n, 1), and (1×1, n, 2) respectively. After branches 1, 2, and 3, a feature dimension expansion operation is added. By integrating the output features of branches 2, 3, and 4, a comprehensive feature set output is obtained, and the total number of features is 3n, and by adding a ReLU activation function, the final convolutional features are output.
[0018] Preferably, the training of the deep convolutional neural network model includes the following steps: Step 1.2.1: Obtain the samples x = {x1(t), x2(t),..., x n (t), (n ∈ N*)} with individual information labels in the underwater target radiated noise signal library; Step 1.2.2: Perform time-frequency transformation preprocessing on the samples x with individual information labels based on the short-time Fourier transform to obtain a time-frequency image sample set, and divide the time-frequency image sample set into independent training sample sets x Train and test sample sets x Test ; Step 1.2.3: Randomly select a reference sample x Train from the training sample set x i , whose corresponding label is a, and the feature calculation result is f(x i ); Then randomly select a sample x j with the label a, and the feature calculation result is f(x j); Randomly select a sample x with a label different from a k , set its label as b, and the feature calculation result as f(x k ); Use the gradient descent algorithm to minimize the loss function J s ,
[0019]
[0020] where δ is a positive number, and iterate and optimize repeatedly to complete the training of the convolutional neural network.
[0021] Preferably, constructing the expert inference rule base includes the following steps: Step 2.1.1: Formulate the matching rule between the unknown feature and the single-template feature. For the unknown feature f(x N ) and the single-template feature f(x R ), use the Euclidean distance and cosine similarity methods to establish a similarity calculation criterion S sig , where μ and λ are weighting coefficients;
[0022]
[0023] Step 2.1.2: Formulate the matching rule between the unknown feature and the multi-template feature. For the unknown feature f(x N ) and the template feature group f(x R,1 ) of the target, f(x R,2 ),..., f(x R,n ), calculate the similarity between the unknown feature and any single-template feature in turn based on Equation (2) to obtain a similarity sequence S sig1 , S sig2 ,..., S sign , and establish a similarity calculation criterion S by comprehensively taking the minimum value and the average value method, where min and avg are the operations of finding the minimum value and the average value respectively, and a and β are weighting coefficients, n > 1,
[0024] S = αmin{S sig1 , S sig2 ,..., S sign} + βavg{S sig1 , S sig2 ,..., S sign} (3)
[0025] Step 2.1.3: Formulate the target identification criterion. By setting a discrimination threshold θ, if S is less than θ, it is determined that the unknown feature matches the template feature.
[0026] Preferably, in step 2, initializing and updating the parameters of the expert inference rule base includes the following steps: Step 2.2.1: Extract features from the training sample set and the test sample set based on the trained deep convolutional neural network model to obtain a training sample feature set and a test sample feature set. Take the training sample feature set as the initial template feature set and the test sample feature set as the initial unknown feature set; Step 2.2.2: Cross-identify the initial template feature set and the initial unknown feature set based on the expert inference rule base to obtain an identification result set R cal ; Step 2.2.3: Based on the identification result set R cal Obtain T cal . When the target template feature and the unknown feature are the same target and the matching result is less than the discrimination threshold, it is set to 1. When the template feature and the unknown feature are different targets and the matching result is greater than the discrimination threshold, it is set to 1, to obtain the set T cal ; Step 2.2.4: Use the genetic algorithm to optimize each weighting coefficient and the discrimination threshold. The objective function is max{T cal}, and the decision variable set is {α, β, μ, λ, θ}; Iteratively optimize the established genetic algorithm model to obtain the best decision variable set.
[0027] Preferably, step 3 is to obtain the target data of interest with a certain duration from the multi-beam tracking time-domain data according to the initial azimuth, generate an initial time-frequency spectrogram sample set based on the short-time Fourier transform after frame division. The total number of samples in the initial time-frequency spectrogram sample set is not less than the lower limit of the total number of samples and not greater than the upper limit N max of the total number of samples. If the total number of samples is greater than N max , then delete the samples from front to back according to the time history until the number of samples meets the requirement of not being greater than N max ; The lower limit of the total number of samples is 10.
[0028] Preferably, in step 7, if the similarity of several unknown features is less than the discrimination threshold, then take the unknown feature with the smallest value as the identification result of the target of interest and track the beam azimuth corresponding to the unknown feature with the smallest value. At the same time, add the corresponding time-frequency spectrogram sample to the initial time-frequency spectrogram sample set; If there is no unknown feature whose similarity is less than the discrimination threshold, then end.
[0029] Preferably, after adding the time-frequency spectrogram sample corresponding to the unknown feature with the smallest value to the initial time-frequency spectrogram sample set, if the total number of samples in the initial time-frequency spectrogram is greater than N max , then delete the samples from front to back according to the time history until the number of samples meets the requirement of not being greater than N max .
[0030] The basic principle of the present invention is as follows: First, aiming at the characteristics of multi-beam time-domain data, starting from the time-frequency domain, time-frequency spectrogram samples are generated. Secondly, a deep convolutional neural network model for time-frequency spectrogram feature extraction is constructed and optimized based on offline data for training. Then, an expert inference rule base is constructed to identify the target of interest from the multi-beam time-frequency spectrogram features. Finally, the on-site multi-beam time-domain data is processed based on the deep convolutional neural network model and the expert inference rule base. On the one hand, according to the known azimuth information of the target of interest, the results of multi-beam feature extraction of the deep convolutional neural network model are comprehensively utilized to initialize and update the expert inference rule base. On the other hand, the expert inference rule base is relied on to identify the new multi-beam features of the deep convolutional neural network model to achieve continuous tracking of the target of interest. During the processing, the expert inference rule base will be adaptively updated according to the identification results of the target of interest / background in each batch to better adapt to the dynamic changes of the target-interference-environment.
[0031] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0032] By comprehensively applying the deep learning model driven by historical data and the expert inference rule driven by on-site data, the present invention can make the whole processing process better adapt to the on-site situation and help improve the continuous tracking ability of underwater targets.
[0033] The present invention comprehensively uses a deep convolutional neural network and an adaptive expert inference rule base to achieve continuous tracking of underwater targets. Compared with the traditional passive target tracking method based on energy and physical characteristics, the present invention has a deeper feature mining level, can effectively enhance the passive target continuous tracking ability under conditions such as multi-targets, strong interference, and low signal-to-noise ratio, has stronger feature comprehensive utilization, and can better adapt to the changes of the on-site environment, and can effectively enhance the target continuous tracking ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 FIG. is the overall principle block diagram for realizing continuous tracking of underwater targets.
[0035] Figure 2 FIG. is the flow chart for continuous tracking of underwater targets.
[0036] Figure 3 FIG. is the structural schematic diagram of 4 ResNet-Inception modules.
[0037] Figure 4 FIG. is the result of continuous target tracking using the method of the present invention.
[0038] Figure 5 FIG. is the construction scheme of the deep convolutional neural network model. DETAILED DESCRIPTION OF THE INVENTION
[0039] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] Referring to Figure 2 As shown, a continuous tracking method for underwater targets based on a deep convolutional neural network and an adaptive expert inference rule specifically includes the following seven steps.
[0041] Step 1: Construct and train a deep convolutional neural network model for beam time-frequency spectrum feature extraction.
[0042] Step 2: Construct an expert inference rule base for identifying target features of unknown features, and initialize and update the parameters of the expert inference rule base;
[0043] Step 3: Obtain an initial time-frequency spectrum sample set of the target of interest.
[0044] Step 4: For each newly obtained batch of multi-beam time-domain data, generate an unknown time-frequency spectrum sample set based on the short-time Fourier transform.
[0045] Step 5: Respectively extract the features of the target of interest and unknown features from the initial time-frequency spectrum sample set and the unknown time-frequency spectrum sample set through the trained deep convolutional neural network model.
[0046] Step 6: Adaptively update the parameters of the expert inference rule base based on the target feature set of the target of interest.
[0047] Step 7: Based on the target feature set of the target of interest, use the expert inference rule base with adaptively updated parameters to respectively identify the target features of each extracted unknown feature.
[0048] In the present invention, the deep convolutional neural network model mainly realizes feature extraction, while the adaptive expert inference rule base mainly realizes the function of identifying target features of the unknown features extracted by the deep convolutional neural network model. The formulation of the expert inference rule base mainly includes the formulation of the matching rule between the unknown feature and the template feature and the formulation of the discrimination threshold.
[0049] Referring to Figure 1, in steps 4-7 of the present invention, time-frequency transformation preprocessing is performed on the multi-beam time-domain data obtained in real time on site to generate spectrogram samples corresponding to each beam. The spectrogram samples are processed based on the trained deep convolutional neural network model to obtain the multi-beam feature extraction results. On the one hand, according to the initial azimuth information of the target of interest, combined with the feature extraction results of the corresponding beam, initial rules are obtained and the expert inference rule base is initialized and updated. On the other hand, relying on the expert inference rule base to identify new multi-beam features, continuous tracking of the target of interest is achieved. During the whole process, the expert inference rule base will be adaptively updated according to the identification results of the target of interest / background in each batch to better adapt to the dynamic changes of the target-interference-environment.
[0050] In step 1, a deep convolutional neural network model for beam spectrogram feature extraction is constructed based on the TensorFlow framework, which mainly includes constructing four basic modules, and constructing the deep convolutional neural network model based on the four constructed basic modules.
[0051] (1.1.1) Construct 4 basic modules, and the specific construction methods are as follows.
[0052] Construct the first ResNet-Inception module. Referring to Figure 3 (a), add 4 parallel branches after the data input layer. Branch 1 is a direct branch without adding any operations. Branch 2 includes 2 convolutional layers. The parameters of convolutional layer 1 are (1×1, n, 1), that is, the convolutional kernel size is 1×1, the number of convolutional kernels is n, which can be specifically set according to the usage requirements, the convolutional stride is 1, and the representation method is the same hereinafter. The parameters of convolutional layer 2 are (3×3, n, 1). Branch 3 includes 4 convolutional layers, and the parameters are (1×1, n, 1), (1×5, n, 1), (5×1, n, 1), and (1×1, n, 1) respectively. Branch 4 includes 4 convolutional layers, and the parameters are (1×1, n, 1), (1×7, n, 1), (7×1, n, 1), and (1×1, n, 1) respectively. After branches 2, 3, and 4, a feature dimension expansion operation is added. By integrating the output features of branches 2, 3, and 4, a comprehensive feature set output is obtained, and the total number of features is 3n. After the output feature of branch 1 and the comprehensive feature set, a direct addition and summation operation is added to obtain the top-level feature of basic module 1, and a ReLU activation function is further added to output the final convolutional feature.
[0053] Construct the second ResNet-Inception module. Referring to Figure 3As shown in (b), three parallel branches are added after the data input layer. Branch 1 is a direct branch without any operations added. Branch 2 includes three convolutional layers with parameters (1×1, n, 1), (1×5, n, 1), and (5×1, n, 1) respectively. Branch 3 contains two parallel sub-branches at the input end. Sub-branch 1 includes three convolutional layers with parameters (1×1, n, 1), (3×3, n, 1), and (3×3, n, 1) respectively. Sub-branch 2 includes three convolutional layers with parameters (1×1, n, 1), (3×1, n, 1), and (1×3, n, 1) respectively. The output ends of Sub-branch 1 and Sub-branch 2 are jointly connected to a convolutional layer with parameters (1×1, n, 1). After Branch 2 and Branch 3, a feature dimension expansion operation is added. By integrating the output features of Branch 2 and Branch 3, a comprehensive feature set output is obtained, and the total number of features is 2n. After the output features of Branch 1 and the comprehensive feature set, a direct addition and summation operation is added to obtain the top-level features of Basic Module 2, and a ReLU activation function is further added to output the final convolutional features.
[0054] Construct the third ResNet-Inception module, referring to Figure 3 As shown in (c), two parallel branches are added after the data input layer. Branch 1 is a direct branch without any operations added. Branch 2 contains two parallel sub-branches at the input end. Sub-branch 1 includes one convolutional layer with parameters (1×1, n, 1). Sub-branch 2 includes three convolutional layers with parameters (1×1, n, 1), (1×3, n, 1), and (3×1, n, 1) respectively. The output ends of Sub-branch 1 and Sub-branch 2 are jointly connected to a convolutional layer with parameters (1×1, 2n, 1). After Branch 1 and Branch 2, a direct addition and summation operation is added to obtain the top-level features of Basic Module 2, and a ReLU activation function is further added to output the final convolutional features.
[0055] Construct the fourth ResNet-Inception module, referring to Figure 3 As shown in (d), three parallel branches are added after the data input layer. Branch 1 includes two convolutional layers with parameters (1×1, n, 1) and (1×1, n, 2) respectively. Branch 2 includes two convolutional layers with parameters (1×3, n, 1) and (3×1, n, 2) respectively. Branch 3 includes three convolutional layers with parameters (1×1, n, 1), (3×3, n, 1), and (1×1, n, 2) respectively. After Branch 1, Branch 2, and Branch 3, a feature dimension expansion operation is added. By integrating the output features of Branch 2, Branch 3, and Branch 4, a comprehensive feature set output is obtained, and the total number of features is 3n. A ReLU activation function is further added to output the final convolutional features.
[0056] In the process of the basic module, the three parameters in the convolution operation are the convolution kernel size, the number of output features, and the stride in sequence. These basic modules all contain multiple parallel branch structures. By configuring the parameters of different convolution operation processes, the adaptability to different scales can be enhanced, thereby improving the ability to insight into data dynamics and the timing of capturing fine features. The activation function used in each convolution layer is set to the ReLU function.
[0057] (1.1.2) Construct the entire convolutional neural network, and its specific construction method is as follows.
[0058] Refer to Figure 5 , add a data input layer with an input data size of 256×256×1; sequentially add convolution-convolution-convolution-pooling-convolution-convolution layers with parameters (5×5, 32, 2), (3×3, 64, 1), (3×3, 128, 1), (3×3, 2), (3×3, 128, 1), (3×3, 128, 1), and (1×1, 256, 2) respectively, where the parameters of the pooling layer are the pooling kernel size and the stride in sequence; add 3 first ResNet-Inception modules; add the fourth ResNet-Inception module; add 3 first ResNet-Inception modules; add 8 second ResNet-Inception modules; add 2 convolution layers with parameters (3×3, 1024, 1) and (1×1, 2048, 2) respectively; add 5 third ResNet-Inception modules; add a global average pooling layer; add a fully connected layer with an output feature dimension of 512.
[0059] Refer to Figure 5 , for the deep convolutional neural network model constructed in step (1.1.2) of the present invention, the size of the input time-frequency spectrogram is 256×256×1, and a sequence with a length of 512 is finally output. The description of the operation parameters is as follows: the convolution parameters are the convolution kernel size, the number of output features, and the stride in sequence; the pooling parameters are the pooling kernel size and the stride in sequence; the parameters of each basic module are the number of module repetitions and the number of output features n in the internal convolution operation.
[0060] (1.1.3) Construct the loss function using the TripletLoss method, and set training parameters such as the optimizer, learning rate, and number of training times during iterative training.
[0061] In step 1, train the constructed deep convolutional neural network model, and the basic process is as follows.
[0062] (1.2.1): Obtain the samples x = {x1(t), x2(t),..., x n (t), (n∈N*)} with individual information labels in the underwater target radiated noise signal library;
[0063] (1.2.2): Perform time-frequency transformation preprocessing on the sample x with individual information labels based on the short-time Fourier transform to obtain a time-frequency image sample set, and divide the time-frequency image sample set into independent training sample sets x Train and test sample sets x Test ;
[0064] (1.2.3): Randomly select a reference sample x Train from the training sample set x i , whose corresponding label is a and the feature calculation result is f(x i ); then randomly select a sample x j with label a, and the feature calculation result is f(x j ); randomly select a sample x k with a label different from a, set its label as b, and the feature calculation result is f(x k ); use the gradient descent algorithm to minimize the loss function J s ,
[0065]
[0066] where δ is a positive number, and iterate and optimize repeatedly to complete the training of the convolutional neural network.
[0067] In step 2, constructing an expert inference rule base for identifying target features from unknown features specifically includes the following 3 sub-steps.
[0068] (2.1.1) Formulate the matching rule between unknown features and single-template features. For unknown feature f(x N ) and single-template feature f(x R ), use the Euclidean distance and cosine similarity methods to establish a similarity calculation criterion S sig , where μ and λ are weighting coefficients;
[0069]
[0070] (2.1.2) Formulate the matching rule between unknown features and multi-template features. For unknown feature f(x N ) and the template feature group f(x R,1 ) of the target, f(x R,2 ),..., f(x R,n ), calculate the similarity between the unknown feature and any single-template feature in turn according to the method in step (2.1.1) to obtain a similarity sequence S sig1 , S sig2 ,..., S sign, a similarity calculation criterion S is established by comprehensively taking the minimum value and the average value method, where min and avg are the operations of finding the minimum value and the average value respectively, α and β are the weighting coefficients respectively, and Ssig can be regarded as a special case of s;
[0071] S = αmin{S sig1 , S sig2 ,..., S sign}+βavg{S sig1 , S sig2 ,..., S sign} (3)
[0072] (2.1.3) Establish a target identification criterion. By setting a discrimination threshold θ, if S is less than θ, it is determined that the unknown feature matches the template feature.
[0073] In step 2, initializing and updating the parameters of the expert inference rule base specifically includes the following 3 sub-steps.
[0074] (2.1.4) Based on the trained deep convolutional neural network model, for the training sample set x Train and the test sample set x Test in step (1.1.2), feature extraction is performed to obtain a training sample feature set and a test sample feature set. The training sample feature set is used as the initial template feature set, and the test sample feature set is used as the initial unknown feature set;
[0075] (2.1.5) Use the methods in steps (2.1.1) and (2.1.2) to perform cross-identification on the initial template feature set and the initial unknown feature set to obtain an identification result set R cal ;
[0076] (2.1.6) According to R cal construct a set T cal . If the target template feature and the unknown feature are the same target and the matching result is less than the discrimination threshold, it is set to 1. If the template feature and the unknown feature are different targets and the matching result is greater than the discrimination threshold, it is set to 1. According to the above results, obtain the set T cal ; Use the genetic algorithm to optimize each weighting coefficient and the discrimination threshold, where the objective function is max{T cal}}, and the decision variable set is {α, β, μ, λ, θ}.
[0077] In step 3, the specific process of obtaining the initial time-frequency spectrogram sample set of the target of interest is as follows:
[0078] Obtain target data of interest with a certain duration from the multi-beam tracking time-domain data according to the initial orientation, and generate an initial time-frequency spectrogram sample set based on the short-time Fourier transform after frame division. The total number in the initial time-frequency spectrogram sample set is not less than the lower limit of the total number of samples (here it is 10), and the total number of samples is not greater than the upper limit N of the total number of samples. max If the total number of samples is greater than N max then delete the samples from front to back according to the time course until the number of samples meets the requirement of not being greater than N. max Take the initial time-frequency spectrogram sample set as the template sample set.
[0079] In step 4, for a newly obtained batch of multi-beam time-domain data, generate unknown time-frequency spectrogram samples corresponding to each beam based on the short-time Fourier transform, and then an unknown time-frequency spectrogram sample set can be obtained.
[0080] In step 5, extract the features of the target of interest from each initial time-frequency spectrogram in the initial time-frequency spectrogram sample set through the trained deep convolutional neural network model, and then an interest target feature set can be obtained. Extract the unknown features from the unknown time-frequency spectrograms in the unknown time-frequency spectrogram sample set through the trained deep convolutional neural network model, and then an unknown feature set can be obtained. The method for the trained deep convolutional neural network to extract unknown features and extract the features of the target of interest is the same.
[0081] In step 6, for the problem of target feature identification of unknown features in the current unknown time-frequency spectrogram sample set, it is necessary to adaptively update the parameters of the expert inference rule base. In step 6 of the present invention, the parameters of the expert inference rule base are adaptively updated based on the interest target feature set. This adaptive update method is the same as the method for initializing and updating the parameters of the expert inference rule base in step 2. The difference is that the initial time-frequency spectrogram sample set is divided into a random template sample set and an unknown sample set according to a certain ratio, and the trained deep convolutional neural network model is used to extract features from the template sample set and the unknown sample set to obtain a template feature set and an unknown feature set. Then, steps (2.1.5) and (2.1.6) are adopted to realize the update of the decision variable set in the expert inference rule base. The ratio division of the template sample set and the unknown sample set can be 1:4, which is common general knowledge in this field.
[0082] Step 7: Based on the interest target feature set, use the expert inference rule base with adaptively updated parameters to identify the target features of the extracted unknown features. If the similarity S of w unknown features is less than the discrimination threshold θ, then take the unknown feature with the smallest similarity value as the identification result of the target of interest and track the beam orientation corresponding to the unknown feature with a small similarity value. At the same time, add the corresponding time-frequency spectrogram sample to the initial time-frequency spectrogram sample set. If there is no unknown feature with a similarity less than the discrimination threshold, then end.
[0083] Step 8: After adding the spectrogram sample corresponding to the unknown feature with the minimum similarity value to the initial spectrogram sample set, determine whether the total number of initial spectrogram samples is greater than N max , if so, delete the samples from front to back according to the time history until the number of samples meets the requirement of not being greater than N max Then perform Step 9, otherwise directly perform Step 9.
[0084] Step 9: Process the next batch of multi-beam time-domain data obtained newly according to Steps 4-7 to obtain the target tracking azimuth corresponding to this next batch of data.
[0085] In Step 9, during the process of target tracking of the next batch of multi-beam time-domain data obtained newly according to Steps 4-7, since during the target tracking process of the previous batch of multi-beam time-domain data, the spectrogram sample corresponding to the unknown feature with the minimum similarity value will be added to the initial spectrogram sample set, and the initial spectrogram sample set is set with an upper limit N for the total number of samples max , so it is necessary to first determine whether the total number of samples in the initial spectrogram sample set is greater than the upper limit of the total number of samples. If so, reduce the scale of the initial spectrogram sample set, that is, when the total number of initial spectrogram samples is greater than N max , delete the samples from front to back according to the time history until the number of samples meets the requirement of not being greater than N max , and then sequentially perform Steps 4-7. Step 4 can be before or after the step of reducing the scale of the initial spectrogram sample set.
[0086] Based on the above method, the present invention processes the continuously obtained multi-beam time-domain data to obtain the continuous tracking result of the target of interest. By continuously updating the expert inference rule base and the template sample set, the adaptability to the continuously changing target-interference-environment is realized.
[0087] For the tracking of a certain target of interest in a certain simulated multi-beam time-domain data, it is processed based on the method proposed by the present invention and the result is compared with the result of the conventional beamforming, as Figure 4 shown. It can be seen that the continuous tracking effect of the target of the method proposed by the present invention is significantly better than the result of the conventional method.
Claims
1. A continuous tracking method for underwater targets based on a deep convolutional neural network and an adaptive expert inference rule, characterized in that The method includes the following steps: Step 1: Construct and train a deep convolutional neural network model for beam time-frequency spectrogram feature extraction; Step 2: Construct an expert inference rule base for target feature identification of unknown features, and initialize and update the parameters of the expert inference rule base; Step 3: Obtain an initial time-frequency spectrogram sample set of the target of interest; Step 4: For each newly obtained batch of multi-beam time-domain data, generate an unknown time-frequency spectrogram sample set based on the short-time Fourier transform; Step 5: Respectively extract the features of the target of interest and unknown features from the initial time-frequency spectrogram sample set and the unknown time-frequency spectrogram sample set through the trained deep convolutional neural network model; Step 6: Adaptively update the parameters of the expert inference rule base based on the target feature set of the target of interest; Step 7: Based on the target feature set of the target of interest, use the expert inference rule base with adaptively updated parameters to identify the target features of the extracted unknown features; In the said Step 1, the method for constructing the deep convolutional neural network model is as follows: successively add an input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a pooling layer, a fourth convolutional layer, and a fifth convolutional layer. The input data size of the input layer is 256*256*1. The convolutional kernel of the first convolutional layer is 5*5*32, and the stride is 2. The convolutional kernel of the second convolutional layer is 3*3*64, and the stride is 1. The convolutional kernels of the third convolutional layer and the fourth convolutional layer are both 3*3*128, and the strides are both 1. The convolutional kernel of the fifth convolutional layer is 1*1*256, and the stride is 2. The pooling kernel of the pooling layer is 3*3*2, and the stride is 1. Add 3 successively connected first ResNet-Inception modules; add 1 fourth ResNet-Inception module; add 3 successively connected first ResNet-Inception modules; add 8 successively connected second ResNet-Inception modules; add a sixth convolutional layer, the convolutional kernel of the sixth convolutional layer is 3*3*1024, and the stride is 1; add a seventh convolutional layer, the convolutional kernel of the seventh convolutional layer is 1*1*2048, and the stride is 2; add 5 successively connected third ResNet-Inception modules; add a global average pooling layer; add a fully connected layer, and the output feature dimension is 512; Construct the loss function of the deep convolutional neural network model and set the training parameters; Among them, the construction of the first ResNet-Inception module is as follows: After the data input layer, 4 parallel branches are added: Branch 1 is a direct branch. Branch 2 includes 2 convolutional layers. The parameters of convolutional layer 1 are (1×1, n, 1), and the parameters of convolutional layer 2 are (3×3, n, 1). Branch 3 includes 4 convolutional layers, and the parameters of each convolutional layer are (1×1, n, 1), (1×5, n, 1), (5×1, n, 1), and (1×1, n, 1) respectively. Branch 4 includes 4 convolutional layers, and the parameters of each convolutional layer are (1×1, n, 1), (1×7, n, 1), (7×1, n, 1), and (1×1, n, 1) respectively. After branches 2, 3, and 4, a feature dimension expansion operation is added. By integrating the output features of branches 2, 3, and 4, a comprehensive feature set output is obtained, and the total number of features is 3n. After the output feature of branch 1 and the comprehensive feature set, a direct addition and summation operation is added to obtain the top-level feature of the first ResNet-Inception module, and by adding a ReLU activation function, the final convolutional feature is output; The construction of the second ResNet-Inception module is as follows: After the data input layer, 3 parallel branches are added. Branch 1 is a direct branch. Branch 2 includes 3 convolutional layers, and the parameters of the 3 convolutional layers are (1×1, n, 1), (1×5, n, 1), and (5×1, n, 1) respectively. Branch 3 includes 2 parallel sub-branches at the input end. Sub-branch 1 includes 3 convolutional layers, and the parameters are (1×1, n, 1), (3×3, n, 1), and (3×3, n, 1) respectively. Sub-branch 2 includes 3 convolutional layers, and the parameters are (1×1, n, 1), (3×1, n, 1), and (1×3, n, 1) respectively. The output ends of sub-branch 1 and sub-branch 2 are commonly connected to 1 convolutional layer, and its parameters are (1×1, n, 1). After branches 2 and 3, a feature dimension expansion operation is added. By integrating the output features of branches 2 and 3, a comprehensive feature set output is obtained, and the total number of features is 2n. After the output feature of branch 1 and the comprehensive feature set, a direct addition and summation operation is added to obtain the top-level feature of the second ResNet-Inception module, and by adding a ReLU activation function, the final convolutional feature is output; The construction of the third ResNet-Inception module is as follows: After the data input layer, 2 parallel branches are added. Branch 1 is a direct branch. Branch 2 includes 2 parallel sub-branches at the input end. Sub-branch 1 includes 1 convolutional layer, and the parameters are (1×1, n, 1). Sub-branch 2 includes 3 convolutional layers, and the parameters are (1×1, n, 1), (1×3, n, 1), and (3×1, n, 1) respectively. The output ends of sub-branch 1 and sub-branch 2 are commonly connected to 1 convolutional layer, and the parameters are (1×1, 2n, 1). After branch 1 and branch 2, a direct addition and summation operation is added to obtain the top-level feature of the third ResNet-Inception module, and by adding a ReLU activation function, the final convolutional feature is output; The construction of the fourth ResNet-Inception module is as follows: After the data input layer, three parallel branches are added. Branch 1 includes two convolutional layers with parameters (1×1, n, 1) and (1×1, n, 2) respectively. Branch 2 includes two convolutional layers with parameters (1×3, n, 1) and (3×1, n, 2) respectively. Branch 3 includes three convolutional layers with parameters (1×1, n, 1), (3×3, n, 1) and (1×1, n, 2) respectively. After branches 1, 2, and 3, a feature dimension expansion operation is added. By integrating the output features of branches 2, 3, and 4, a comprehensive feature set output is obtained, with the total number of features being 3n. And by adding a ReLU activation function, the final convolutional features are output; In step 1, training the deep convolutional neural network model includes the following steps: Step 1.2.1: Obtain the samples \(x = \{x_1(t), x_2(t), \cdots, x_n(t), (n\in N^*)\}\) with individual information tags in the target radiated noise signal library of water; n (t), (n\in N^*)} Step 1.2.2: Perform time-frequency transformation preprocessing on the sample x with individual information tags based on the short-time Fourier transform to obtain a time-frequency image sample set, and divide the time-frequency image sample set into independent training sample sets x Train and test sample sets x Test ; Step 1.2.3: From the training sample set x Train Randomly select a reference sample x from i , its corresponding label is a, and the feature calculation result is f(x i ); then randomly select a sample x with label a j , the characteristic calculation result is f(x j ); Randomly select a sample x with a different label k , set its label to b, and the feature calculation result to f(x k );Use gradient descent algorithm to minimize the loss function J s , where δ is a positive number, and through repeated iterative optimization, the training of the convolutional neural network is completed; In step 2, constructing the expert inference rule base includes the following steps: Step 2.1.1: Establish the matching rules between unknown features and single-template features. For the unknown feature f(x N ) and the single-template feature f(x R ), use the Euclidean distance and cosine similarity methods to establish a similarity calculation criterion S sig , where μ and λ are weighting coefficients; Step 2.1.2: Establish the matching rules between the unknown feature and the multi-template features. For the unknown feature f(x N ) and the target template feature groups f(x R,1 ), f(x R,2 ),..., f(x R,n ), calculate the similarity between the unknown feature and any single-template feature in sequence based on Equation (2) to obtain the similarity sequence S sig1 , S sig2 ,..., S sign . Establish a similarity calculation criterion S by comprehensively taking the minimum value and the average value method, where min and avg are the operations of finding the minimum value and the average value respectively, and a and β are the weighting coefficients respectively, n > 1 S = α min{S sig1 , S sig2 ,..., S sign} + β avg{S sig1 , S sig2 ,..., S sign} (3) Step 2.1.3: Formulate a target identification criterion. By setting a discrimination threshold θ, if S is less than θ, it is determined that the unknown feature matches the template feature; In step 2, initializing and updating the parameters of the expert inference rule base includes the following steps: Step 2.2.1: Based on the trained deep convolutional neural network model, extract features from the training sample set and the test sample set to obtain the training sample feature set and the test sample feature set. Take the training sample feature set as the initial template feature set and the test sample feature set as the initial unknown feature set; Step 2.2.2: Based on the expert inference rule base, cross-identify the initial template feature set and the initial unknown feature set to obtain the identification result set R cal ; Step 2.2.3: Based on the recognition result set R cal Obtain T cal , if the target template feature and the unknown feature are the same target and the matching result is less than the discrimination threshold, it is set to 1; if the template feature and the unknown feature are different targets and the matching result is greater than the discrimination threshold, it is set to 1, to obtain the set T cal ; Step 2.2.4: Optimize each weighting coefficient and discrimination threshold using the genetic algorithm, with the objective function being max{T cal}, and the decision variable set being {α, β, μ, λ, θ}; perform iterative optimization on the established genetic algorithm model to obtain the optimal decision variable set; Step 3 is to obtain target data of interest with a certain duration from the multi-beam tracking time-domain data according to the initial orientation, generate an initial time-frequency spectrogram sample set based on the short-time Fourier transform after frame division, where the total number of samples in the initial time-frequency spectrogram sample set is not less than the lower limit of the total number of samples and not greater than the upper limit N of the total number of samples max , if the total number of samples is greater than N max , then sample deletion is performed from front to back according to the time history until the number of samples meets the requirement of not being greater than N max ; In step 7, if the similarity of several unknown features is less than the discrimination threshold, the unknown feature with the smallest value is taken as the identification result of the target of interest, and the beam azimuth corresponding to the unknown feature with the smallest value is tracked. At the same time, the corresponding time-frequency spectrogram sample is added to the initial time-frequency spectrogram sample set. If there is no unknown feature with a similarity less than the discrimination threshold, it ends.
2. The continuous tracking method for underwater targets based on a deep convolutional neural network and an adaptive expert inference rule according to claim 1, characterized in that In step 7, the lower limit of the total number of samples is 10.
3. A continuous tracking method for underwater targets based on a deep convolutional neural network and adaptive expert inference rules according to claim 2, characterized in that, After adding the spectrogram sample corresponding to the unknown feature with the smallest value to the initial spectrogram sample set, if the total number of initial spectrogram samples is greater than N max , then delete the samples from front to back according to the time history until the number of samples meets the requirement of not being greater than N max .
Citation Information
Patent Citations
Continuous tracking method suitable for target grabbing of underwater robot
CN111105444A
System and method for autonomous joint detection-classification and tracking of acoustic signals of interest
US9869752B1