Underwater target identification method and system based on software and hardware cooperation
The hardware circuit module preprocesses and feature extraction of underwater audio data and combines with neural network models to identify it, which solves the problems of poor real-time performance and high power consumption in the prior art, and achieves efficient underwater target recognition.
Patent Information
- Application Number
- CN202510319291.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-01
AI Technical Summary
In the existing underwater target recognition methods, the real-time performance of signal processing is poor and the power consumption is high. This is mainly because the preprocessing and feature recognition of water audio data are all completed at the software level, resulting in excessive dependence on high-performance processors.
Using a method based on software and hardware collaboration, the hardware circuit module is used to preprocess and feature extraction of water audio data, including analog-to-digital conversion, filtering, sampling and other operations, and the feature data is identified and fused through the neural network model, and the target recognition is completed in combination with hardware and software.
The real-time nature of underwater target recognition is improved and power consumption is reduced. After extracting audio feature data through hardware circuits, it uses the software part to further process it to achieve efficient underwater target recognition.
Smart Images

Figure CN120236188A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of underwater target recognition, and more particularly, to an underwater target recognition method and system based on software and hardware cooperation. Background Art
[0002] The recognition of underwater environmental sound signal sources is an important research direction in the field of underwater acoustic signal processing, and is widely used in fields such as marine exploration, underwater communication, and environmental monitoring. Its core task is to analyze the sound signals collected underwater and classify them into different signal sources, such as marine organisms, ships, environmental noise, etc. It is precisely because of the complex underwater environment that has seriously affected the reliability of underwater communication and the accuracy of underwater target recognition.
[0003] Since underwater audio signals are variable and contain a lot of noise, when studying underwater audio, many algorithms need to be developed to extract the original signals and convert them into available information. This process includes several key stages, and each stage uses a unique set of technologies. The general processing method is to use a computer, an embedded system, or a digital signal processing (DSP) program to process the audio signals, such as preprocessing such as filtering, noise reduction, and echo cancellation, and then combining machine learning algorithms to complete the classification and recognition of the signal sources.
[0004] Traditional audio signal preprocessing operations are mostly completed at the software level. Existing models mostly use a two-stage model at the software level. The first stage is to extract effective feature information from underwater acoustic signals, and the second stage is to use a classifier to identify the target category.
[0005] Among them, in the existing methods, the preprocessing and feature recognition of underwater audio data in the first stage, and the recognition of underwater targets in the second stage are all processed by software. In this way, there is a problem of poor real-time signal conversion, and at the same time, complex calculations are required by the CPU / GPU, resulting in high energy consumption. The more complex the calculation, the greater the dependence on high-performance processors. Summary of the Invention
[0006] Aiming at the defects of poor real-time performance and high power consumption in the existing underwater target recognition, the present invention provides an underwater target recognition method and system based on software and hardware cooperation.
[0007] According to a first aspect of the present invention, there is provided an underwater target method based on software and hardware cooperation, including:
[0008] Obtain underwater audio data, preprocess the underwater audio data based on a hardware circuit module, and extract various audio feature data in the underwater audio data. The various audio feature data include LOFAR spectrum images, Mel frequency cepstrum coefficient spectrum images, continuous wavelet transform spectrum images, and DEMON spectrum images;
[0009] Input each of the extracted audio feature data into a neural network model, and output the underwater target recognition result corresponding to each of the audio feature data;
[0010] Perform weighted fusion on the underwater target recognition results corresponding to each of the audio feature data to obtain the final underwater target recognition result.
[0011] Based on the above technical solutions, the present invention can also be improved as follows.
[0012] Optionally, the preprocessing of the underwater audio data and the extraction of various audio feature data from the underwater audio data based on the hardware circuit module include:
[0013] Preprocess the underwater audio data based on the hardware circuit module, and the preprocessing includes analog-to-digital conversion, signal amplification, filtering, sampling, quantization, encoding, and anti-aliasing filtering of the underwater audio data;
[0014] Extract various audio feature data based on the preprocessed underwater audio data.
[0015] Optionally, the hardware circuit module includes a memory and a preprocessing module, the preprocessing module includes a data reading module, a configuration module, and a calculation unit, and both the data reading module and the configuration module communicate with the memory through a bus;
[0016] The configuration module is used to read configuration information from the memory through the APB bus, and configure the calculation unit based on the configuration information, and the configuration information includes the number of calculation points, calculation mode, data loading start address, data output address, and number of calculation rounds;
[0017] The data reading module is used to read the data to be calculated from the memory through the AXI bus according to the data loading start address, and load the data to be calculated into the calculation unit;
[0018] The calculation unit is used to preprocess the data to be calculated and extract audio feature data, and store the extracted audio feature data in the corresponding storage address of the memory according to the data output address to complete one calculation; and complete multiple calculations according to the number of calculation rounds.
[0019] Optionally, the inputting each of the extracted audio feature data into a neural network model and outputting the underwater target recognition result corresponding to each of the audio feature data includes:
[0020] For any one of the spectrogram images, extract global features of multiple scales and local features of multiple levels of the spectrogram image based on the neural network model;
[0021] Fuse the global features at multiple levels and the local features at multiple levels to obtain fused features;
[0022] Based on the fused features, identify underwater targets.
[0023] Optionally, the neural network model includes a convolutional kernel, a global branch network, a local branch network, a fusion network, and a classifier network. The global branch network includes multiple globally connected feature extraction modules in series, the local branch network includes multiple locally connected feature extraction modules in series, and the fusion network includes multiple feature fusion modules in series; wherein, the number of the global feature extraction modules, the number of the local feature extraction modules, and the number of the feature fusion modules are all equal;
[0024] Divide the spectral image into multiple sub-image blocks through the convolutional kernel, and input the multiple sub-image blocks into the global branch network and the local branch network respectively;
[0025] Extract multiple levels of global features of each sub-image block through the multiple global feature extraction modules in the global branch network, and extract multiple levels of local features of each sub-image block through the multiple local feature extraction modules in the local branch network;
[0026] Fuse the global features at multiple levels and the local features at multiple levels through the multiple feature fusion modules to obtain the fused features of each sub-image block;
[0027] Obtain the fused features of the spectral image according to the fused features of each sub-image block;
[0028] Perform underwater target recognition on the fused features through the classifier network.
[0029] Optionally, the classifier network is a four-channel classifier; performing underwater target recognition on the fused features through the classifier network includes:
[0030] Input the fused features into the four-channel classifier, and output four target types and the prediction probability of each target type.
[0031] Optionally, the weighted fusion of the underwater target recognition results corresponding to each audio feature data to obtain the final underwater target recognition result includes:
[0032] Perform weighted summation on the probabilities of the same target type recognized from the four spectral images to obtain the prediction probability of each target type;
[0033] Take the target type with the highest predicted probability as the final underwater target recognition result.
[0034] According to the second aspect of the present invention, there is provided an underwater target recognition system based on software and hardware cooperation, including a hardware circuit module and a software module, and a neural network model is deployed in the software module;
[0035] The hardware circuit module is used to obtain underwater audio data, preprocess the underwater audio data and extract various audio feature data from the underwater audio data, and the audio feature data includes LOFAR spectrum images, Mel frequency cepstrum coefficient spectrum images, continuous wavelet transform spectrum images, and DEMON spectrum images;
[0036] The neural network model is used to input each of the extracted audio feature data into the neural network model of the software module, output the underwater target recognition result corresponding to each of the audio feature data; and perform weighted fusion on the underwater target recognition results corresponding to each of the audio feature data to obtain the final underwater target recognition result.
[0037] The underwater target recognition method and system based on software and hardware cooperation provided by the present invention use the hardware circuit to preprocess the underwater audio data and extract the audio feature data, and only use the software to process the target recognition part. Compared with the existing method of using software to process all processing stages, the real-time performance of underwater target recognition is high, and the power consumption of the software module is low. Description of the Drawings
[0038] Figure 1 It is a flowchart of an underwater target recognition method based on software and hardware cooperation provided by an embodiment of the present invention;
[0039] Figure 2 It is a schematic framework diagram of underwater target recognition based on software and hardware cooperation according to an embodiment of the present invention;
[0040] Figure 3 It is a schematic structural diagram of the hardware circuit module provided by an embodiment of the present invention;
[0041] Figure 4 It is a schematic calculation process diagram of the hardware circuit part provided by an embodiment of the present invention;
[0042] Figure 5 It is a schematic diagram of the multi-level fusion network structure according to an embodiment of the present invention;
[0043] Figure 6 It is a schematic diagram of weighted prediction of various target recognition results;
[0044] Figure 7Schematic diagram of a structure of an underwater target recognition system based on software and hardware cooperation provided by an embodiment of the present invention. Detailed implementation manners
[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. In addition, the technical features in each embodiment or a single embodiment provided by the present invention can be combined with each other arbitrarily to form a feasible technical solution. Such combination is not restricted by the order of steps and / or the pattern of structural composition, but must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0046] Based on the fact that all processing stages of underwater audio data in the prior art are implemented by software, the present invention uses an underwater target recognition method. After the audio signal preprocessing and audio feature data extraction are completed on a dedicated hardware circuit, the software part is used to perform the next processing on the collected and transformed feature data to complete the classification and recognition of underwater targets. Combining the hardware and software parts, the present invention proposes a method for underwater audio signal processing and target recognition based on software and hardware cooperation.
[0047] Figure 1 Flowchart of a method for underwater target recognition based on software and hardware cooperation provided by the present invention, as Figure 1 and Figure 2 shown. The method includes the following steps:
[0048] Step 1, obtain underwater audio data, perform preprocessing on the underwater audio data based on a hardware circuit module, and extract various audio feature data in the underwater audio data. The audio feature data includes LOFAR spectrum images, Mel frequency cepstrum coefficient spectrum images, continuous wavelet transform spectrum images, and DEMON spectrum images.
[0049] It can be understood that the embodiments of the present invention perform preprocessing on underwater audio data and extraction of audio feature data based on a hardware circuit module. Among them, performing preprocessing on the underwater audio data based on the hardware circuit module mainly includes performing analog-to-digital conversion, signal amplification, filtering, sampling, quantization, encoding, and anti-aliasing filtering on the underwater audio data.
[0050] Specifically, first, underwater audio signals are collected. The collected audio signals are analog signals, which are not conducive to subsequent digital processing. Therefore, it is necessary to perform data conversion on the sampled signals to convert the analog signals into binary data.
[0051] According to the characteristics of underwater target recognition, four spectral images, namely, the LOFAR (Low-Frequency Analysis and Recording) spectrum, Mel-Frequency Cepstral Coefficients (MFCC), Continuous Wavelet Transform (CWT) spectrum, and DEMON (Demodulation of Envelope Modulation on Noise) spectrum, are selectively extracted from the digital underwater audio data.
[0052] The hardware circuit part of the embodiment of the present invention adopts a dedicated audio processing hardware circuit, and its circuit structure can be referred to Figure 3 , in some embodiments of the present invention, the hardware circuit module includes a memory and a preprocessing module. The preprocessing module includes a data reading module, a configuration module, and a calculation unit. The data reading module and the configuration module are both communicatively connected to the memory through a bus.
[0053] The configuration module is configured to read configuration information from the memory through the APB bus and configure the calculation unit based on the configuration information. The configuration information includes the number of calculation points, calculation mode, data loading start address, data output address, and number of calculation rounds;
[0054] The data reading module is configured to read the data to be calculated from the memory through the AXI bus according to the data loading start address and load the data to be calculated into the calculation unit;
[0055] The calculation unit is configured to preprocess the data to be calculated and extract audio feature data, and store the preprocessed audio feature data in the corresponding storage address of the memory according to the data output address to complete one calculation; and complete multiple calculations according to the number of calculation rounds.
[0056] Among them, specifically, for the data reading part of the hardware circuit, it mainly loads the collected digital audio signals into the preprocessing module for subsequent calculations; the mode configuration part reads the current mode configuration information through the APB bus, such as the number of calculation points, the current calculation mode, the starting address of data loading, the data output address, and the number of calculation rounds, etc.; the calculation unit undertakes the core calculations, such as FFT, complex number operations, etc. Through the calculation module, continuous wavelet transform, LOFAR spectrum, MFCC spectrum, and DEMON spectrum calculations can be completed; the address management part undertakes the tasks of saving and converting the intermediate calculation data addresses to ensure that the data required for each step in the calculation process can be correctly sent to the corresponding calculation unit. The data output part stores the calculated data according to the saving address given by the configuration. The specific calculation flowchart of this hardware part is as Figure 4 shown, and mainly includes the following steps:
[0057] 401: Start the preprocessing module to ensure that the operating environment supports the module to run calculations.
[0058] 402: After starting the module, complete the basic information configuration of the module execution through the pre-set configuration information file. The specific steps are as shown in 403 to 406.
[0059] 403 - 406: Complete the configuration of data input, output data address, number of calculation points, number of loop calculations, and calculation mode by means of register assignment. The setting options for the number of calculation points include 128 and 256 points, with the default being 256 points. The value of the number of loops is default set to 2. The calculation mode includes the following LOFAR, MFCC, CWT, and DEMON options. Select different modes and start the module to calculate the corresponding feature data.
[0060] 407: When starting the module, it is necessary to reset the module to the waiting calculation state, complete all state machine jumps to the initial state, clear the counter, etc.
[0061] 408: After the previous step is completed, obtain the starting address of the data to be calculated through the APB bus.
[0062] 409: Subsequently, complete the data transmission and transfer through the AXI bus and load the data to be calculated.
[0063] 410: The internal state machine of the module jumps from the initial state to the data reading state, jumps to the calculation state after the data is loaded, and starts the corresponding calculation unit. Confirm the number of calculation points, the number of loop calculations, and the calculation mode through the read configuration information, and start calculating the feature data.
[0064] 411: Wait for the calculation result to be completed and output the calculation result.
[0065] 412: Obtain the starting address of the pre-set output result.
[0066] 413: Transfer and save the calculation settlement result through the AXI bus, and at the same time, the calculation ends.
[0067] 414: Pull up the end signal after the calculation result is completed to indicate that the current calculation has been completed.
[0068] 415: Print the end information and corresponding configuration information, such as the calculation mode and the number of calculation points.
[0069] 416: A single calculation is completed. If you want to continue the calculation, you can complete the calculation of the next mode through a soft reset, or repeat step 401.
[0070] Step 2: Input each of the extracted audio feature data into the neural network model, and output the underwater target recognition result corresponding to each of the audio feature data.
[0071] It can be understood that in step 1, the audio feature data is extracted from the original underwater audio data through the hardware circuit module, that is, four spectral images are extracted. In this step, the neural network model is used to perform feature extraction and underwater target recognition on the four spectral images.
[0072] In an embodiment of the present invention, the inputting each of the extracted audio feature data into the neural network model and outputting the underwater target recognition result corresponding to each of the audio feature data includes: for any one of the spectral images, extracting global features of multiple scales and local features of multiple levels of the spectral image based on the neural network model; fusing the global features of multiple levels and the local features of multiple levels to obtain a fused feature; and recognizing the underwater target based on the fused feature.
[0073] Among them, in some embodiments of the present invention, the neural network model includes a convolution kernel, a global branch network, a local branch network, a fusion network, and a classifier network. The global branch network includes multiple serially connected global feature extraction modules, the local branch network includes multiple serially connected local feature extraction modules, and the fusion network includes multiple serially connected feature fusion modules; wherein, the number of the global feature extraction modules, the number of the local feature extraction modules, and the number of the feature fusion modules are all equal;
[0074] The spectral image is divided into multiple sub-image blocks through the convolution kernel, and the multiple sub-image blocks are respectively input into the global branch network and the local branch network;
[0075] Extracting multiple levels of global features of each sub-image block through multiple global feature extraction modules in the global branch network, and extracting multiple levels of local features of each sub-image block through multiple local feature extraction modules in the local branch network;
[0076] Fusing the global features of multiple levels and the local features of multiple levels through multiple feature fusion modules to obtain the fused features of each sub-image block;
[0077] Obtaining the fused features of the spectral image according to the fused features of each sub-image block;
[0078] Performing underwater target recognition on the fused features through the classifier network.
[0079] Specifically, after extracting the feature data of the underwater audio data through the hardware circuit module, the multi-scale global and local features of the audio feature data are obtained by combining with the multi-level fusion network structure. In the embodiment of the present invention, new global and local feature blocks are introduced in the processing of the underwater target recognition scenario, such as Figure 5 As shown, the global features and local features of each type of audio image are extracted in parallel, and the global and local representations of different semantic scales are fused through a multi-level fusion (MLF, Multi-level Fusion) structure, which makes it more advantageous than classifying by simply extracting feature data.
[0080] The working process of the neural network model is as follows: First, the original spectral image of 224×224×3 is converted into an appropriate size, and the original image is divided into 4×4 sub-images through a convolutional kernel with a stride of 4 and 4×4. The H×W of each sub-image is 56×56, and then they are respectively input into the globally branched network after linear embedding and the locally branched network after layer normalization. At the same time, in order to improve the accuracy of the image classification model and fuse the local features and global features of different levels, a parallel network structure for hierarchical feature fusion is introduced, and the overall structure is as Figure 5 shown.
[0081] Among them, the local branch network is used to extract the local features of the image, and 3×3 depth convolution is used, and the number of groups is equal to the number of channels. Then, cross-channel information interaction is realized through a linear layer. Finally, the extracted local features are input into the MLF block.
[0082] The global branch network is used to extract the global semantic representation of images. A Windows Multi-head Self-Attention (WMSA) module is introduced in the global feature extraction branch, which can divide the feature map into a specified size of M×M. At each stage, by incorporating the Patching operation into the global feature block, the feature map passes through layer normalization and enters the WNSA module, and then passes through a linear layer with a GELU activation function. A residual connection is applied after each module.
[0083] The global branch is downsampled through Patch Merging and then input into the global feature block for feature transformation. The local branch is downsampled through a convolution with a stride of 2 and a kernel size of 2. The advantage of this parallel structure is that it can maximally retain local and global features without interfering with each other.
[0084] The MLF block is used to fuse the local and global features at each stage and connect the output of the previous stage. It includes spatial attention, channel attention, an inverse residual MLP, and a shortcut for adaptively fusing the semantic information between various scale features of each branch. The adaptive hierarchical feature fusion block can adaptively fuse local features, global features, and the semantic information after fusion of the previous layer according to the input features. The formula is as follows:
[0085]
[0086]
[0087] Among them, G i represents the feature matrix output by the global feature block, L i represents the feature matrix output by the local feature block, F i-1 represents the feature matrix output by the MLF of the previous stage, F i represents the feature matrix generated by the fusion of the MLF at this stage. The output of each level of local features obtained through spatial attention is combined with the output of each level of global features obtained through channel attention and input into the IRMLP module.
[0088] The fusion features of each extracted spectral image are classified using a global average pooling and layer normalization linear classifier. Among them, the classifier network is a four-channel classifier; underwater target recognition is performed on the fusion features through the classifier network, including:
[0089] The fusion features of each spectral image are input into the four-channel classifier, and four target types and the prediction probability of each target type are output.
[0090] Step 3: According to the underwater target recognition results corresponding to each type of the audio feature data, perform weighted fusion to obtain the final underwater target recognition result.
[0091] In some embodiments of the present invention, the performing weighted fusion according to the underwater target recognition results corresponding to each type of the audio feature data to obtain the final underwater target recognition result includes:
[0092] Perform weighted summation on the probabilities of the same target type recognized from the four spectral images to obtain the prediction probability of each target type; and take the target type with the maximum prediction probability as the final underwater target recognition result.
[0093] Specifically, after the MFCC spectrum, LOFAR spectrum, DEMON spectrum, and CWT spectrum feature vectors pass through the classifier, the probability of each predicted type is obtained, and a direct average weighting method is adopted for the four spectral prediction results to make a weighted decision on the recognition results. As Figure 6 shown, this method defaults to assign weights of 1:1:1:1 to the above four results, and then multiplies the model prediction results by the corresponding weights and sums them to obtain the final prediction result. The weighted decision model is as follows:
[0094]
[0095]
[0096] where F is the final prediction result of the weighted decision model; N is the number of channels, and in the embodiments of the present invention, N is 4; w i is the weighted weight of the i-th channel; F i is the prediction result of the classifier of the i-th channel.
[0097] See Figure 7 , which is an underwater target recognition system based on software and hardware cooperation provided by the embodiments of the present invention, including a hardware circuit module and a software module, and a neural network model is deployed in the software module;
[0098] The hardware circuit module is used to acquire underwater audio data, preprocess the underwater audio data, and extract multiple audio feature data from the underwater audio data, where the audio feature data includes LOFAR spectrum images, Mel frequency cepstrum coefficient spectrum images, continuous wavelet transform spectrum images, and DEMON spectrum images;
[0099] The neural network model is used to input each type of the extracted audio feature data into the neural network model of the software module, output the underwater target recognition result corresponding to each type of the audio feature data; and perform weighted fusion on the underwater target recognition results corresponding to each type of the audio feature data to obtain the final underwater target recognition result.
[0100] The underwater target recognition method and system based on software and hardware cooperation provided by the present invention have the following beneficial effects:
[0101] (1) When processing audio signals, the hardware circuit adopted by the present invention to process audio data has the advantages of low delay, unrestricted hardware performance, and low power consumption.
[0102] (2) In the target recognition network, the main model is based on ordinary convolution to extract the features of underwater acoustic signals. However, due to the mixing of ocean environmental noise and partial information of underwater acoustic target features, ordinary convolution operations are extremely likely to lose some effective features of underwater acoustic targets and wrongly retain ocean environmental noise, thus reducing the ability of the underwater acoustic target recognition model to extract effective features. The multi-level feature fusion network introduced by the present invention for underwater target recognition scenarios can fuse the advantages of Transformer and CNN from multiple scales without destroying their respective modeling, thereby improving the classification accuracy of various spectral images. The parallel hierarchical structure of local and global feature blocks can effectively extract local features and global features at different semantic scales, and has the flexibility to model at different scales.
[0103] (3) Most studies use a certain spectral diagram to extract the features of underwater acoustic signals, and the obtained feature information is relatively single. They do not consider the differences between different feature information and the defects brought by single feature information in the face of complex underwater environments. The present invention combines the method of weighted decision-making, multiplies the prediction results of four classifiers obtained by corresponding weights and sums them to obtain the final prediction result, reduces the error brought by a single feature information source, excavates the complementary information of audio features, and effectively improves the recognition accuracy.
[0104] It should be noted that in the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0105] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0106] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded computers, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more flows Figure 1 or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0107] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more flows Figure 1 or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows Figure 1 or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0109] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.
[0110] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. An underwater target recognition method based on software and hardware collaboration, characterized in that: include: Acquire underwater acoustic audio data, pre-process the underwater acoustic audio data based on a hardware circuit module and extract multiple audio feature data from the underwater acoustic audio data, wherein the multiple audio feature data include LOFAR spectrum image, Mel frequency cepstral coefficient spectrum image, continuous wavelet transform spectrum image and DEMON spectrum image; Inputting each of the extracted audio feature data into a neural network model, and outputting an underwater target recognition result corresponding to each of the audio feature data; The underwater target recognition results corresponding to each type of audio feature data are weighted and fused to obtain a final underwater target recognition result.
2. The underwater target recognition method based on software and hardware collaboration according to claim 1 is characterized in that: The hardware circuit module is used to preprocess the underwater acoustic audio data and extract a variety of audio feature data from the underwater acoustic audio data, including: Preprocessing the underwater acoustic audio data based on a hardware circuit module, wherein the preprocessing includes analog-to-digital conversion, signal amplification, filtering, sampling, quantization, encoding, and anti-aliasing filtering of the underwater acoustic audio data; Based on the pre-processed underwater audio data, a variety of audio feature data are extracted.
3. The underwater target recognition method based on software and hardware collaboration according to claim 1 is characterized in that: The hardware circuit module includes a memory and a preprocessing module, the preprocessing module includes a data reading module, a configuration module and a calculation unit, and the data reading module and the configuration module both communicate with the memory through a bus; The configuration module is used to read configuration information from the memory through the APB bus, and configure the computing unit based on the configuration information, wherein the configuration information includes calculation points, calculation mode, data loading first address, data output address and calculation rounds; The data reading module is used to read the data to be calculated from the memory through the AXI bus according to the data loading first address, and load the data to be calculated into the calculation unit; The calculation unit is used to preprocess the calculation data and extract the audio feature data, and store the extracted audio feature data to the corresponding storage address of the memory according to the data output address to complete one calculation; and complete multiple calculations according to the calculation rounds.
4. The underwater target recognition method based on software and hardware collaboration according to claim 1 is characterized in that: The step of inputting each of the extracted audio feature data into a neural network model and outputting an underwater target recognition result corresponding to each of the audio feature data comprises: For any spectrum image, extracting global features of multiple scales and local features of multiple levels of the spectrum image based on a neural network model; Fusing the global features at multiple levels and the local features at multiple levels to obtain fused features; Based on the fusion features, underwater targets are identified.
5. The underwater target recognition method based on software and hardware collaboration according to claim 4 is characterized in that: The neural network model includes a convolution kernel, a global branch network, a local branch network, a fusion network and a classifier network, the global branch network includes a plurality of global feature extraction modules connected in series, the local branch network includes a plurality of local feature extraction modules connected in series, and the fusion network includes a plurality of feature fusion modules connected in series; wherein the number of the global feature extraction modules, the number of the local feature extraction modules and the number of the feature fusion modules are all equal; Dividing the spectrum image into a plurality of sub-image blocks by using the convolution kernel, and inputting the plurality of sub-image blocks into the global branch network and the local branch network respectively; Extracting multiple levels of global features of each of the sub-image blocks through the multiple global feature extraction modules in the global branch network, and extracting multiple levels of local features of each of the sub-image blocks through the multiple local feature extraction modules in the local branch network; fusing the global features at multiple levels and the local features at multiple levels through the multiple feature fusion modules to obtain the fusion features of each of the sub-image blocks; Obtaining a fusion feature of the spectrum image according to the fusion feature of each of the sub-image blocks; The fused features are used to identify underwater targets through the classifier network.
6. The underwater target recognition method based on software and hardware collaboration according to claim 5 is characterized in that: The classifier network is a four-channel classifier; Performing underwater target recognition on the fusion features through the classifier network includes: The fused features are input into a four-channel classifier, which outputs four target types and the predicted probability of each target type.
7. The underwater target recognition method based on software and hardware collaboration according to claim 6 is characterized in that: The weighted fusion of the underwater target recognition results corresponding to each type of the audio feature data to obtain the final underwater target recognition result includes: Performing a weighted summation on the probabilities of the same target type identified according to the four spectrum images to obtain a predicted probability of each target type; The target type with the highest predicted probability is taken as the final underwater target recognition result.
8. An underwater target recognition system based on software and hardware collaboration, characterized in that: It includes a hardware circuit module and a software module, wherein a neural network model is deployed in the software module; The hardware circuit module is used to obtain underwater acoustic audio data, pre-process the underwater acoustic audio data and extract multiple audio feature data from the underwater acoustic audio data, wherein the audio feature data includes LOFAR spectrum image, Mel frequency cepstral coefficient spectrum image, continuous wavelet transform spectrum image and DEMON spectrum image; The neural network model is used to input each of the extracted audio feature data into the neural network model of the software module, output the underwater target recognition result corresponding to each of the audio feature data; and weightedly fuse the underwater target recognition results corresponding to each of the audio feature data to obtain the final underwater target recognition result.
Citation Information
Cited By
Miniaturized underwater target real-time monitoring system based on software and hardware collaborative optimization and data interaction
CN120932678A