A Smart Classification Method for Surface and Underwater Targets Based on Grazing Angle-Range Domain Interferometry

By constructing a deep convolutional neural network model and utilizing deep-sea interference structures to automatically learn target depth features, the robustness and accuracy issues of surface and underwater target classification in existing technologies have been resolved, achieving end-to-end intelligent classification.

CN122090249APending Publication Date: 2026-05-26THE 715TH RES INST OF CHINA SHIPBUILDING IND CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE 715TH RES INST OF CHINA SHIPBUILDING IND CORP
Filing Date
2026-01-15
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies for classifying surface and underwater targets in complex deep-sea environments rely on human experience, feature extraction is cumbersome and lacks robustness, making it difficult to achieve high-precision end-to-end classification.

Method used

A deep convolutional neural network model is constructed, which utilizes the interferometric structure received by a vertical array in a deep-sea environment. Through multi-scale cross-stage feature fusion, the nonlinear mapping relationship between the interferometric structure and the target category is automatically learned, thereby achieving end-to-end intelligent classification.

Benefits of technology

The process was simplified, deep features were automatically mined, the accuracy of classification and environmental adaptability were improved, and high-precision classification of surface and underwater targets was achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090249A_ABST
    Figure CN122090249A_ABST
Patent Text Reader

Abstract

This invention relates to the field of underwater acoustic signal processing and target recognition technology, and in particular to an intelligent classification method for surface and underwater targets based on a grazing angle-range domain interferometric structure. The method first constructs a standardized sample set of vertical time-recorded maps containing depth features of both surface and underwater targets through acoustic field simulation. Then, it designs and trains a multi-scale, cross-stage feature fusion classification model. This model uses multi-scale convolution operators to extract features from different receptive fields and integrates shallow and deep-layer features through a cross-stage fusion module, achieving deep mining of depth information in the interferometric structure. Finally, the trained model is used to classify unknown target data. This invention avoids the limitations of manual feature design, realizes end-to-end mapping from raw data to category decisions, and improves the accuracy and robustness of surface and underwater target classification in complex deep-sea environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of underwater acoustic signal processing and target recognition technology, specifically to an intelligent classification method for surface and underwater targets based on a grazing angle-range domain interference structure. Background Technology

[0002] Accurate classification of surface and underwater targets is one of the core technologies in underwater acoustic detection, marine surveillance, and early warning systems. Traditional methods for classifying surface and underwater targets mainly rely on the analysis of the target's radiated noise characteristics, such as spectral features, modulation spectrum features, and line spectrum components. However, in real marine environments, especially under complex deep-sea channel conditions, the radiated noise characteristics of targets are easily affected by propagation attenuation, multipath effects, and environmental noise, posing challenges to the stability and robustness of acoustic signature-based classification methods.

[0003] Besides the difference in radiated noise, surface ships and underwater vehicles differ fundamentally in their operating depth. Typical surface ships typically operate at shallower depths, generally less than 20 meters, while submarines and other underwater targets usually operate at greater depths, often between 100 and 200 meters or even deeper. Therefore, target depth information can serve as a key characteristic for distinguishing their category. In the deep-sea environment, both surface and underwater targets can be considered near-surface sound sources relative to the deep-sea acoustic channel. During sound wave propagation, interference occurs between the waves reflected from the sea surface and the direct waves, forming an interference structure with alternating bright and dark fringes. This interference structure is extremely sensitive to the depth of the sound source, and its fringe pattern exhibits regular changes with factors such as the sound source depth, the location of the receiving array, and the sea depth.

[0004] Based on the above principles, existing technologies have developed methods for estimating or classifying target depth using interference structures. One type of method is based on feature matching, which involves pre-establishing a template library of interference structures corresponding to sound sources at different depths through simulation or experimentation, and then matching the measured data with the templates to estimate the target depth or category. Another type of method is based on feature extraction, which involves manually designing a feature extractor to extract depth-related features from the interference structure, and then using a traditional classifier for classification. However, the above methods also have some drawbacks, such as: high dependence on human experience and prior knowledge, cumbersome feature extraction process, difficulty in ensuring the optimality of constructed features and limited generalization ability; the use of a segmented processing flow, separating feature extraction from classification decision, and the non-end-to-end mode is prone to information loss and makes it difficult to guarantee global optimality; poor adaptability to complex and changing environments, and insufficient robustness of models based on fixed templates or simple features when the interference structure is distorted, which can easily lead to a decrease in classification accuracy.

[0005] In summary, there is an urgent need for a method that can automatically and deeply mine the target depth features contained in deep-sea interference structures and achieve end-to-end, high-precision intelligent classification of surface and underwater targets. Summary of the Invention

[0006] To address the shortcomings of existing technologies, the present invention aims to provide an intelligent classification method for surface and underwater targets based on grazing angle-range domain interferometric structures. This method utilizes the interferometric structures contained in the vertical time record diagram formed by receiving target radiation signals from a vertical array in a deep-sea environment to construct a deep convolutional neural network model. The model autonomously learns the complex nonlinear mapping relationship between the interferometric structure and the target category (surface / underwater), thereby achieving end-to-end intelligent classification.

[0007] To achieve the above objectives, the present invention provides the following technical solution: an intelligent classification method for surface and underwater targets based on a grazing angle-range domain interferometric structure, comprising the following steps:

[0008] S100. Construct a standardized training sample set. Based on the set deep-sea environment parameters, receiving array parameters, and depth parameters of surface and underwater targets, simulate the movement of the target in the horizontal plane through sound field simulation, generate a vertical time record map containing target depth category labels, and segment and standardize the vertical time record map to obtain a sample set of fixed size.

[0009] S200. Construct a multi-scale cross-stage feature fusion classification model. Construct a deep convolutional neural network comprising an initial convolutional layer, at least two feature extraction stages, and a classification layer connected sequentially. At least one feature extraction stage includes multiple cascaded multi-scale basic modules, each with multiple convolutional branches operating in parallel, and the kernel sizes of these branches are different. At least one feature extraction stage is connected to at least one previous stage via a cross-stage feature fusion module, which fuses feature maps from different depth stages.

[0010] S300. Train the multi-scale cross-stage feature fusion classification model by using the standardized training sample set to perform supervised training on the multi-scale cross-stage feature fusion classification model to obtain a trained classification model.

[0011] S400. Classify the target data to be identified, obtain the vertical time record data of the target to be identified and perform the same standardization process as in step S100, input the trained classification model, and obtain the classification result of the water surface or underwater.

[0012] Preferably, step S100 includes:

[0013] S110. Set simulation parameters: Set sea depth, sound velocity profile, and seabed parameters; Set the receiving array to a vertical uniform linear array, and set its number of array elements and element spacing; Set the reference depth z0 for the surface target and the reference depth z1 for the underwater target; Set the target radiation signal to low-frequency broadband noise.

[0014] S120. Generate target motion trajectory: Within a region centered on the receiving array and with a set radius, N target motion trajectories are randomly generated. Each trajectory is defined by an initial position and velocity, and it is assumed that the target's depth remains constant during the motion.

[0015] S130. Simulation generation of vertical time recording map: For each trajectory, based on the set target depth, the sound field transfer function is calculated using the sound field model, the received signal of the receiving array is simulated, and continuous narrowband beamforming is performed on the received signal to obtain the corresponding vertical time recording map p. n ;

[0016] S140. Constructing the sample set: Record graph p for each vertical time period. n According to the set sample length l sample and overlap length l overlap Sliding window segmentation is performed to obtain multiple fixed-size image blocks. Each image block inherits the depth label of the original vertical time record map, forming a standardized sample set.

[0017] Preferably, in step S120, the target always moves toward the receiving array.

[0018] Preferably, the multi-scale basic module includes at least three parallel branches. The first branch sequentially includes a first-size convolutional layer, a first normalization layer, and a first activation function layer; the second branch sequentially includes a second-size convolutional layer, a second normalization layer, and a second activation function layer; and the third branch sequentially includes a third-size convolutional layer, a third normalization layer, and a third activation function layer. The first, second, and third sizes are all different. The multi-scale basic module also includes an adder that sums the outputs of the at least three parallel branches.

[0019] Preferably, the first size is 3×3, the second size is 5×5, and the third size is 7×7; the normalization layer is a layer normalization layer; and the activation function layer is a ReLU function layer.

[0020] Preferably, the multi-scale basic module further includes a fourth branch, which is an identity mapping layer or a 1×1 convolutional layer, and the output of the fourth branch is also input to the adder.

[0021] Preferably, the cross-stage feature fusion module has two inputs: a first input from the output feature map of the previous stage, and a second input from the output feature map of an even earlier stage; the cross-stage feature fusion module includes:

[0022] The first adjustment convolutional layer is used to adjust the number of channels and / or size of the first input;

[0023] The second adjustment convolutional layer is used to adjust the number of channels and / or size of the second input;

[0024] Learnable weight coefficients are used to weight the feature maps processed by the second adjusted convolutional layer;

[0025] The fusion unit is used to add the feature map processed by the first adjusted convolutional layer to the weighted feature map, and output it after passing it through an activation function.

[0026] Preferably, the structure of the multi-scale cross-stage feature fusion classification model includes:

[0027] Initial layer: consists of a 7×7 convolutional layer, a layer normalization layer, a ReLU activation function layer, and a max pooling layer;

[0028] The first stage includes a 1×1 convolutional layer and at least one of the aforementioned multi-scale basic modules;

[0029] The second stage includes a cross-stage feature fusion module that fuses the feature map output from the initial layer, and at least one of the multi-scale basic modules.

[0030] The third stage includes a cross-stage feature fusion module that fuses the feature map output from the first stage, and at least one of the multi-scale basic modules.

[0031] The fourth stage includes a cross-stage feature fusion module that fuses the feature map output from the second stage, and at least one of the multi-scale basic modules.

[0032] Classification layer: includes global average pooling layer and fully connected layer.

[0033] Preferably, in step S300, a two-stage training strategy is adopted:

[0034] S310, First stage training: Train the model using a stochastic gradient descent optimizer until the loss function converges;

[0035] S320, Second Stage Training: Using the Adam optimizer, the model is fine-tuned and trained at a learning rate lower than that of the first stage until the loss function converges.

[0036] Compared with existing technologies, the beneficial effects of this invention are as follows: It constructs a complete deep learning model that directly maps from the original VTR image to the target category, eliminating the tedious and potentially suboptimal intermediate steps of manual feature design and template matching in traditional methods, thus simplifying the processing flow; it utilizes the powerful feature learning capabilities of deep convolutional neural networks to automatically extract deep and abstract features related to the target depth from interference structure images. These features may exceed the scope of manual design, thereby making fuller use of data information; the network model is trained with a large number of samples containing random motion trajectories and sound field changes, learning the inherent laws of interference structure changes, and has good adaptability to changes in target motion state and fluctuations in environmental parameters within a certain range. Attached Figure Description

[0037] Figure 1 This is a schematic diagram of the deep-sea environment in which the method of the present invention is applied;

[0038] Figure 2 This is a flowchart illustrating the overall signal processing of the method of the present invention;

[0039] Figure 3 This is a schematic diagram of the multi-scale basic module structure used in the network architecture of this invention;

[0040] Figure 4 This is a schematic diagram of the cross-stage feature fusion module structure used in the network architecture of this invention;

[0041] Figure 5 This is a schematic diagram of the network structure constructed in this invention;

[0042] Figure 6 The examples used for model training are (a) VTR images of surface targets and (b) VTR images of underwater targets. Detailed Implementation

[0043] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings, so that those skilled in the art can more clearly understand how to practice the present invention. Although the present invention has been described in conjunction with its preferred embodiments, these embodiments are merely illustrative and not intended to limit the scope of the invention.

[0044] See Figure 1-6 In one embodiment of the present invention, an intelligent classification method for surface and underwater targets based on a grazing angle-range domain interference structure is provided. This method is applicable to deep-sea environments. In deep-sea environments, a vertical uniform linear array is placed near the seabed. Due to the existence of reliable acoustic paths in the deep sea, the vertical array can capture target radiation signals at medium and long distances.

[0045] The specific implementation method is as follows:

[0046] (1) Construction of standard sample set: This step simulates the sound propagation process of targets at different depths in the deep sea environment through sound field simulation, generating a large number of VTR map samples with category labels. The main process is as follows:

[0047] (1.1) Simulation parameter settings: The depth of the surface target is set to z0=10m, and the depth of the underwater target is set to z1=100m. Both the surface and underwater targets radiate low-frequency broadband noise. The receiving array is a vertical uniform linear array. The receiving array is placed near the seabed at a depth of 4000m. The vertical array has 32 elements and the element spacing is 4m. The target moves within a radius R=20km centered on the vertical array. The target always moves at a constant speed and its depth remains unchanged during the movement.

[0048] (1.2) Randomly generate N target motion trajectories, and establish a two-dimensional rectangular coordinate system with the vertical matrix as the origin. The initial coordinates of the targets are x0=[x0,y0]. T The initial ordinate of the target is randomly selected within the range of [0, 15] km, and the initial abscissa of the target is... The target velocity v takes random values ​​within the range [5, 10]kn; the target always moves in the negative x-axis direction, and the target's movement time within the observation area is... At any time k, the two-dimensional position of the target can be represented as x. k =[x k ,y k The target's horizontal distance from the vertical array is... ;

[0049] (1.3) Randomly select z0 and z1 as the target navigation depth. Simulate the vertical array received signal based on the pre-generated target motion trajectory and the sound field information obtained from the sound field model. Perform continuous narrowband beamforming on the vertical array received signal to obtain the corresponding target's vertical time record (VTR) (Figure p). n The VTR atlas was obtained by simulating N trajectories. The target motion time varies, while the time interval between each narrowband beamforming is the same, resulting in different VTR map lengths. The length of the nth VTR map is denoted by l. n This indicates that the sample length is set to l. sample The sample overlap length is l overlap The VTR atlas was examined sequentially. The sample set is obtained by segmentation. Where n represents the nth VTR image, m n This indicates the number of times the nth VTR graph is segmented.

[0050] (2) Construction of a multi-scale, cross-stage feature fusion classification model: This application designs a dedicated deep convolutional neural network architecture that integrates multi-scale convolution and cross-stage feature fusion mechanisms to enhance the model's ability to extract multi-resolution features from VTR maps and fuse features at different levels of abstraction. The model includes the following parts in sequence:

[0051] (2.1) Construct a basic module with the number of output channels settable by parameter x. Add four parallel branches: Branch 1 includes a convolutional layer (3×3, x, [1,1]), an LN normalized layer, and a ReLU activation function. The convolutional layer parameter (3×3, x, [1,1]) indicates that the kernel size is 3×3, the number of convolutional channels is x, and the horizontal and vertical strides of the convolutional kernel are both 1. Branch 2 includes a convolutional layer (5×5, x, [1,1]), an LN normalized layer, and a ReLU activation function. Branch 3 includes a convolutional layer (7×7, x, [1,1]), an LN normalized layer, and a ReLU activation function. Branch 4 is a direct connection layer that adds the convolutional features output from the four branches in the channel dimension to obtain the final output of the module. In other words, the outputs of the four branches are added in the channel dimension to form the final output of the module. This design allows the network to capture local features under different receptive fields simultaneously.

[0052] (2.2) Construct a cross-stage feature fusion module, such as Figure 4 The cross-stage module has two inputs: input 1 is the output of the previous stage, and input 2 is the cross-stage output. The number of channels and dimensions of input 1 are different from those of input 2. A (1×1, x, [2,2]) convolutional layer is added to input 1; a (1×1, x, [S1,S2]) convolutional layer is added to input 2. The result of input 2 is multiplied by the parameter α and added to the result of input 1 in the channel dimension. Then, the ReLU activation function is used. This mechanism realizes the complementarity and fusion of deep and shallow features.

[0053] (2.3) Construct a multi-scale, cross-stage feature fusion classification network. The specific process is as follows:

[0054] Construct the initial layer: Add a convolutional layer (7×7,64,[2,2]), add an LN normalization layer, a ReLU activation function, and a max pooling layer (3×3,64,[2,2]);

[0055] Construct stage1: Add a convolutional layer (1×1, 256, [1,1]), add 3 basic modules, and set x to 256;

[0056] Build stage2: Add 1 feature fusion module, set x to 512, set S1 and S2 to 2, add 4 basic modules, set x to 512;

[0057] Build stage3: Add 1 feature fusion module, set x to 1024, set S1 and S2 to 4, add 6 basic modules, and set x to 1024;

[0058] Build stage4: Add 1 feature fusion module, set x to 2048, set S1 and S2 to 4, add 3 basic modules, set x to 2048;

[0059] Construct a classification layer: Add a global average pooling layer (7×7,2048,[1,1]) to flatten the feature vectors, and add a fully connected layer. The input dimension of the fully connected layer is determined by the dimension of the flattened feature vectors, and the output dimension is 2.

[0060] (3) Use the training sample set constructed in step (1) to perform supervised training on the network model constructed in step (2). The training process is divided into two stages to improve the convergence effect and model performance:

[0061] (3.1) Conduct Phase 1 training on the multi-scale cross-stage feature fusion classification model, setting the optimizer, learning rate, and data block size N during iterative training. bz The training parameters are set as follows, where the optimizer is set to SGD, the learning rate is set to 0.001, and N... bz =32. N is randomly selected with replacement from the training dataset. bz For each set of samples, forward inference calculation is performed based on a multi-scale cross-stage feature fusion classification model to obtain the target classification result. Combined with the sample labels, the loss value is obtained based on the cross-entropy loss function. Then, the structural parameters of the multi-scale cross-stage feature fusion classification model are optimized using the set SGD optimizer. The model is trained in multiple rounds until the loss function basically converges.

[0062] (3.2) Conduct Phase 2 training on the multi-scale cross-stage feature fusion classification model, setting the optimizer, learning rate, and data block size N during iterative training. bz The training parameters are set as follows, where the optimizer is set to Adam, the learning rate is set to 0.0005, and N... bz =16, N is randomly selected with replacement from the training dataset. bz For each set of samples, forward inference calculation is performed based on a multi-scale, cross-stage feature fusion classification model to obtain the target classification result. Combined with the sample labels, the loss value is obtained based on the cross-entropy loss function. Then, the structural parameters of the multi-domain feature deep fusion recognition model are optimized using the set Adam optimizer. The model is trained in multiple rounds until the loss function basically converges.

[0063] (4) Classification of unknown underwater acoustic targets, the basic process is as follows:

[0064] Using the sample simulation conditions and process in step (1), a test sample set is generated through simulation. The multi-scale cross-stage feature fusion classification model constructed in step (2) and trained in step (3) is used to perform forward inference on the VTR sample to obtain the classification result of the underwater acoustic target data. Specifically, the measured VTR data of the target to be classified at unknown depth is preprocessed in the same way as the training samples, such as size normalization. Then, it is input into the multi-scale cross-stage feature fusion classification model that has been trained. The model performs forward propagation calculation, and the final output layer gives the probability values ​​of the two categories. The category with the higher probability value is taken as the classification result of the unknown target (surface or underwater).

[0065] Model testing and performance evaluation:

[0066] The trained final model was evaluated using a reserved test set (approximately 4500 samples): test samples were input into the model to obtain the probability of each sample belonging to the surface and underwater categories; the overall classification accuracy was calculated, and the classification accuracy for targets at depths of 10m (surface) and 100m (underwater) was statistically analyzed respectively; the test results show that the model achieved excellent performance on the test set, with a classification accuracy of 93.34% for surface targets (10m) and 94.68% for underwater targets (100m), and an overall accuracy of approximately 94.01%. This indicates that the method proposed in this application can effectively learn depth-related interference structure features in VTR maps and achieve high-precision classification of surface and underwater targets.

[0067] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for intelligent classification of surface and underwater targets based on a grazing angle-range domain interferometric structure, characterized in that, Includes the following steps: S100. Construct a standardized training sample set. Based on the set deep-sea environment parameters, receiving array parameters, and depth parameters of surface and underwater targets, simulate the movement of the target in the horizontal plane through sound field simulation, generate a vertical time record map containing target depth category labels, and segment and standardize the vertical time record map to obtain a sample set of fixed size. S200. Construct a multi-scale cross-stage feature fusion classification model. Construct a deep convolutional neural network comprising an initial convolutional layer, at least two feature extraction stages, and a classification layer connected sequentially. At least one feature extraction stage includes multiple cascaded multi-scale basic modules, each with multiple convolutional branches operating in parallel, and the kernel sizes of these branches are different. At least one feature extraction stage is connected to at least one previous stage via a cross-stage feature fusion module, which fuses feature maps from different depth stages. S300. Train the multi-scale cross-stage feature fusion classification model by using the standardized training sample set to perform supervised training on the multi-scale cross-stage feature fusion classification model to obtain a trained classification model. S400. Classify the target data to be identified, obtain the vertical time record data of the target to be identified and perform the same standardization process as in step S100, input the trained classification model, and obtain the classification result of the water surface or underwater.

2. The method according to claim 1, characterized in that, Step S100 includes: S110. Set simulation parameters: Set sea depth, sound velocity profile, and seabed parameters; Set the receiving array to a vertical uniform linear array, and set its number of array elements and element spacing; Set the reference depth z0 for the surface target and the reference depth z1 for the underwater target; Set the target radiation signal to low-frequency broadband noise. S120. Generate target motion trajectory: Within a region centered on the receiving array and with a set radius, N target motion trajectories are randomly generated. Each trajectory is defined by an initial position and velocity, and it is assumed that the target's depth remains constant during the motion. S130. Simulation generation of vertical time recording map: For each trajectory, based on the set target depth, the sound field transfer function is calculated using the sound field model, the received signal of the receiving array is simulated, and continuous narrowband beamforming is performed on the received signal to obtain the corresponding vertical time recording map. p n ; S140. Constructing the sample set: Recording graphs for each vertical time period. p n According to the set sample length l sample and overlap length l overlap Sliding window segmentation is performed to obtain multiple fixed-size image blocks. Each image block inherits the depth label of the original vertical time record map, forming a standardized sample set.

3. The method according to claim 2, characterized in that: In step S120, the target always moves toward the receiving array.

4. The method according to claim 1, characterized in that: The multi-scale basic module includes at least three parallel branches. The first branch includes, in sequence, a first-size convolutional layer, a first normalization layer, and a first activation function layer; the second branch includes, in sequence, a second-size convolutional layer, a second normalization layer, and a second activation function layer; and the third branch includes, in sequence, a third-size convolutional layer, a third normalization layer, and a third activation function layer. The first, second, and third sizes are all different. The multi-scale basic module also includes an adder that sums the outputs of the at least three parallel branches.

5. The method according to claim 1, characterized in that: The first size is 3×3, the second size is 5×5, and the third size is 7×7; the normalization layer is a layer normalization layer; and the activation function layer is a ReLU function layer.

6. The method according to claim 4 or 5, characterized in that: The multi-scale basic module also includes a fourth branch, which is an identity mapping layer or a 1×1 convolutional layer, and the output of the fourth branch is also input to the adder.

7. The method according to claim 1, characterized in that, The cross-stage feature fusion module has two inputs: the first input comes from the output feature map of the previous stage, and the second input comes from the output feature map of an even earlier stage. The cross-stage feature fusion module includes: The first adjustment convolutional layer is used to adjust the number of channels and / or size of the first input; The second adjustment convolutional layer is used to adjust the number of channels and / or size of the second input; Learnable weight coefficients are used to weight the feature maps processed by the second adjusted convolutional layer; The fusion unit is used to add the feature map processed by the first adjusted convolutional layer to the weighted feature map, and output it after passing it through an activation function.

8. The method according to claim 1, characterized in that: The structure of the multi-scale, cross-stage feature fusion classification model includes: Initial layer: consists of a 7×7 convolutional layer, a layer normalization layer, a ReLU activation function layer, and a max pooling layer; The first stage includes a 1×1 convolutional layer and at least one of the aforementioned multi-scale basic modules; The second stage includes a cross-stage feature fusion module that fuses the feature map output from the initial layer, and at least one of the multi-scale basic modules. The third stage includes a cross-stage feature fusion module that fuses the feature map output from the first stage, and at least one of the multi-scale basic modules. The fourth stage includes a cross-stage feature fusion module that fuses the feature map output from the second stage, and at least one of the multi-scale basic modules. Classification layer: includes global average pooling layer and fully connected layer.

9. The method according to claim 1, characterized in that: In step S300, a two-stage training strategy is adopted: S310, First stage training: Train the model using a stochastic gradient descent optimizer until the loss function converges; S320, Second Stage Training: Using the Adam optimizer, the model is fine-tuned and trained at a learning rate lower than that of the first stage until the loss function converges.