A radar-based method, medium, and equipment for identifying small targets on the sea surface based on CNN-GCN dual-modal fusion.

CN122568461APending Publication Date: 2026-08-14JINLING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

同时,同一类型目标在不同海况条件下的回波呈现出非平稳、非线性等特性,进一步增加了分类识别的难度

Benefits of technology

[0017]本发明的有益效果是:本发明将傅里叶同步压缩变换作为CNN的输入,时域波形的可视图邻接矩阵与图单跳算子作为GCN输入,通过构建CNN-GCN双模态网络实现对如漂浮小球和漂浮船只的分类识别。与单模态相比,本发明有效融合了观测信号的时频分布特性及其拓扑特征,并通过引入GCN对时频结构关联及时间拓扑演化过程进行建模,增强了在复杂海杂波背景下对微弱目标特征的表征能力。实验结果表明,该方法在识别精度上优于现有模型。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122568461A_ABST
    Figure CN122568461A_ABST
Patent Text Reader

Abstract

This invention proposes a radar-based method, medium, and device for identifying small targets on the sea surface based on CNN-GCN dual-modal fusion, belonging to the field of signal processing technology. The method includes: acquiring the echo sequence received by the radar as the observation signal; extracting the time-domain amplitude sequence and FSST time-frequency map from the observation signal; constructing a visual graph from the time-domain amplitude sequence and defining graph single-hop filters as node features; feeding the adjacency matrix and node features of the FSST time-frequency map and the visual graph into a CNN branch network and a GCN branch network respectively for automatic feature extraction; and then concatenating the extracted features to complete the identification. This invention increases the feature differences between different targets by introducing a CNN-GCN network model. Experimental results show that the identification performance of this invention is superior to traditional CNN-based algorithms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of signal processing technology, specifically relating to a radar method, medium, and device for identifying small targets on the sea surface based on CNN-GCN dual-modal fusion. Background Technology

[0002] In radar-based maritime detection scenarios, target types are diverse, ranging from large ships, islands, reefs, and icebergs to small and medium-sized floating objects, ice floes, and even divers. Therefore, whether for military applications like shore-based radar monitoring the maritime environment or civilian needs like maritime search and rescue, different target types often require different processing strategies and decision-making methods. Thus, the classification and identification of maritime targets has significant research and application value. Typically, maritime target detection relies mainly on clutter and the number of target samples. However, maritime target classification relies primarily on the variety of target samples. In maritime detection, clutter is often easily acquired while targets are difficult to obtain—a phenomenon known as class imbalance. Therefore, maritime target classification must consider algorithm design issues even with a limited number of target samples. Furthermore, the echoes of the same type of target under different sea conditions exhibit non-stationary and non-linear characteristics, further increasing the difficulty of classification and identification. Summary of the Invention

[0003] To address the shortcomings of existing technologies, this invention provides a radar-based method, medium, and device for identifying small targets on the sea surface based on CNN-GCN dual-modal fusion.

[0004] To achieve the above objectives, the present invention adopts the following technical solution:

[0005] In a first aspect, the present invention provides a radar-based method for identifying small targets on the sea surface based on CNN-GCN dual-modal fusion, comprising the following steps: Step 1: Obtain the echo sequence received by the radar as the observation signal, and extract the time-domain amplitude sequence and FSST time-frequency diagram from the observation signal; Step 2: Construct a visual graph from the time-domain amplitude sequence and define the graph single-hop filter as a node feature; Step 3: The adjacency matrix and node features of the FSST time-frequency map and visual map are fed into the CNN branch network and GCN branch network respectively for abstract feature extraction. The extracted features are then concatenated to complete the recognition.

[0006] Optionally, in step 1, the time-domain amplitude sequence is:

[0007] In the formula, Represents the time-domain amplitude sequence. Indicates the observed signal, Indicates the sample number. Indicates the number of samples; The FSST time-frequency diagram is represented as follows:

[0008]

[0009] In the formula, FSST time-frequency representation, , Indicates continuous , and Indicates frequency and time delay. Represents the window function. Let be the Dirac impulse function. These are the compressed frequency coordinates. Indicates the instantaneous frequency operator. This indicates taking the real part. Indicates to Find the partial derivative.

[0010] Optionally, in step 2, constructing the time-domain amplitude sequence into a visual diagram specifically involves: A visual transformation is performed on the time-domain amplitude sequence. In the visual transformation, the number of samples in the time-domain amplitude sequence... The number of vertices in the visible view ,satisfy The edge is defined by the visibility between two samples: (denotes...) , , The sample value at time t is , , The corresponding coordinates in a two-dimensional coordinate system are , , , If the coordinates of the three nodes satisfy If these three nodes are considered visible and have edges, then the adjacency matrix can be obtained. .

[0011] Optionally, in step 2, the definition of the graph single-hop filter as a node feature specifically refers to: Visual have The unsigned Laplace matrix of the given vertices is: , For degree matrix, Given an adjacency matrix, the corresponding graph signal is: , If the observed signal is represented, then the graph single-hop filter is defined as follows: .

[0012] Optionally, in step 3, the FSST time-frequency map is fed into a CNN branch network for abstract feature extraction. The CNN branch network uses ResNet18 as the main network to extract features from the FSST time-frequency map, and compresses the dimension of the output feature vector of ResNet18 to 256×1 through a fully connected layer.

[0013] Optionally, in step 3, the adjacency matrix and its node features of the visual image are put into the GCN branch network for abstract feature extraction. The GCN branch network extracts features from the adjacency matrix and its node features through GAT and residual structure, and at the same time, self-attention pooling is used to perform structure-adaptive downsampling on the extracted features. Finally, a 256×1 feature vector is obtained through global aggregation.

[0014] Optionally, in step 3, the feature vectors output by the CNN branch network and the GCN branch network are concatenated to obtain a 512×1 feature vector. The concatenated feature vector is then input into the fully connected layer to complete classification and recognition.

[0015] In a second aspect, the present invention provides a computer-readable storage medium storing a computer program that causes a computer to execute the radar sea surface small target identification method based on CNN-GCN dual-modal fusion as described in the first aspect.

[0016] Thirdly, the present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the radar sea surface small target identification method based on CNN-GCN dual-modal fusion as described in the first aspect.

[0017] The beneficial effects of this invention are as follows: This invention uses the Fourier synchronous compressed transform as the input to a CNN, and the visual adjacency matrix of the time-domain waveform and the graph single-hop operator as the input to a GCN. By constructing a CNN-GCN dual-modal network, it achieves the classification and recognition of objects such as floating balls and floating vessels. Compared with single-modal methods, this invention effectively integrates the time-frequency distribution characteristics and topological features of the observed signals, and enhances the ability to represent weak target features in complex sea clutter backgrounds by introducing a GCN to model the time-frequency structural correlation and temporal topological evolution process. Experimental results show that this method outperforms existing models in terms of recognition accuracy. Attached Figure Description

[0018] Figure 1 This is a flowchart of a radar-based method for identifying small targets on the sea surface based on CNN-GCN dual-modal fusion.

[0019] Figure 2 It is the topology of the view generated by the floating ball.

[0020] Figure 3 It is a graph signal mapping generated by the floating ball.

[0021] Figure 4 It is the topology of the view generated by the floating boat.

[0022] Figure 5 It is a graph signal mapping generated by the floating boat.

[0023] Figure 6 This is a time-frequency distribution representation of the FSST generated by the floating ball.

[0024] Figure 7 This is a time-frequency distribution representation of FSST generated by a floating boat.

[0025] Figure 8 It is a CNN-GCN dual-modal fusion network model.

[0026] Figure 9 This is a schematic diagram of the GCN branch network structure.

[0027] Figure 10 This is a schematic diagram of the structure of a CNN branch network.

[0028] Figure 11 This is a schematic diagram comparing the confusion matrix of the recognition accuracy of the method of this invention with that of existing algorithms. Detailed Implementation

[0029] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0030] In one embodiment, the present invention proposes a radar-based method for identifying small targets on the sea surface based on CNN-GCN dual-modal fusion, such as... Figure 1 As shown, it includes the following steps: Step 1: Obtain the echo sequence received by the radar as the observation signal, and extract the time-domain amplitude sequence and FSST time-frequency distribution representation from the observation signal.

[0031] Step 1.1: Calculate the observed signal Time-domain amplitude sequence:

[0032] In the formula, Represents the time-domain amplitude sequence. Indicates the observed signal, Indicates the sample number. Indicates the number of samples.

[0033] Step 1.2: Extract the observation signal FSST time-frequency representation:

[0034]

[0035] In the formula, FSST time-frequency representation, , Indicates continuous , and Indicates frequency and time delay. Represents the window function. Let be the Dirac impulse function. These are the compressed frequency coordinates. Indicates the instantaneous frequency operator. This indicates taking the real part. Indicates to Find the partial derivative.

[0036] Figure 6 and Figure 7 FSST time-frequency distribution representations generated for floating balls and floating boats.

[0037] Step 2: Construct a visual graph from the time-domain amplitude sequence and define the graph single-hop filter as a node feature.

[0038] Step 2.1: Visual transformation treats each sample point in the time series as a vertex of the graph. At the same time, if the line connecting any two sample points is geometrically visible (i.e., there are no other sample points obscuring it), then it is considered that there is an edge between them.

[0039] In graph transformation, the number of samples in the sequence That is, the number of vertices in the visible view. ,satisfy And the edge is defined by the visibility between two samples: Let the edge be defined by the visibility between two samples. The sample value at time t is The corresponding coordinates in a two-dimensional coordinate system are... Correspondingly, there are also and ,satisfy If the coordinates of the three nodes satisfy:

[0040] These three nodes are considered visible and therefore have edges. From this, the adjacency matrix can be obtained. .

[0041] Step 2.2: Graphical Single-Hop Filter. Assume a graphical single-hop filter. have The unsigned Laplace matrix of the given vertices is: The corresponding image signal is Then the graph single-hop filter is defined as:

[0042] in, It is a degree matrix.

[0043] Figures 2 to 5 It is a visual topology and graph signal mapping generated by floating balls and floating boats.

[0044] Step 3: Input the FSST time-frequency map and the adjacency matrix of the visual map and its node features into the CNN and GCN networks respectively for abstract feature extraction, and then concatenate the extracted features to complete the recognition.

[0045] like Figure 8 As shown, the CNN-GCN dual-modal fusion network architecture consists of three parts: feature extraction, feature fusion, and classification decision. In feature extraction, the network inputs are the adjacency matrix of the visual graph, node features composed of graph single-hop filtering operators, and the time-frequency graph. Specifically, graph domain features are extracted by the graph neural network, while the time-frequency graph features are extracted by the ResNet18 network.

[0046] like Figure 9 As shown, the GCN branch network achieves stable extraction of high-order features of nodes through GAT and residual structure; at the same time, it adopts self-attention pooling (SAGPooling) for structure adaptive downsampling to retain key node information, and finally obtains a more discriminative global graph-level representation through global aggregation to extract the features of the entire graph and obtain a 256×1 feature vector.

[0047] like Figure 10 As shown, the CNN branch network uses ResNet18 as the main network to extract features from time-frequency represented images. ResNet18 effectively alleviates the vanishing and exploding gradient problems through residual connections, improving the stability of feature extraction and recognition accuracy. Simultaneously, this structure significantly reduces the number of model parameters, accelerating the network's training and convergence process. Finally, a fully connected layer compresses the dimension of the ResNet18 output feature vector to 256×1. Two feature vectors of the same length are obtained through the two branch networks respectively. These two feature vectors are concatenated using Concat to obtain a 512×1 feature vector, achieving feature fusion. Finally, the concatenated feature vector is fed into a fully connected layer to complete classification and recognition.

[0048] Figure 11The following is a comparison of the detection probabilities of the method in this embodiment with those of classic CNN and existing network recognition algorithms. (a) is SPWVD+AlexNet, (b) is SPWVD+VGG16, (c) is STFT+MBRF, and (d) is the CNN / GCN bimodal network model proposed in this embodiment. (1) Training and test set partitioning: Each segment is 512 bytes long. Data sets are constructed for floating balls and floating boats. The partitioning of training and test sets yields 6450 training sets and 1866 test sets for floating balls, and 3420 training sets and 904 test sets for floating boats. (2) Cross-entropy loss function and AdamW optimizer are used. The batch size and initial learning rate are set to 32 and 0.0001, respectively, and the learning rate is adjusted using a cosine annealing strategy. As can be seen from the figure, the detection probability of this method is better than that of existing algorithms in most datasets, which means that the detection performance of the method in this embodiment is superior. Therefore, the method in this embodiment can be well applied to the identification of small targets on the sea surface under different radars, and the identification probability is significantly improved compared with the existing methods.

[0049] In another embodiment, the present invention provides a computer-readable storage medium storing a computer program that enables a computer to execute the radar surface small target identification method based on CNN-GCN dual-modal fusion of the foregoing embodiments.

[0050] In another embodiment, the present invention proposes an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the radar surface small target recognition method based on CNN-GCN dual-modal fusion of the aforementioned embodiments.

[0051] In the embodiments disclosed in this application, a computer storage medium may be a tangible medium that may contain or store programs for use by or in conjunction with an instruction execution system, apparatus, or device. The computer storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of computer storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CDROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0052] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0053] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should be considered within the scope of protection of the present invention.

Claims

1. A radar-based method for identifying small targets on the sea surface based on CNN-GCN dual-modal fusion, characterized in that, Includes the following steps: Step 1: Obtain the echo sequence received by the radar as the observation signal, and extract the time-domain amplitude sequence and FSST time-frequency diagram from the observation signal; Step 2: Construct a visual graph from the time-domain amplitude sequence and define the graph single-hop filter as a node feature; Step 3: The adjacency matrix and node features of the FSST time-frequency map and visual map are fed into the CNN branch network and GCN branch network respectively for abstract feature extraction. The extracted features are then concatenated to complete the recognition.

2. The radar-based small target identification method for the sea surface based on CNN-GCN dual-modal fusion as described in claim 1, characterized in that: In step 1, the time-domain amplitude sequence is: In the formula, Represents the time-domain amplitude sequence. Indicates the observed signal, Indicates the sample number. Indicates the number of samples; The FSST time-frequency diagram is represented as follows: In the formula, FSST time-frequency representation, , Indicates continuous , and Indicates frequency and time delay. Represents the window function. Let be the Dirac impulse function. These are the compressed frequency coordinates. Indicates the instantaneous frequency operator. This indicates taking the real part. Indicates to Find the partial derivative.

3. The radar-based small target identification method for the sea surface based on CNN-GCN dual-modal fusion as described in claim 1, characterized in that: In step 2, the specific steps of constructing a visual chart from the time-domain amplitude sequence are as follows: A visual transformation is performed on the time-domain amplitude sequence. In the visual transformation, the number of samples in the time-domain amplitude sequence... The number of vertices in the visible view ,satisfy ; An edge is defined by the visibility between two samples: (denotes...) , , The sample value at time t is , , The corresponding coordinates in a two-dimensional coordinate system are , , , If the coordinates of the three nodes satisfy If these three nodes are considered visible and have edges, then the adjacency matrix can be obtained. .

4. The radar-based small target identification method for the sea surface based on CNN-GCN dual-modal fusion as described in claim 1, characterized in that: In step 2, the definition of the graph single-hop filter as a node feature specifically refers to: Visual have The unsigned Laplace matrix of the given vertices is: , For degree matrix, Given an adjacency matrix, the corresponding graph signal is: , If the observed signal is represented, then the graph single-hop filter is defined as follows: .

5. The radar-based small target identification method for the sea surface based on CNN-GCN dual-modal fusion as described in claim 1, characterized in that: In step 3, the FSST time-frequency map is fed into a CNN branch network for abstract feature extraction. The CNN branch network uses ResNet18 as the main network to extract features from the FSST time-frequency map, and compresses the dimension of the output feature vector of ResNet18 to 256×1 through a fully connected layer.

6. The radar-based small target identification method for the sea surface based on CNN-GCN dual-modal fusion as described in claim 5, characterized in that: In step 3, the adjacency matrix and node features of the visual image are put into the GCN branch network for abstract feature extraction. The GCN branch network extracts features from the adjacency matrix and node features through GAT and residual structure. At the same time, self-attention pooling is used to perform structure-adaptive downsampling on the extracted features. Finally, a 256×1 feature vector is obtained through global aggregation.

7. The radar-based small target identification method for the sea surface based on CNN-GCN dual-modal fusion as described in claim 6, characterized in that: In step 3, the feature vectors output by the CNN branch network and the GCN branch network are concatenated to obtain a 512×1 feature vector. The concatenated feature vector is then input into the fully connected layer to complete the classification and recognition.

8. A computer-readable storage medium storing a computer program, characterized in that, The computer program causes the computer to execute the radar surface small target identification method based on CNN-GCN dual-modal fusion as described in any one of claims 1-7.

9. An electronic device, characterized in that, include: The system includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the radar surface small target identification method based on CNN-GCN dual-modal fusion as described in any one of claims 1-7.