Multi-node multi-source heterogeneous information fusion and situation information generation method and system
Through the multi-node multi-source heterogeneous information fusion method, the improved YOLOV8 model and D-S evidence theory are used to combine visible light and SAR images to generate high confidence ship detection and motion information, solving the problem of ship detection and recognition in complex marine environments and achieving efficient situation information generation.
Patent Information
- Application Number
- CN202510787281.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-13
AI Technical Summary
In complex marine scenarios, single load imaging cannot achieve high confidence ship detection and recognition. It is disturbed by factors such as sea surface islands and reefs, sea surface waves, clouds, fog, rain, snow, etc., and is seriously affected by meteorological factors, making it difficult to obtain ship motion information with high update rate.
The multi-node multi-source heterogeneous information fusion method is adopted, and the visible light image and SAR image combined with D-S evidence theory is used to generate high confidence target recognition results through the improved YOLOV8 model and multi-branch deep feature fusion network, and the target heading, speed and track are extracted to form situation information.
It realizes high confidence detection of ship targets and rapid extraction of motion information in complex marine environments, assists in port management decisions, and improves target recognition rate and information integrity.
Smart Images

Figure CN120298918A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing image processing methods and systems, and specifically, to a method and system for multi-node multi-source heterogeneous information fusion and situation information generation. Background Art
[0002] With the rapid development of the global marine economy and port transportation industry, marine shipping has become an important pillar of international trade. However, the complexity of maritime activities has been increasing continuously. Especially in busy port areas and important sea lanes, the types and quantities of vessels have been continuously increasing. Using satellite payload imaging has gradually become an important method for obtaining vessel information and assisting maritime traffic management.
[0003] Under high-resolution data source conditions, the number of pixels occupied by vessel targets is large. Generally, the outlines are clear and the detailed features are rich. Single-payload target detection and recognition can detect and recognize vessels with high confidence according to various obvious vessel features.
[0004] However, in complex marine scenarios, factors such as sea reefs and sea spray on the sea surface interfere with vessel detection, and meteorological factors such as clouds, fog, rain, and snow seriously affect payload imaging. It is impossible to achieve high-confidence detection and recognition of vessels relying on a single payload.
[0005] For information elements such as vessels with a high update rate requirement, the information related to target movement is crucial. The core means of obtaining such information lies in continuously observing key targets using multiple payloads, fusing multi-source information, and quickly generating relevant situation information. Therefore, there is an urgent need to design a method and system for multi-node multi-source heterogeneous information fusion and situation information generation. By performing fusion processing and intelligent correlation analysis on multi-dimensional heterogeneous monitoring information, important information is extracted from a large amount of marine vessel information to achieve intelligent perception of the maritime activity situation and assist ports in efficiently completing management decisions. Summary of the Invention
[0006] Aiming at the defects in the prior art, the purpose of the present invention is to provide a method and system for multi-node multi-source heterogeneous information fusion and situation information generation.
[0007] According to a multi-node multi-source heterogeneous information fusion and situation information generation method provided by the present invention, it includes: Step S1: Perform target detection on visible light images and SAR images respectively to generate visible light image slices and SAR image slices; Step S2: Input the visible light image slices and SAR image slices into a multi-branch deep feature fusion network for target recognition; Step S3: Use the D-S evidence theory to perform decision-level fusion on the recognition results and electronic reconnaissance data to further obtain high-confidence recognition results; Step S4: Generate the entire situation base map, extract the target course and speed based on the situation base map according to the target recognition result, generate a track based on the target course and speed, and map the track onto the situation base map to generate situation information.
[0008] Preferably, the said step S1 includes: Step S1.1: Use the improved YOLOV8 model based on the C2f-PKI module to perform target detection on the visible light image to generate visible light image slices; Step S1.2: Use the improved YOLOV8 model based on the CBAM-SPPF module to perform target detection on the SAR image to generate SAR image slices; The improved YOLOV8 model based on the C2f-PKI module combines the C2f module for extracting and transforming input features with the PKI module for capturing texture features of various scales to form the C2f-PKI module. The C2f-PKI module performs feature extraction through global average pooling and multiple depth convolution kernels of different sizes to achieve the purpose of detecting and capturing texture features of targets of different sizes; The improved YOLOV8 model based on the CBAM-SPPF module adds the CBAM module for enhancing the feature representation ability in the convolutional neural network after the SPPF module for fusing large-scale global information to form the CBAM-SPPF module, enhancing the network's feature extraction ability from both spatial and channel dimensions, so as to learn more target feature information and position information.
[0009] Preferably, the multi-branch deep feature fusion network includes: two independent feature extractors, one feature fusion module, and three independent classifiers; Use two independent feature extractors to perform feature extraction on the visible light image slices and SAR image slices respectively to obtain optical modality features and SAR modality features; The optical modality features and SAR modality features obtain fused modality features through the feature fusion module; Use three independent classifiers to perform target recognition on the optical modality features, SAR modality features, and fused modality features respectively; Among them, the feature extractor is the improved ResNet-50; the improved ResNet-50 removes the last two layers of ResNet-50, namely the global average pooling layer and the fully connected layer, and adjusts the output to a high-dimensional feature map that meets the preset requirements, retaining the spatial information of the image; The feature fusion module splices the optical modality features and SAR modality features, uses a linear layer for fusion, and finally reduces the dimension to the dimension of a single modality; The classifier uses the softmax activation function to convert the output into a probability.
[0010] Preferably, the method further includes: calculating a total loss through weighted calculation based on the cross-entropy loss functions of each branch, and evaluating the performance of the multi-branch depth feature fusion network by using the total loss; Wherein, the total loss includes:
[0011] Wherein, represents the cross-entropy loss of the optical modality; represents the cross-entropy loss of the SAR modality; represents the cross-entropy loss of the fusion modality.
[0012] Preferably, the step S3 includes: Setting the decision space as target, non-target, and uncertain, represented by (1, 0, -1); a types of source types participating in the fusion: e1, e2,... e a ; The correct detection rates of the detection targets corresponding to a types of sources: d1, d2,... d a , its practical meaning is the probability that the detected suspicious targets are confirmed as targets, and it is defined as: Detection target correct rate = number of correctly detected targets / total number of detected targets = detection rate / (detection rate + false alarm rate); Detection rate = number of correctly detected targets / total number of correct targets; False alarm rate = number of incorrectly detected targets / total number of correct targets; The target recognition confidence levels of a types of sources: c1, c2,... c a , then the basic probability assignment function of the i-th source is:
[0013]
[0014]
[0015] The mixed basic probability assignment function is:
[0016]
[0017]
[0018] Wherein, K is the conflict coefficient; Through the above process, the mixed basic probability assignment function is calculated through the visible light image target information, SAR image target information, and electronic reconnaissance image data, and the decision branch with the largest value is selected as the decision result of whether the target exists.
[0019] Preferably, the step S4 includes: When the target appears in the image at the first moment, its speed is default set to 0; for the i-th moment, where i>1, the speed is deduced according to the speed calculation formula; First, determine the driving distance d of the target between the i-th moment and the (i-1)-th moment; the specific calculation method is as follows:
[0020] where R is the radius of the earth, and are the latitudes at the (i-1)-th moment and the i-th moment respectively, and are the longitudes at the (i-1)-th moment and the i-th moment respectively; After calculating the distance, calculate the speed of the current target. The specific calculation formula is:
[0021] where v is the speed, d is the driving distance between the i-th moment and the (i−1)-th moment, is the time interval between the i-th moment and the (i−1)-th moment; Calculate the coordinate difference of the connection line: Let the coordinates at the (i-1)-th moment be ( , ), and the coordinates at the i-th moment be ( , ); the specific calculation formula is as follows:
[0022]
[0023] Calculate the heading angle: Calculate the angle between the connection line and the positive y-axis direction through the arctangent function; since the heading is the angle relative to the north, use Δx and Δy to calculate the target heading θ;
[0024] The calculation result of the heading angle is in degrees, and the result needs to be converted to a preset range; the specific conversion formula is: Heading = (θ + 360) mod 360 Connect the targets at different moments in the direction of the calculated heading and map them to the situation base map according to the coordinates to form situation information.
[0025] According to a multi-node multi-source heterogeneous information fusion and situation information generation system provided by the present invention, it includes: Module M1: Perform target detection on the visible light image and the SAR image respectively to generate visible light image slices and SAR image slices; Module M2: Input visible light image slices and SAR image slices into a multi-branch deep feature fusion network for target recognition; Module M3: Use the D-S evidence theory to perform decision-level fusion on the recognition results and electronic reconnaissance data to further obtain high-confidence recognition results; Module M4: Generate the entire situation base map, extract the target heading and speed based on the target recognition results according to the situation base map, generate a track based on the target heading and speed, and map the track onto the situation base map to generate situation information.
[0026] Preferably, the module M1 includes: Module M1.1: Use an improved YOLOV8 model based on the C2f-PKI module to perform target detection on visible light images to generate visible light image slices; Module M1.2: Use an improved YOLOV8 model based on the CBAM-SPPF module to perform target detection on SAR images to generate SAR image slices; The improved YOLOV8 model based on the C2f-PKI module combines the C2f module for extracting and transforming input features with the PKI module for capturing texture features of various scales to form the C2f-PKI module. The C2f-PKI module performs feature extraction through global average pooling and multiple depth convolution kernels of different sizes to achieve the purpose of detecting and capturing texture features of targets of different sizes.
[0027] Preferably, the multi-branch deep feature fusion network includes: two independent feature extractors, a feature fusion module, and three independent classifiers; Use two independent feature extractors to perform feature extraction on visible light image slices and SAR image slices respectively to obtain optical modality features and SAR modality features; The optical modality features and SAR modality features obtain fused modality features through the feature fusion module; Use three independent classifiers to classify the optical modality features, SAR modality features, and fused modality features respectively; Among them, the feature extractor is an improved ResNet-50; the improved ResNet-50 removes the last two layers of ResNet-50, namely the global average pooling layer and the fully connected layer, and adjusts the output to a high-dimensional feature map that meets the preset requirements, retaining the spatial information of the image; The feature fusion module performs feature stitching on the optical modality features and SAR modality features, uses a linear layer for fusion, and finally reduces the dimension to the dimension of a single modality; The classifier uses the softmax activation function to convert the output into a probability; The cross-entropy loss function based on each branch is weighted to obtain the total loss, and the total loss is used to evaluate the performance of the multi-branch deep feature fusion network; Among them, the total loss includes:
[0028] Among them, represents the cross-entropy loss of the optical modality; represents the cross-entropy loss of the SAR modality; represents the cross-entropy loss of the fusion modality.
[0029] Preferably, the module M4 includes: When the target appears in the image at the first moment, its speed is default set to 0; for the i-th moment, i > 1, the speed is then deduced according to the speed calculation formula; First, determine the driving distance d of the target between the i-th moment and the (i - 1)-th moment; the specific calculation method is as follows:
[0030] Among them, R is the radius of the earth, and are the latitudes at the (i - 1)-th moment and the i-th moment respectively, and are the longitudes at the (i - 1)-th moment and the i-th moment respectively; After calculating the distance, calculate the speed of the current target, and the specific calculation formula is:
[0031] Among them, v is the speed, d is the driving distance between the i-th moment and the (i - 1)-th moment, is the time interval between the i-th moment and the (i - 1)-th moment; Calculate the coordinate difference of the connecting line: Let the coordinates at the (i - 1)-th moment be ( , ), and the coordinates at the i-th moment be ( , ); the specific calculation formula is as follows:
[0032]
[0033] Calculate the heading angle: Calculate the angle between the connecting line and the positive y-axis through the arctangent function; since the heading is the angle relative to the north, use Δx and Δy to calculate the target heading θ;
[0034] The calculation result of the heading angle is in degrees, and the result needs to be converted to a preset range; the specific conversion formula is: Heading = (θ + 360) mod 360 Connect the targets at different times in the direction of the calculated heading and map them to the situation base map according to the coordinates to form situation information.
[0035] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention simultaneously performs end-to-end fast extraction of multi-modal data to form target data with high information density, and solves the problem that multi-source heterogeneous payload data contains a large amount of interference information; 2. The present invention combines deep feature fusion and decision fusion to solve the problem of low target recognition rate caused by the influence of factors such as meteorology, geography, and environment on single payload acquisition, and greatly improves the confidence of the target; 3. Since the ships in the port and on the sea are strongly correlated with the movement information and the target position needs to be updated frequently, the present invention extracts the speed, heading, and track of the target to achieve the purpose of quickly extracting key information from a large amount of target data to assist management decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] By reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent: Figure 1 It is a flowchart of a multi-node multi-source heterogeneous information fusion and situation information generation method.
[0037] Figure 2 It is a structural diagram of the YOLOV8 model improved based on the C2f-PKI module.
[0038] Figure 3 It is the detection result of the YOLOV8 model improved based on the C2f-PKI module for visible light images.
[0039] Figure 4 It is a structural diagram of the YOLOV8 model improved based on the CBAM-SPPF module.
[0040] Figure 5 It is the detection result of the YOLOV8 model improved based on the CBAM-SPPF module for SAR images.
[0041] Figure 6 It is a structural diagram of a multi-branch deep feature fusion network.
[0042] Figure 7 It is a flowchart of decision fusion based on the D-S evidence theory.
[0043] Figure 8This is a partial screenshot of the decision fusion result based on the D-S evidence theory.
[0044] Figure 9 This is a result graph of the situation information generation.
[0045] Figure 10 This is a schematic diagram of a multi-node multi-source heterogeneous information fusion and situation information generation system. Specific implementation manners
[0046] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0047] Embodiment 1 A multi-node multi-source heterogeneous information fusion and situation information generation method provided by the present invention includes: Step 1: Use an improved YOLOV8 model based on the C2f-PKI module to perform target detection on visible light images to generate visible light image slices; use an improved YOLOV8 model based on the CBAM-SPPF module to perform target detection on SAR images to generate SAR image slices; Step 2: Input the visible light image slices and SAR image slices into a multi-branch deep feature fusion network to identify the fine types of targets; Step 3: Use the D-S evidence theory to perform decision-level fusion on the recognition results and electronic reconnaissance data to further obtain high-confidence recognition results; Step 4: Generate the entire situation base map, and based on the situation base map, extract the target heading and speed according to the recognition results to generate a track and map it onto the situation base map to generate situation information.
[0048] Specifically, the improved YOLOV8 model based on the C2f-PKI module combines the C2f module for extracting and transforming input features with the PKI module for capturing texture features of various scales to form the C2f-PKI module.
[0049] Specifically, the improved YOLOV8 model based on the CBAM-SPPF module adds the CBAM module for enhancing the feature representation ability in the convolutional neural network after the SPPF module for fusing large-scale global information.
[0050] Specifically, the multi-branch deep feature fusion network independently obtains the features of visible light image slices and SAR image slices through two feature extractors, fuses the multi-modal features to obtain a joint representation, jointly calculates the loss through a multi-task classifier, and independently classifies the optical modality, SAR modality, and the fused features.
[0051] Specifically, the two feature extractors have the same structure, both of which are obtained by deleting the global average pooling layer and the fully connected layer of the last two layers of ResNet-50. The multi-modal feature fusion concatenates the features of the optical modality and the SAR modality and uses a linear layer for fusion.
[0052] Specifically, the multi-task classifier independently processes the features of the optical modality, SAR modality, and fusion modality, uses the softmax activation function to convert the output of the model into probabilities, and uses the cross-entropy loss function to evaluate the performance of the multi-branch deep feature fusion network. Finally, the total loss is calculated through weighted calculation.
[0053] Specifically, the D-S evidence theory calculates the mixed basic probability assignment function, selects the decision branch with the largest value, and determines whether the target exists.
[0054] Specifically, the speed of the target is obtained by calculating the actual distance of the target's movement from the longitude and latitude differences of the target at two fixed times, and then dividing by the fixed time. The heading is calculated by the arctangent function to obtain the angle between the position connection line of the target within the fixed time and the due north direction. The track connects the targets at different times in the direction of the heading and maps them to the situation base map stitched by optical images according to the coordinates.
[0055] The present invention also provides a multi-node multi-source heterogeneous information fusion and situation information generation system. The multi-node multi-source heterogeneous information fusion and situation information generation system can be implemented by executing the process steps of the multi-node multi-source heterogeneous information fusion and situation information generation method. That is, those skilled in the art can understand the multi-node multi-source heterogeneous information fusion and situation information generation method as a preferred implementation manner of the multi-node multi-source heterogeneous information fusion and situation information generation system.
[0056] Example 2 Example 2 is a preferred example of Example 1 As Figure 1 shown, according to a multi-node multi-source heterogeneous information fusion and situation information generation method provided by the present invention, it includes: Step 1: Use an improved YOLOV8 model based on the C2f-PKI module to perform target detection on visible light images to generate visible light image slices; use an improved YOLOV8 model based on the CBAM-SPPF module to perform target detection on SAR images to generate SAR image slices; Step 2: Input the visible light image slices and SAR image slices into a multi-branch deep feature fusion network to identify the fine types of targets; Step 3: Use the D-S evidence theory to perform decision-level fusion on the recognition results and electronic reconnaissance data to further obtain high-confidence recognition results; Step 4: Generate the entire situation base map, and based on the situation base map, extract the target heading and speed according to the recognition results to generate a track and map it onto the situation base map to generate situation information.
[0057] Furthermore, as Figure 2 shown, this embodiment provides an improved YOLOV8 model structure diagram based on the C2f-PKI module.
[0058] The improved YOLOV8 model based on the C2f-PKI module combines the C2f module for extracting and transforming input features with the PKI module for capturing various scale texture features to form the C2f-PKI module.
[0059] The PKI module is an Inception-style module that captures local information through a small kernel convolution and then uses a series of parallel depth convolutions to capture context information across multiple scales.
[0060] The C2f-PKI module uses global average pooling and 1×1 bar convolutions to enhance the features in the central region, capture long-range context information. Different from methods relying on large convolution kernels or dilated convolutions, it uses multiple depth convolution kernels of different sizes to obtain multi-scale texture features in different receptive fields without dilation.
[0061] Furthermore, as Figure 3 shown, this embodiment provides the detection results of an improved YOLOV8 model based on the C2f-PKI module for visible light images. The improved YOLOV8 model based on the C2f-PKI module was verified using a meter-level visible light remote sensing image dataset; the detection results of 16 randomly selected pictures were used, and the ship targets were marked with bounding boxes of different colors, and the type of the ship and the confidence of the target detection result were indicated above.
[0062] Furthermore, as Figure 4 shown, this embodiment provides an improved YOLOV8 model structure diagram based on the CBAM-SPPF module.
[0063] The improved YOLOV8 model based on the CBAM-SPPF module adds the CBAM module, which enhances the feature representation ability in the convolutional neural network, after the SPPF module that fuses large-scale global information. It enhances the network's feature extraction ability from both spatial and channel dimensions, suppresses the interference of background clutter, and strengthens the target feature extraction ability in SAR images, enabling the network to learn more target feature information and location information and improving the performance of the detection algorithm.
[0064] The CBAM-SPPF module is an attention mechanism module designed to enhance the feature representation ability in the convolutional neural network. It focuses on more informative features by integrating spatial and channel attention mechanisms, thereby improving the performance of the model.
[0065] The channel attention mechanism aims to analyze the correlation between feature channels, increase the proportion of effective channels, and reduce the proportion of ineffective channels, thereby improving the feature expression ability of the target. The CBAM-SPPF module aggregates the spatial information in the feature map through global average pooling and global max pooling operations, then processes these two pooling results through a multi-layer perceptron (MLP) with shared weights respectively, and finally adds these two signals and passes through a Sigmoid function to obtain the attention weight for each channel.
[0066] The spatial attention mechanism focuses on the effective information in the feature map by analyzing the internal spatial relationship. Channel attention focuses on the correlation between each channel to determine the weight of each channel, while spatial attention focuses on which spatial positions in the image are the most informative. After channel attention, the CBAM-SPPF module applies max pooling and average pooling (along the channel direction) to the feature map and stacks the results, then generates an attention map in the spatial dimension through a convolutional layer and a sigmoid activation function, which highlights which spatial positions are more important.
[0067] Furthermore, as Figure 5 shown, this embodiment provides the detection results of an improved YOLOV8 model based on the CBAM-SPPF module for SAR images. The improved YOLOV8 model based on the CBAM-SPPF module is verified using the RSDD-SAR public dataset. The detection results of 16 randomly selected images are used, and the ship targets are marked with blue bounding boxes, and "ship" and the confidence level of the target detection result are marked above.
[0068] Furthermore, as Figure 6As shown in the figure, this embodiment provides a structural diagram of a multi-branch deep feature fusion network. By inputting visible light target slices and SAR target slices, the network first extracts the features of the two different modal target image slices through two branches respectively, then fuses the features of the optical modality and the SAR modality to obtain a joint representation, and finally jointly calculates the loss through a multi-task classifier, and independently classifies the optical modality, the SAR modality and the fused features, so as to realize the recognition of the fine types of targets.
[0069] The network architecture can support the simultaneous training of the fusion classification task and the single-modal classification task. When the network receives only the input data of one modality, it outputs the classification result of the single-modal image. When the network inputs multi-modal images simultaneously, it extracts features respectively and performs deep feature fusion.
[0070] The feature extractor of the network uses an improved ResNet-50, removing the last two layers of ResNet-50, namely the global average pooling layer and the fully connected layer. The adjusted output is a high-dimensional feature map, retaining the spatial information of the image. Each modality configures an independent feature extractor. Although the structures are the same, they are independently trained for data of different modalities.
[0071] Batch normalization is performed on the output high-dimensional feature map to accelerate the training convergence of the model, alleviate the problems of gradient vanishing and gradient explosion to a certain extent, and process the value range differences of the high-dimensional feature maps on the same scale.
[0072] Feature fusion concatenates the features of the optical modality and the SAR modality, and uses a linear layer for fusion, finally reducing the dimension to the dimension of a single modality. The linear layer is a lightweight method that does not introduce additional model complexity and only performs feature transformation by learning a matrix.
[0073] The multi-branch deep feature fusion network processes the optical modality, the SAR modality and the fused modality features through independent classifiers for single-modal classification and fused-modal classification respectively. Each classifier uses the softmax activation function to convert the model output into probabilities, and finally calculates the total loss.
[0074] softmax activation function: Converts the output of each classifier into a probability distribution:
[0075] where zi is the output value of the i-th class, is the normalization factor for all classes.
[0076] Cross-entropy loss function: Cross-entropy is used to measure the difference between the predicted probability and the true label,
[0077] where, The one-hot encoding for the true label is the probability that sample i is predicted as class j. N represents the sample size, and C represents that there are C types in total for the target.
[0078] Multi-classifier loss design: For the optical modality loss LH and the SAR modality loss LS, optimize them separately, and use the fused modality loss LF to optimize the classification after combining the two modalities. The total loss is the weighted sum of the single-modal loss and the fused modality loss:
[0079] Construct an optical-SAR two-modal civilian ship image dataset using the MSAW remote sensing data of SpaceNet6. This dataset contains 11 different types of civilian ships. Randomly divide the dataset into a training set and a test set according to a ratio of 2:8. Resize the images uniformly to 256×128, set the output dimension of the feature converter to 2028, select SGD as the optimizer, set the initial learning rate to 0.003, set the number of training epochs to 50, and the training batch size to 1, and finally measure the recognition accuracy.
[0080] Table 1 Results of the multi-branch deep feature fusion network
[0081] It can be seen that in the single-task mode of blocking one of the branches, the recognition accuracy of the optical single modality is 91.84%, and the accuracy of the SAR single modality is 75.84%; in the multi-task mode of dual-modal input, the recognition accuracy of the dual-modal fusion branch is 93.12%, the recognition accuracy of the optical modality branch is 92.32%, and the recognition accuracy of the SAR modality branch is 80.48%. Using the multi-branch deep feature fusion network, the recognition accuracy of SAR and visible light images has been improved.
[0082] Furthermore, as Figure 7 shown, this embodiment provides a decision fusion flowchart based on the D-S evidence theory.
[0083] Set the decision space (target, non-target, uncertain), represented by (1, 0, -1); a types of information source types participating in the fusion: e1, e2,... e a ; the correct detection rates of the detection targets corresponding to a types of information sources: d1, d2,... d a , its actual meaning is the probability that the suspected targets detected are confirmed as targets, and define: Correct detection rate of detection target = Number of correctly detected targets / All detected targets = Detection rate / (Detection rate + False alarm rate); Detection rate = Number of correctly detected targets / All correct targets; False alarm rate = Number of detected false targets / Number of all correct targets; Target recognition confidence of source a: c1, c2,... c a , then the basic probability assignment function of the i-th source is:
[0084]
[0085]
[0086] The mixed basic probability assignment function is:
[0087]
[0088]
[0089] Where K is the conflict coefficient.
[0090] Through the above process, based on the target information of visible light images, SAR image target information, and electronic reconnaissance image data, the mixed basic probability assignment function is calculated, and the decision branch with the largest value is selected as the decision result of whether the target exists.
[0091] Furthermore, as Figure 8 shown, this embodiment provides a decision fusion result graph based on D-S evidence theory. Decision fusion is performed on the results of deep feature fusion of the optical-SAR two-modal civilian ship image dataset on Ascend 310, and the recognition rate of 95.08% and the classification results of different types of civilian ships are printed.
[0092] Furthermore, as Figure 9 shown, this embodiment provides a situation information generation result graph.
[0093] The situation base map of 1024×22551 pixels is formed by stitching visible light images of the port, and the longitude and latitude information of the lower left corner and the upper right corner of the base map is recorded.
[0094] When the target appears in the image at the first moment, its speed is default set to 0 because there is no position information of the previous moment to calculate the actual speed. For other moments (i.e., the i-th moment, i > 1), the speed is estimated according to the speed calculation formula. The calculation method is to first determine the travel distance d of the target between the i-th moment and the (i - 1)-th moment. That is, according to the longitude and latitude differences between two consecutive moments, the actual distance traveled by the entity during this period is obtained. The specific calculation method is as follows:
[0095] Among them, R is the radius of the earth, usually taken as 6371 Km, and are the latitudes at the (i - 1)-th moment and the i-th moment respectively (expressed in radians), and are the longitudes at the (i - 1)-th moment and the i-th moment respectively (expressed in radians).
[0096] After calculating the distance, we can calculate the speed of the current target. The specific calculation formula is:
[0097] Among them, v is the speed, d is the driving distance between two consecutive moments (the i-th moment and the (i - 1)-th moment), is the time interval between the i-th moment and the (i - 1)-th moment. Substituting the calculated d and into the speed formula can obtain the speed at the i-th moment.
[0098] Calculate the coordinate difference of the connection line: Let the coordinates at the (i - 1)-th moment be ( , ), and the coordinates at the i-th moment be ( , ). The specific calculation formula is as follows:
[0099]
[0100] Calculate the heading angle: Calculate the angle between the connection line and the positive y-axis (due north direction) through the arctangent function. Since the heading is the angle relative to the north, we use Δx and Δy to calculate the target heading θ.
[0101]
[0102] The calculation result of the heading angle is in degrees, and the result needs to be converted to an appropriate range (for example, 0° to 360°) to represent different directions. The specific conversion formula is: Heading = (θ + 360) mod 360 Connect the targets at different moments in the direction of the calculated heading and map them according to the coordinates onto the situation base map stitched by optical images to form situation information.
[0103] As Figure 10 shown, this embodiment also provides a multi-node multi-source heterogeneous information fusion and situation information generation system, which mainly includes: A visible light image processing module, which is used to deploy an improved YOLOV8 model based on the C2f-PKI module to perform target detection on visible light images and generate visible light image slices; The SAR image processing module is used to deploy the improved YOLOV8 model based on the CBAM-SPPF module to perform target detection on SAR images and generate SAR image slices; The fusion processing module is used to perform feature fusion on visible light image slices, SAR image slices and electronic reconnaissance data; The situation information generation module is used to extract the target heading and speed to generate a track and map it onto the situation base map to generate situation information.
[0104] Those skilled in the art know that in addition to implementing the systems, devices and their respective modules provided by the present invention in the form of pure computer-readable program codes, the method steps can be logically programmed to enable the systems, devices and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same program. Therefore, the systems, devices and their respective modules provided by the present invention can be regarded as a kind of hardware components, and the modules included therein for implementing various programs can also be regarded as the structures within the hardware components; the modules for implementing various functions can also be regarded as either software programs for implementing methods or structures within hardware components.
[0105] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A multi-node multi-source heterogeneous information fusion and situation information generation method, characterized in that Including: Step S1: Use an improved YOLOV8 model based on the C2f-PKI module to perform object detection on visible light images to generate visible light image slices; Step S2: Use an improved YOLOV8 model based on the CBAM-SPPF module to perform object detection on SAR images to generate SAR image slices; Step S3: Input the visible light image slices and SAR image slices into a multi-branch deep feature fusion network for object recognition; Step S4: Use the D-S evidence theory to perform decision-level fusion on the recognition results and electronic reconnaissance data to further obtain a high-confidence recognition result; Step S5: Generate the entire situation base map, extract the target heading and speed based on the target recognition result based on the situation base map, generate a track based on the target heading and speed, and map the track onto the situation base map to generate situation information; The improved YOLOV8 model based on the C2f-PKI module combines the C2f module for extracting and transforming input features with the PKI module for capturing texture features of various scales to form the C2f-PKI module. The C2f-PKI module performs feature extraction through global average pooling and multiple depth convolutional kernels of different sizes to achieve the purpose of detecting and capturing texture features of targets of different sizes; The improved YOLOV8 model based on the CBAM-SPPF module adds the CBAM module for enhancing the feature representation ability in the convolutional neural network after the SPPF module for fusing large-scale global information to form the CBAM-SPPF module, enhancing the network's feature extraction ability from both spatial and channel dimensions, so as to learn more target feature information and location information.
2. The multi-node multi-source heterogeneous information fusion and situation information generation method according to claim 1, wherein The multi-branch deep feature fusion network includes: two independent feature extractors, a feature fusion module, and three independent classifiers; Use two independent feature extractors to perform feature extraction on the visible light image slices and SAR image slices respectively to obtain optical modality features and SAR modality features; The optical modality features and SAR modality features obtain fused modality features through the feature fusion module; Use three independent classifiers to perform object recognition on the optical modality features, SAR modality features, and fused modality features respectively; Among them, the feature extractor is an improved ResNet-50; the improved ResNet-50 removes the last two layers of ResNet-50, namely the global average pooling layer and the fully connected layer, adjusts the output to a high-dimensional feature map that meets the preset requirements, and retains the spatial information of the image; The feature fusion module splices the optical modality features and SAR modality features, uses a linear layer for fusion, and finally reduces the dimension to the dimension of a single modality; The classifier uses the softmax activation function to convert the output into a probability.
3. The multi-node multi-source heterogeneous information fusion and situation information generation method according to claim 1, characterized in that The method further includes: obtaining the total loss through weighted calculation based on the cross-entropy loss function of each branch, and using the total loss to evaluate the performance of the multi-branch deep feature fusion network; Among them, the total loss includes: Among them, represents the optical modality cross-entropy loss; represents the SAR modality cross-entropy loss; represents the fusion modality cross-entropy loss.
4. The multi-node multi-source heterogeneous information fusion and situation information generation method according to claim 1, characterized in that The said step S4 includes: The decision-making space is set as target, non-target, and uncertain, represented by (1, 0, -1); there are a types of information sources participating in the fusion: e1, e2, … e a ; the correct detection rates of the a types of information sources corresponding to the detection target: d1, d2, … d a , its actual meaning is the probability that the detected suspicious targets are confirmed as targets, and it is defined as: Detection target accuracy = Number of correctly detected targets / Total number of detected targets = Detection rate / (Detection rate + False alarm rate); Detection rate = Number of correctly detected targets / Total number of correct targets; False alarm rate = Number of incorrectly detected targets / Total number of correct targets; The target recognition confidence of source a: c1, c2, … c a , then the basic probability assignment function of the i-th source is: The mixed basic probability assignment function is as follows: where K is the conflict coefficient; Through the above process, using the target information in visible light images, SAR image target information, and electronic reconnaissance image data, the mixed basic probability assignment function is calculated, and the decision branch with the largest value is selected as the decision result of whether the target exists.
5. The multi-node multi-source heterogeneous information fusion and situation information generation method according to claim 1, characterized in that, The step S5 includes: When the target appears in the image at the first moment, its speed is default set to 0; for the i-th moment, where i > 1, the speed is deduced according to the speed calculation formula; First, determine the driving distance d between the target at the i-th moment and the (i - 1)-th moment; the specific calculation method is as follows: where R is the radius of the Earth, and are the latitudes at the (i - 1)-th and i-th moments respectively, and are the longitudes at the (i - 1)-th and i-th moments respectively; After calculating the distance, calculate the speed of the current target. The specific calculation formula is: where v is the speed, d is the driving distance between the i-th moment and the (i-1)-th moment, and is the time interval between the i-th moment and the (i-1)-th moment; Calculate the coordinate difference of the connection line: Let the coordinates at the (i - 1)-th moment be ( , ), and the coordinates at the i-th moment be ( , ); The specific calculation formula is as follows: Calculate the heading angle: Calculate the angle between the connecting line and the positive y-axis direction through the arctangent function; since the heading is the angle relative to the north, use Δx and Δy to calculate the target heading θ; The calculation result of the heading angle is in degrees, and the result needs to be converted to a preset range; the specific conversion formula is: Heading = (θ + 360) mod 360 Connect the targets at different moments in the direction of the calculated heading and map them to the situation base map according to the coordinates to form situation information.
6. A multi-node multi-source heterogeneous information fusion and situation information generation system, characterized in that, It includes: Module M1: Use the improved YOLOV8 model based on the C2f-PKI module to perform target detection on visible light images to generate visible light image slices; Module M2: Use the improved YOLOV8 model based on the CBAM-SPPF module to perform target detection on SAR images to generate SAR image slices; Module M3: Input the visible light image slices and SAR image slices into a multi-branch deep feature fusion network for target recognition; Module M4: Use the D-S evidence theory to perform decision-level fusion on the recognition results and electronic reconnaissance data to further obtain high-confidence recognition results; Module M5: Generate the entire situation base map, extract the target heading and speed based on the situation base map according to the target recognition results, generate a track based on the target heading and speed, and map the track to the situation base map to generate situation information; The improved YOLOV8 model based on the C2f-PKI module combines the C2f module for extracting and transforming input features with the PKI module for capturing texture features of various scales to form the C2f-PKI module. The C2f-PKI module performs feature extraction through global average pooling and multiple depth convolution kernels of different sizes to achieve the purpose of detecting and capturing texture features of targets of different sizes; The improved YOLOV8 model based on the CBAM-SPPF module adds the CBAM module for enhancing the feature representation ability in the convolutional neural network after the SPPF module for fusing large-scale global information to form the CBAM-SPPF module, enhancing the network's feature extraction ability from both spatial and channel dimensions, so as to learn more target feature information and location information.
7. The multi-node multi-source heterogeneous information fusion and situation information generation system according to claim 6, characterized in that The multi-branch deep feature fusion network includes: two independent feature extractors, a feature fusion module, and three independent classifiers; Two independent feature extractors are used to extract features from visible light image slices and SAR image slices respectively to obtain optical modality features and SAR modality features; The optical modality features and SAR modality features obtain fused modality features through a feature fusion module; Three independent classifiers are used to classify the optical modality features, SAR modality features, and fused modality features respectively; Among them, the feature extractor is an improved ResNet-50; the improved ResNet-50 removes the last two layers of ResNet-50, namely the global average pooling layer and the fully connected layer, and adjusts the output to a high-dimensional feature map that meets the preset requirements, retaining the spatial information of the image; The feature fusion module splices the optical modality features and SAR modality features, uses a linear layer for fusion, and finally reduces the dimension to the dimension of a single modality; The classifier uses the softmax activation function to convert the output into probabilities.
8. The multi-node multi-source heterogeneous information fusion and situation information generation system according to claim 6, characterized in that The system further includes: calculating the total loss through weighted calculation based on the cross-entropy loss functions of each branch, and using the total loss to evaluate the performance of the multi-branch deep feature fusion network; Among them, the total loss includes: Among them, represents the optical modal cross-entropy loss; represents the SAR modal cross-entropy loss; represents the fusion modal cross-entropy loss.
9. The multi-node multi-source heterogeneous information fusion and situation information generation system according to claim 6, wherein The module M4 includes: Set the decision space as target, non-target, and uncertain, represented by (1, 0, -1); there are a types of information sources participating in the fusion: e1, e2, … e a ; the correct detection rates of the a types of information sources corresponding to the detection target: d1, d2, … d a , its actual meaning is the probability that the detected suspicious targets are confirmed as targets, and it is defined as: Detection target accuracy = number of correctly detected targets / total number of detected targets = detection rate / (detection rate + false alarm rate); Detection rate = number of correctly detected targets / total number of correct targets; False alarm rate = number of incorrectly detected targets / total number of correct targets; The target recognition confidence of source a: c1, c2, … c a , then the basic probability assignment function of the i-th source is: The mixed basic probability assignment function is: Among them, K is the conflict coefficient; Through the above process, based on the target information of visible light images, SAR image target information, and electronic reconnaissance image data, the mixed basic probability assignment function is calculated, and the decision branch with the largest value is selected as the decision result of whether the target exists.
10. The multi-node multi-source heterogeneous information fusion and situation information generation system according to claim 6, characterized in that, The module M5 includes: When the target appears in the image at the first moment, its speed is default set to 0; for the i-th moment, i > 1, the speed is then deduced according to the speed calculation formula; First, determine the driving distance d of the target between the i-th moment and the (i - 1)-th moment; the specific calculation method is as follows: where R is the radius of the Earth, and are the latitudes at the (i - 1)-th and i-th moments respectively, and are the longitudes at the (i - 1)-th and i-th moments respectively; After calculating the distance, calculate the speed of the current target, and the specific calculation formula is: where v is the speed, d is the driving distance between the i-th moment and the (i - 1)-th moment, and is the time interval between the i-th moment and the (i - 1)-th moment; Calculate the coordinate differences of the connection lines: Let the coordinates at the (i - 1)-th moment be ( , ), and the coordinates at the i-th moment be ( , ); The specific calculation formula is as follows: Calculate the heading angle: calculate the angle between the connecting line and the positive y-axis direction through the arctangent function; since the heading is the angle relative to the north, use Δx and Δy to calculate the target heading θ; The calculation result of the heading angle is in degrees, and the result needs to be converted to a preset range; the specific conversion formula is: Heading = (θ + 360) mod 360 Connect the targets at different moments in the direction of the calculated heading and map them to the situation base map according to the coordinates to form situation information.
Citation Information
Patent Citations
High-precision bunching type bistatic SAR space synchronization angle calculation method and device
CN111766581A
Vehicle-mounted phased array communication-in-motion antenna tracking method
CN118012133A
Modal information fusion method based on multi-source features of visible light and infrared images
CN118710516A
Cited By
SAR and optical image-based ship target fusion detection method in polar region navigation channel
CN121921498A