An AI-based multi-modal remote sensing image maritime target recognition method

By applying feature extraction and instance segmentation models for images with dense navigation and target overlap, combined with time series compensation method and motion compensation technology, the problem that deep learning models are difficult to identify and segment targets in dense ship overlap environments is solved, achieving higher recognition accuracy and robustness.

CN119274062BActive Publication Date: 2025-06-03SHANDONG JIMU SPACE TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411484960.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-23
Publication Date
2025-06-03
Estimated Expiration
2044-10-23

AI Technical Summary

Technical Problem

In dense ship navigation environments, multiple maritime targets may overlap in remote sensing images, making it difficult for deep learning models to accurately identify and segment targets, especially when the target size is small and relatively close, the overlap phenomenon significantly increases the difficulty of classification.

Method used

By feature extraction of images with dense navigation and target overlap, dense target test datasets and occlusion test datasets were constructed, segmentation masks were generated using instance segmentation models and segment consistency and occlusion resistance were evaluated. Combining the time series compensation method and motion compensation technology, the position and shape of the obstructed part are speculated to further distinguish and identify dense and overlapping targets.

Benefits of technology

The target recognition accuracy of deep learning models in dense ship overlap environments is improved, and the robustness and occlusion resistance of the model are enhanced, ensuring that maritime targets can be accurately identified in complex overlap and occlusion scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119274062B_ABST
    Figure CN119274062B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for identifying maritime targets in multi-modal remote sensing images based on AI, specifically relating to the technical field of maritime target identification; by preprocessing and feature extraction of remote sensing images of dense ship overlapping scenes, evaluating the consistency of target segmentation based on an instance segmentation model, and measuring the deviation between the detection box and the true target area in an occlusion test dataset, comprehensively analyzing the segmentation performance and anti-occlusion ability of the deep learning model; in the case of inaccurate target identification, the time series compensation method is adopted to combine the motion information of the front and rear frames to infer the target position and shape of the occluded part, so as to improve the segmentation and identification accuracy of the target in a dense ship overlapping environment, and finally achieve effective discrimination and accurate identification of overlapping targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of maritime target recognition, and particularly to an AI-based multi-modal remote sensing image maritime target recognition method. Background Art

[0002] The AI-based multi-modal remote sensing image maritime target recognition method collects multi-source image data of ocean areas through remote sensing technology, combines multi-modal image processing and AI analysis technology to accurately identify and detect maritime targets such as ships, oil tankers, icebergs, etc. This recognition method relies on multi-spectral, radar or high-resolution images obtained by satellites, drones or other aerial sensors, and through an AI model to fuse information of different modalities, automatically or semi-automatically perform target detection, classification and recognition. This technology is widely used in ocean monitoring, shipping management, maritime safety and environmental protection, improving the efficiency and accuracy of maritime target recognition.

[0003] In the prior art, when performing maritime target recognition based on AI-based multi-modal remote sensing images, first, preprocess the remote sensing images, including denoising, enhancing contrast, geometric correction, etc. Next, extract features related to the target through a feature extraction algorithm, such as shape, texture, spectral characteristics, etc. Then, use a deep learning model to analyze and classify these features to identify the targets in the images. However, in a dense ship navigation environment, multiple targets may appear in the same area of the remote sensing image and overlap with each other. This situation will make it difficult for the model to separately segment and identify the targets. Especially when the size of the targets is small and relatively close, the overlapping phenomenon will significantly increase the difficulty of classification. Long-distance shooting makes the target details blurred, and it is difficult to effectively separate overlapping targets through conventional methods. And when the model cannot accurately identify the targets, it may misidentify multiple targets as a single target, or completely fail to identify some of the targets. Summary of the Invention

[0004] The purpose of the present invention is to provide an AI-based multi-modal remote sensing image maritime target recognition method to solve the deficiencies in the background art.

[0005] To achieve the above purpose, the present invention provides the following technical solution: An AI-based multi-modal remote sensing image maritime target recognition method, comprising the following steps:

[0006] S1: Obtain several images with dense ship navigation and target overlap from the remote sensing images, and the images with target overlap include ships with different densities and different types of overlap situations;

[0007] S2: Extract features from the images with target overlap, and respectively construct a dense target test data set and an occlusion test data set based on the extracted feature data of the images with target overlap;

[0008] S3: Use the instance segmentation model to run the model on the dense target test dataset to generate a pixel-level segmentation mask for each target. For overlapping targets, calculate the average boundary coincidence rate between the segmentation mask of each target and its true boundary, and evaluate the consistency of target segmentation after analyzing the change of the average boundary coincidence rate.

[0009] S4: Run the model on the occlusion test dataset to generate a detection box and a segmentation mask for each target. Measure the deviation between the detection box of the occluded target and the true target area, and evaluate the occlusion resistance of the deep learning model after analyzing the abnormal change of the deviation within a fixed time period.

[0010] S5: Comprehensively analyze the consistency of target segmentation and the occlusion resistance of the deep learning model, and evaluate the accuracy of target recognition of the deep learning model in the dense ship overlapping environment.

[0011] S6: When the deep learning model is inaccurate in target recognition in the dense ship overlapping environment, based on the time series compensation method, use the motion compensation technology of the front and back frames to infer the position and shape of the occluded part in the current frame, and further distinguish and recognize dense and overlapping targets by combining the time information of consecutive frames.

[0012] Preferably, in S3, after analyzing the change of the average boundary coincidence rate in the target overlapping area, generate the average boundary coincidence rate fluctuation index. The method for obtaining the average boundary coincidence rate fluctuation index is as follows:

[0013] Obtain the average boundary coincidence rate IoU of N target overlapping areas, and construct the corresponding time series {IoU 1 , IoU 2 , …, IoU N}; and set the sliding window size to w, that is, the window contains w consecutive IoU values. For each position i in the sequence, the sliding window starting from IoU i contains the next w IoU values IoU i , IoU i+1 , …, IoU i+w-1 ; the IoU values in each window are calculated by weighted average. Set the weights of each window to w 1 , w 2 , …, w w , and the weighted average IoU within the sliding window is: where μ i is the weighted average value within the sliding window starting from the i-th IoU, IoU i+k-1 represents the k-th IoU value in the sliding window, and w kis the weight of the k-th IoU value, and the global IoU mean value μ is calculated global , and the expression is: For the weighted average IoU of each sliding window i , calculate its deviation Δ global from the global IoU mean value μ i , and the expression is: Δ i = μ i - μ global ; calculate the average boundary coincidence rate fluctuation index, and the expression is: In the formula, N - w + 1 is the total number of times calculated by the sliding window, and WD is the average boundary coincidence rate fluctuation index.

[0014] Preferably, in S4, after analyzing the abnormal change situation of the deviation within a fixed time period, a mask deviation anomaly index is generated, and the acquisition method of the mask deviation anomaly index is:

[0015] Collect the mask deviation data obtained within the W time period and establish a corresponding time series set {MD 1 , MD 2 , …, MD n}; where MD n is the mask deviation value at the n-th time point; calculate the mean value of the mask deviation sequence. For a one-dimensional mask deviation, the covariance matrix is a scalar; for a multi-dimensional mask deviation, the covariance matrix is a d×d matrix, where d is the dimension of the deviation vector, and the calculation formula of the covariance matrix Σ is: T is the matrix transpose. For the mask deviation of each time point MD i , calculate the distance D M MD i from each time point to the mean value μ, and the expression is: MD i - μ is the difference between the i-th deviation value and the mean value, Σ -1 is the inverse matrix of the covariance matrix, MD i - μ T is the transpose of the vector MD i - μ; compare the distance D M MD i obtained for each time point from the mean value μ with the standard distance in the preset normal state in the historical data, take the data points greater than or equal to the standard distance as abnormal points, and calculate the proportion of abnormal points in all data points, that is, calculate the mask deviation anomaly index.

[0016] Preferably, in S5, after comprehensively analyzing the consistency of the target segmentation and the anti-occlusion ability of the deep learning model, the accuracy of the deep learning model in target recognition in a dense ship overlapping environment is evaluated, specifically as follows:

[0017] Convert the average boundary coincidence rate fluctuation index and the mask deviation anomaly index into a first feature vector, use the first feature vector as the input of the machine learning model, and use the machine learning model to predict the accuracy value label of target recognition in a dense ship overlapping environment for each group of first feature vectors as the prediction target, and minimize the sum of the prediction errors of the accuracy value labels of target recognition in all dense ship overlapping environments as the training target, train the machine learning model until the sum of the prediction errors reaches convergence and then stop the model training, and determine the accuracy value of target recognition in a dense ship overlapping environment according to the model output result, where the machine learning model is a polynomial regression model.

[0018] Preferably, compare the obtained accuracy value of target recognition in a dense ship overlapping environment with the gradient standard threshold. The gradient standard threshold includes a first standard threshold and a second standard threshold, and the first standard threshold is less than the second standard threshold. Compare the accuracy value of target recognition in a dense ship overlapping environment with the first standard threshold and the second standard threshold respectively;

[0019] If the accuracy value of target recognition in a dense ship overlapping environment is greater than the second standard threshold, it indicates that the accuracy of the deep learning model in target recognition in a dense ship overlapping environment is high. At this time, generate a high-accuracy recognition signal and mark it as accurate recognition;

[0020] If the accuracy value of target recognition in a dense ship overlapping environment is greater than or equal to the first standard threshold and less than or equal to the second standard threshold, it indicates that the accuracy of the deep learning model in target recognition in a dense ship overlapping environment is average. At this time, generate a medium-accuracy recognition signal and mark it as medium-accurate recognition;

[0021] If the accuracy value of target recognition in a dense ship overlapping environment is less than the first standard threshold, it indicates that the accuracy of the deep learning model in target recognition in a dense ship overlapping environment is low. At this time, generate a low-accuracy recognition signal and mark it as inaccurate recognition.

[0022] Preferably, in S6, when the deep learning model is inaccurate in target recognition in a dense ship overlapping environment, based on the time series compensation method, use the motion compensation technology of the front and rear frames to infer the position and shape of the occluded part in the current frame, specifically as follows:

[0023] When the deep learning model is inaccurate in target recognition in a dense ship overlapping environment, that is, the accuracy value of target recognition in a dense ship overlapping environment is less than the first standard threshold, for the subsequent fixed time period, construct an image sequence {It-1 , I t , I t+1}; where I t is the current frame, I t-1 and I t+1 are the images of the previous frame and the next frame respectively, and there are multiple overlapping and dense targets in each frame of the image;

[0024] Position of the target in the previous frame: Let the position of the i-th target in the previous frame I t-1 be: p i,t-1 = (x i,t-1 , y i,t-1 ); Position of the target in the current frame: The position of the target in the current frame I t is: p i,t = (x i,t , y i,t ); Position of the target in the next frame: The position of the target in the next frame I t+1 is: p i,t+1 = (x i,t+1 , y i,t+1 ); Calculate its displacement through the positions of the previous and next frames. The displacement formula is: The displacement vector v i of the target is expressed as: v i = p i,t - p i,t-1 = (x i,t - x i,t-1 , y i,t - y i,t-1 ); Velocity formula: The velocity u i of the target is the displacement per unit time: where Δt is the time interval; Based on the displacement information of the previous and next frames, predict the position of the occluded part of the target in the current frame The expression is: The formula is based on the displacement information of the previous two frames to infer the position of the occluded part of the target in the current frame.

[0025] Preferably, when the target is partially occluded in the current frame I t , perform motion compensation through its motion trajectory in I t-1 and I t+1 to infer the position and shape of the occluded part. Specifically:

[0026] Within the time interval Δt, approximate the motion trajectory with the optical flow equation. The expression is: I x u + I y v + I t = 0; I x , I y are the brightness gradients of the image in the x and y directions, u and V are the horizontal and vertical components of the optical flow, representing the velocity components of the target, It is the change of image brightness in the time direction; based on the optical flow method, the motion compensation formula for the target is: where, Δv i =(u i , v i ) is the motion component obtained by optical flow calculation, which is used to compensate the position of the occluded part of the target; the shape S of the occluded part in the current frame i is estimated by the shapes of the front and back frames, and the expression is: S i,t-1 and S i,t+1 are the shapes of the target in the front and back frames, indicating that the average value of the shapes of the front and back frames is used to infer the shape of the occluded part in the current frame.

[0027] In the above technical solution, the technical effects and advantages provided by the present invention are:

[0028] 1. The present invention extracts features from dense navigation and target overlapping images, constructs a dense target test data set and an occlusion test data set, uses an instance segmentation model to generate a segmentation mask and evaluates the segmentation consistency and anti-occlusion ability, so as to comprehensively analyze the performance of the deep learning model. In addition, the average boundary coincidence rate fluctuation index and the mask deviation anomaly index are introduced, the target recognition accuracy of the model is evaluated through a polynomial regression model, and the accuracy of the model is classified according to the set threshold to generate a corresponding recognition signal, so that the deep learning model can more accurately handle complex overlapping and occlusion scenarios.

[0029] 2. When the model recognition is inaccurate, the present invention uses the time series compensation method, combined with the motion compensation technology of the front and back frames, to infer the position and shape of the occluded target in the current frame, further improving the accuracy of target recognition. The optical flow method is used for motion trajectory analysis and shape compensation to ensure that the target can be correctly recognized even in the case of dense targets and partial occlusion. Through dynamic time series information, the robustness of the deep learning model is effectively enhanced, and the accuracy of maritime target recognition is improved, especially in an environment with dense and severely overlapping ships. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.

[0031] Figure 1 is the method flow chart of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0033] For the embodiments, please refer to Figure 1 As shown in the figure, a method for identifying maritime targets in multi-modal remote sensing images based on AI according to this embodiment includes the following steps:

[0034] S1: Obtain several images in which ships are densely sailing and there are overlapping targets from the remote sensing images. The images with overlapping targets include ships with different densities and different types of overlapping situations.

[0035] S2: Extract features from the images with overlapping targets, and respectively construct a dense target test data set and an occlusion test data set based on the extracted feature data of the images with overlapping targets.

[0036] S3: Use an instance segmentation model to run the model on the dense target test data set to generate a pixel-level segmentation mask for each target. For overlapping targets, calculate the average boundary coincidence rate between the segmentation mask of each target and its true boundary, and evaluate the consistency of target segmentation after analyzing the change situation of the average boundary coincidence rate.

[0037] S4: Run the model in the occlusion test data set to generate a detection box and a segmentation mask for each target. Measure the deviation between the detection box of the occluded target and the true target area, and evaluate the anti-occlusion ability of the deep learning model after analyzing the abnormal change situation of the deviation within a fixed time period.

[0038] S5: After comprehensively analyzing the consistency of target segmentation and the anti-occlusion ability of the deep learning model, evaluate the accuracy of the deep learning model in identifying targets in a dense ship overlapping environment.

[0039] S6: When the deep learning model is inaccurate in identifying targets in a dense ship overlapping environment, based on the time series compensation method, use the motion compensation technology of the front and rear frames to infer the position and shape of the occluded part in the current frame, and further distinguish and identify dense and overlapping targets by combining the time information of consecutive frames.

[0040] Among them, in S1, several images in which ships are densely sailing and there are overlapping targets are obtained from the remote sensing images. The images with overlapping targets include ships with different densities and different types of overlapping situations.

[0041] In a dense navigation scenario, the density of ships can present different distribution patterns according to the navigation environment and the type of ships. According to the distribution density of ships, it can be divided into the following typical situations:

[0042] Low-density navigation: The distance between ships is large, and each ship usually occupies a relatively independent area in the image, and there is little overlap between targets. It usually appears in open sea areas, and the sailing routes between ships are scattered, which is suitable for long-distance ocean transportation scenarios. Although there is less target overlap, due to the small size of the ships and the long distance between them, the model may miss small targets due to resolution issues.

[0043] Medium-density navigation: The distance between ships is small, and some targets overlap. Especially when the ships are sailing in the same direction, the outlines of the front and rear ships may overlap. This is common in sea areas with many ships or at the entrances of large ports. Although the ships are densely distributed, there is still a certain amount of space for some targets to be independently visible. The overlap of some targets increases the difficulty of segmenting the overlapping area, especially when the hull structures of multiple ships overlap, which may lead to incorrect boundary division.

[0044] High-density navigation: The distance between ships is very close, even close to each other or staggered, and multiple ship targets in the image overlap seriously, forming a dense target group. It occurs in large ports, berthing areas, or scenes where fleets work together. This type of dense navigation also occurs in scenes of emergency rescue or collective fishing operations. In such dense distribution situations, the model needs to be able to segment and distinguish a large number of overlapping targets, which is extremely difficult. The model is prone to merging multiple targets into one large target or ignoring the existence of small targets.

[0045] In remote sensing images, the overlapping of ships will present different types of overlapping due to different ocean environments, shooting angles and ship navigation status. The following are several common target overlapping situations:

[0046] Overlap in the front and rear directions (longitudinal overlap): When the ships are sailing in the same direction, the projections of the front and rear ships may partially overlap in the image. The front ship may block part of the area of ​​the rear ship, especially the overlap between the bow and the stern. Usually occurs in ports or waterways, when multiple ships enter or leave the same route one after another, typical longitudinal overlap occurs in the image. Front-to-back overlap affects the boundary extraction of the target, especially when the ships are of similar size, the model may regard these ships as one target, resulting in misidentification or missed detection.

[0047] Overlap in the left - right direction (horizontal overlap): When ships are sailing side by side, the side profiles of the ships may overlap, making it difficult to accurately distinguish the boundaries of the targets. In this case, the horizontal overlap of the ships makes it difficult for the model to distinguish the specific boundaries and shapes of each ship. Horizontal overlap usually occurs when large fleets are sailing in formation or side by side. In port areas, such horizontal overlap is also likely to occur when ships are berthing. Horizontal overlap not only affects the extraction of boundaries but also interferes with the classification of ships (for example, different types of ships may be misclassified as the same type because of their similar appearances).

[0048] Partial overlap (local overlap): The ships are not completely overlapped, but some areas (such as the bow or stern) are blocked by other ships. This is usually because the ships are sailing at different depths or speeds, resulting in local overlap of some parts of the hulls. This situation usually occurs in an environment where ships are relatively freely distributed and at different speeds, especially in narrow waterways where ships are likely to overlap locally when shuttling. Partial overlap will cause the model to ignore the blocked parts during feature extraction, or separate some targets into multiple independent small targets, affecting the overall recognition accuracy.

[0049] Complete overlap (severe occlusion): Some ships are completely blocked by other targets, with only part of them visible or almost invisible. This situation is especially common in fleets at anchor or sailing in close proximity. In mooring areas, shipyards, or near large fleets with a large number of ships, the phenomenon of complete occlusion is more common. Large ships may completely block the medium - sized and small - sized ships behind them. Completely overlapped ships will cause the model to miss detecting targets or misclassify the hidden targets as the background. The model must use other means (such as context awareness or motion trajectory speculation) to infer the existence of the occluded targets.

[0050] Complex mixed overlap: It includes a mixed situation of longitudinal, horizontal, and local overlaps, where multiple targets occlude or partially occlude each other, resulting in an extremely complex recognition situation. This situation often appears in emergency rescue scenarios. For example, in bad weather conditions, when multiple rescue ships, fishing boats, or other ships are mixed in navigation, mooring, and support, the arrangement order and angles of the ships are variable. The complex mixed overlap situation makes it extremely difficult to segment and classify targets. The model needs to be able to handle multiple different overlap methods simultaneously and ensure that each target is correctly recognized and segmented.

[0051] S2: Extract features from the images with target overlap, and respectively construct a dense target test data set and an occlusion test data set based on the extracted feature data of the images with target overlap.

[0052] Before feature extraction of images, it is necessary to preprocess remote sensing images to ensure the quality and consistency of the data. Methods such as Gaussian filtering and median filtering are used to eliminate the noise in the remote sensing images, ensuring that the images are not interfered by noise when extracting features. Through histogram equalization or adaptive contrast enhancement (CLAHE) technology, the contrast of the images is enhanced, making the boundaries and shapes of the targets clearer. If there are multiple data sources (such as optical, radar images, etc.), different types of images are aligned and fused to increase the richness of the image information and enhance the expression of target features.

[0053] Geometric correction is performed on the remote sensing images to eliminate the geometric distortion caused by the sensor's perspective and terrain changes, aligning the images with the actual geographical coordinates. Based on the target distribution, the overlapping parts of the targets in the images are cropped out to form the region of interest (ROI), so as to focus on the feature extraction of the dense target areas.

[0054] Feature extraction is to obtain the information useful for target recognition and segmentation from the images. When processing images with overlapping targets, multiple features need to be extracted to ensure that the model can accurately distinguish the targets. Edge detection algorithms such as Sobel operator and Canny operator are used to extract the edges of the target objects. These edge information is very important for the segmentation and the distinction of overlapping targets. Calculate the contours of the ship targets to obtain geometric features such as the perimeter, area, and curvature of the contours of the targets. These features can help the model distinguish targets with similar sizes but different shapes. Analyze the texture information in the images through local binary pattern, especially in the target overlapping areas. The texture features can help the model distinguish the boundaries between different targets. Calculate the gray-level co-occurrence matrix of the target areas to obtain the texture features of the targets in the images, including contrast, correlation, entropy, etc. This helps to distinguish the texture differences of overlapping targets, especially when the appearances of the ships are similar.

[0055] Calculate the color distribution of the target areas. Especially in optical images, the color features of the ships can be used to distinguish targets with similar shapes but different colors. If the image is a multi-spectral remote sensing image, the spectral features of different targets can be obtained by analyzing the data of different bands (such as infrared, visible light). This is very effective for distinguishing overlapping targets with different optical properties. To process overlapping targets of different sizes and scales, multi-scale feature extraction methods can be used, combined with architectures such as feature pyramid network (FPN) to obtain the multi-scale information of the targets. This can ensure that even when there are large and small targets overlapping in the image, the model can still capture their respective features.

[0056] Based on the extracted feature data, first construct a dense target test dataset to test the performance of the model in dense target and complex overlapping scenarios. Use annotation tools to perform pixel-level annotation on each target. Especially in the target overlapping area, ensure that each target has accurate boundary labels, which can be used for training and testing the model. Define the criteria for dense targets. For example, when the distance between targets is less than a certain threshold (such as the target boundaries are less than 20 pixels apart), they are considered dense targets. Group the targets in the image according to the density to ensure that the dataset contains enough dense target instances. Extract ship navigation images with different densities from the processed remote sensing images according to different densities. Ensure that the dataset contains scenes from low density to high density to comprehensively evaluate the performance of the model. Save the feature extraction results (such as geometric features, texture features, depth features, etc.) of each dense target into the feature library for subsequent testing and evaluation. To enhance the robustness of the dataset, data augmentation operations such as rotation, scaling, and translation can be performed on the dense target images to generate more dense scene samples.

[0057] Construct an occlusion test dataset: Extract the ships in the image that are partially occluded by other targets, such as the bow or stern being blocked by other ships. This type of data is used to evaluate the model's ability to handle partially occluded targets. Select the ships that are completely occluded by other targets or only a small part is exposed. This type of data poses a higher requirement for the model's reasoning ability. Divide the targets in the image into partially occluded and severely occluded according to the degree of occlusion. Ensure that the dataset contains targets with different degrees of occlusion to comprehensively evaluate the model's performance in various occlusion situations. When sampling, consider various types of occlusion scenarios, such as targets being occluded by other ships, by waves or floating objects, etc., to ensure the diversity of the occlusion test dataset. Manually annotate the boundaries of the occluded targets, especially the partially occluded targets. Although some areas are occluded, complete segmentation labels still need to be generated for these targets to evaluate whether the model can accurately infer and recover the occluded targets. For severely occluded targets, generate special marks for the occluded areas and judge whether the model can handle these difficult scenarios in subsequent evaluations.

[0058] To enable the model to better learn the features of occluded targets, the dataset can be extended by simulating different degrees of occlusion to enhance the model's adaptability to occlusion phenomena. For example, by randomly adding occluders to the image to generate new training samples. Ensure that the feature sets of the occluded targets (such as geometric, texture, depth features) are consistent with those in the dense target test dataset so that the performance of the model in different scenarios can be compared during evaluation.

[0059] S3: Use the instance segmentation model to run the model on the dense target test dataset to generate a pixel-level segmentation mask for each target. For overlapping targets, calculate the average boundary coincidence rate between the segmentation mask of each target and its true boundary, and evaluate the consistency of target segmentation after analyzing the change of the average boundary coincidence rate.

[0060] Use common instance segmentation models such as Mask R-CNN, DeepLabV3, etc., and load the pre-trained model or the model specially trained for the sea dense target scenario. Input the constructed dense target test dataset into the model. The images in this dataset contain the situation where ships are densely sailing and there are overlapping targets. The model performs pixel-level segmentation on the input images to generate a segmentation mask for each target. The mask is a binary image, where each pixel value indicates whether the pixel belongs to a part of the target. When processing the overlapping target area, the model will try to distinguish the boundaries of different targets even if they overlap in some areas. Each target will have an independent segmentation mask.

[0061] Each image in the dense target test dataset has been manually annotated with accurate true segmentation boundaries. These annotation information can be used as the standard for evaluating the segmentation effect of the model. The segmentation mask of each target generated by the instance segmentation model represents the predicted segmentation result of the model for that target.

[0062] Calculate the average boundary coincidence rate IoU to measure the coincidence degree between the predicted mask and the true boundary. Its calculation formula is: In the formula, Aintersection is the intersection area of the predicted mask and the true boundary, and A union is their union area. For each target, calculate the intersection area and the union area between the predicted mask and the true boundary respectively. Divide the intersection area by the union area to get the IoU of each target. Calculate the IoU for all targets in each image, and take the average of all IoUs in the image to get the average boundary coincidence rate of the image.

[0063] For the overlapping target area, the model will generate overlapping part masks of multiple targets during the prediction process. It is necessary to calculate the IoU of these overlapping areas separately and focus on analyzing the boundary coincidence rate of these areas. In the overlapping area, due to the blurred target boundary, the IoU value may decrease. Therefore, the generation of the mask can be optimized by adjusting the threshold of the model or through post-processing (such as Soft-NMS).

[0064] After analyzing the change of the average boundary coincidence rate in the overlapping target area, generate the average boundary coincidence rate fluctuation index. The method for obtaining the average boundary coincidence rate fluctuation index is:

[0065] Obtain the average boundary coincidence rate IoU of N target overlapping regions and construct the corresponding time series {IoU 1 , IoU 2 ,..., IoU N}; and set the sliding window size to w (usually 3, 5, or 7), that is, the window contains w consecutive IoU values; for each position i in the sequence, the sliding window starting from IoU i contains the next w IoU values (IoU i , IoU i+1 ,..., IoU i+w-1 ); the IoU values in each window are calculated by weighted average; set the weights of each window to w 1 , w 2 ,..., w w (generally satisfying ). The weighted average IoU within the sliding window is: where μi is the weighted average value within the sliding window starting from the i-th IoU, IoU i+k-1 represents the k-th IoU value in the sliding window, and w k is the weight of the k-th IoU value. To measure the difference between the weighted average IoU within each window and the overall IoU mean, first calculate the global IoU mean μ global , and the expression is: For the weighted average IoU i of each sliding window, calculate its deviation Δ global from the global IoU mean μ i , and the expression is: Δ i = |μ i - μ global |; calculate the average boundary coincidence rate fluctuation index, and the expression is: In the formula, N - w + 1 is the total number of times the sliding window is calculated, because N - w + 1 weighted average IoUs of the windows can be calculated respectively from i = 1 to i = N - w + 1, and WD is the average boundary coincidence rate fluctuation index.

[0066] When the average boundary coincidence rate fluctuation index is larger, it means that in different target overlapping regions, the segmentation performance of the model fluctuates greatly. Specifically, it means that the model's segmentation is more accurate in some regions, while the segmentation effect is poor in other regions. This inconsistency reflects the lack of stability of the model in dealing with complex scenarios, and it may be disturbed when dealing with dense and overlapping targets, resulting in unstable boundaries of the segmentation results or difficulty in accurately distinguishing target overlapping regions. Therefore, the larger the fluctuation index, the worse the target segmentation consistency of the model, and larger segmentation errors are likely to occur when dealing with complex scenarios.

[0067] When the average boundary coincidence rate fluctuation index is smaller, it indicates that the segmentation performance of the model in different target overlapping regions is more consistent. Specifically, whether it is dense targets or overlapping regions, the segmentation effect of the model is relatively stable. A smaller average boundary coincidence rate fluctuation index shows that the model can better handle different scenarios and target types. A lower fluctuation index means that the segmentation results of the model for target boundaries are more accurate and consistent, and will not be significantly affected by target overlap or environmental complexity. Therefore, the smaller the fluctuation index, the better the target segmentation consistency, and the more reliable and robust the model's performance in various scenarios.

[0068] S4: Run the model on the occlusion test dataset to generate the detection boxes and segmentation masks for each target; measure the deviation between the detection boxes of the occluded targets and the true target regions, and analyze the abnormal changes of the deviation within a fixed time period to evaluate the anti-occlusion ability of the deep learning model.

[0069] Use the constructed occlusion test dataset and input it into the pre-trained deep learning models (such as MaskR-CNN, YOLO, etc.). These images contain different degrees of target occlusion, such as partial occlusion and complete occlusion.

[0070] The model first generates detection boxes (bounding boxes) for the targets in each image. These boxes are the preliminary localization results of the model for each target. Output of the detection box: The rectangular region of each target boxed by the model, which contains the main part of the target.

[0071] Based on the generated detection boxes, the model generates pixel-level segmentation masks (segmentation masks) for each target. The mask represents the specific shape and boundary of the target. For occluded targets, the model may only generate partial masks. The unoccluded parts will be clearer, while the masks of the occluded parts may be incomplete or have errors.

[0072] Each target in the occlusion test dataset has been manually annotated with accurate ground truth bounding boxes and ground truth segmentation masks. These annotated data are used as the benchmark for evaluating the model performance.

[0073] The bounding box deviation BBD is used to measure the deviation between the detection boxes generated by the model and the ground truth detection boxes. The calculation expression of the bounding box deviation is: In the formula, A pred is the area of the detection box generated by the model, and A true is the area of the ground truth detection box of the target. The larger the BBD value, the greater the difference between the detection box generated by the model and the ground truth box. Calculate the IoU of the predicted detection box and the ground truth detection box. The expression is: Where, A intersection is the overlapping area between the predicted bounding box and the ground truth bounding box, and A union is their union area. A lower IoU indicates that there is a large error in the generation of the detection bounding box for the occluded area. By comparing the differences between the generated segmentation mask and the ground truth mask, the accuracy of the model in segmentation under occlusion is measured. The mask deviation MD can be calculated by the pixel-level error, and the expression is: P pred is the number of pixels in the predicted segmentation mask, and P true is the number of pixels in the ground truth mask. A higher MD indicates a larger deviation between the mask generated by the model and the ground truth target mask.

[0074] After analyzing the abnormal changes in the deviation within a fixed time period, the mask deviation anomaly index is generated. The method for obtaining the mask deviation anomaly index is as follows:

[0075] Collect the mask deviation data obtained within the W time period and establish the corresponding time series set {MD 1 , MD 2 ,..., MD n}; where MD n is the mask deviation value at the nth time point; for the mask deviation sequence, calculate its mean value. The covariance matrix is used to measure the linear relationship between the mask deviations in each dimension. For the single-dimensional mask deviation, the covariance matrix is a scalar; for the multi-dimensional mask deviation, the covariance matrix is a d×d matrix (where d is the dimension of the deviation vector). The calculation formula for the covariance matrix ∑ is: T is the matrix transpose. For the mask deviation at each time point MD i , calculate the distance D M (MD i ) between each time point and the mean value μ, and the expression is: MD i -μ is the difference between the ith deviation value and the mean value, Σ -1 is the inverse matrix of the covariance matrix, and (MD i -μ) T is the transpose of the vector MD i -μ; the larger the distance, the farther the mask deviation at this time point is from the center of the overall distribution, indicating that this deviation point may be an outlier. Compare the distance D M (MD i ) between each obtained time point and the mean value μ with the standard distance under the preset normal state in the historical data, and regard the data points greater than or equal to the standard distance as outliers, and calculate the proportion of outliers in all data points, that is, calculate the mask deviation anomaly index.

[0076] When the mask deviation anomaly index is larger, it indicates that there are significant fluctuations in the segmentation performance of the deep learning model when dealing with occluded targets, and there are many abnormal situations. This means that the prediction of the segmentation boundary or shape of the target by the model in the occlusion scenario has a large difference from the true value, and false detection or missed detection is likely to occur. When the deviation anomaly index is large, the model may not be able to effectively handle complex occlusion situations, showing poor anti-occlusion ability. Especially when the target is partially occluded or completely occluded, the deviation of the segmentation mask of the model increases significantly, affecting the robustness and reliability of the model.

[0077] When the mask deviation anomaly index is smaller, it indicates that the segmentation performance of the model is relatively stable under different occlusion degrees, with small deviation fluctuations and no obvious abnormal situations. A smaller anomaly index indicates that the model can maintain a high accuracy when dealing with occlusion scenarios, and the segmentation of the shape and boundary of the occluded target is closer to the true value, showing strong anti-occlusion ability. Even when the target is partially occluded or interfered by a complex background, the model can still effectively identify the complete contour of the target and maintain a low segmentation error, indicating its good robustness in complex scenarios.

[0078] S5: After comprehensively analyzing the consistency of target segmentation and the anti-occlusion ability of the deep learning model, evaluate the accuracy of target recognition of the deep learning model in the dense ship overlapping environment.

[0079] Convert the average boundary coincidence rate fluctuation index and the mask deviation anomaly index into a first feature vector, and use the first feature vector as the input of the machine learning model. The machine learning model takes the prediction of the accuracy value label of target recognition in the dense ship overlapping environment for each group of first feature vectors as the prediction target, and takes minimizing the sum of the prediction errors of the accuracy value labels of all target recognitions in the dense ship overlapping environment as the training target. Train the machine learning model until the sum of the prediction errors reaches convergence and then stop the model training. Determine the accuracy value of target recognition in the dense ship overlapping environment according to the model output result, where the machine learning model is a polynomial regression model.

[0080] The method for obtaining the accuracy value of target recognition in the dense ship overlapping environment is: obtain the corresponding function expression from the first feature vector training data of the trained machine learning model: WR = F(WD, EQ); where F is the output function of the model, WD is the average boundary coincidence rate fluctuation index, EQ is the mask deviation anomaly index, and WR is the accuracy value of target recognition in the dense ship overlapping environment.

[0081] Compare the accuracy value of target recognition in the dense ship overlapping environment obtained with the gradient standard thresholds, where the gradient standard thresholds include a first standard threshold and a second standard threshold, and the first standard threshold is less than the second standard threshold, and compare the accuracy value of target recognition in the dense ship overlapping environment with the first standard threshold and the second standard threshold respectively;

[0082] If the accuracy value of target recognition in the dense ship overlapping environment is greater than the second standard threshold, it indicates that the deep learning model has high accuracy in target recognition in the dense ship overlapping environment. At this time, generate a high-accuracy recognition signal and mark it as accurate recognition;

[0083] If the accuracy value of target recognition in the dense ship overlapping environment is greater than or equal to the first standard threshold and less than or equal to the second standard threshold, it indicates that the deep learning model has general accuracy in target recognition in the dense ship overlapping environment. At this time, generate a medium-accuracy recognition signal and mark it as medium-accurate recognition;

[0084] If the accuracy value of target recognition in the dense ship overlapping environment is less than the first standard threshold, it indicates that the deep learning model has low accuracy in target recognition in the dense ship overlapping environment. At this time, generate a low-accuracy recognition signal and mark it as inaccurate recognition.

[0085] S6: When the deep learning model fails to accurately recognize targets in the dense ship overlapping environment, based on the time series compensation method, use the motion compensation technology of the front and back frames to infer the position and shape of the occluded part in the current frame, and further distinguish and recognize dense and overlapping targets by combining the time information of consecutive frames.

[0086] When the deep learning model fails to accurately recognize targets in the dense ship overlapping environment, that is, the accuracy value of target recognition in the dense ship overlapping environment is less than the first standard threshold. For the subsequent fixed time period, construct an image sequence {I t-1 , I t , I t+1}; where I t is the current frame, and I t-1 and I t+1 are the images of the previous frame and the next frame respectively. There are multiple overlapping and dense targets (such as ships) in each frame of the image, and the motion of these targets can be modeled by the displacement between frames.

[0087] Position of the target in the previous frame: Let the position of the i-th target in the previous frame I t-1 be: p i,t-1 = (x i,t-1 , y i,t-1 ); Position of the target in the current frame: The position of this target in the current frame I t is: p i,t=(x i,t , y i,t ); The position of the target in the next frame: The position of the target in the next frame I t+1 is: p i,t+1 =(x i,t+1 , y i,t+1 ); The movement of the target can be represented by speed or displacement. Assuming that the movement speed of the target is constant, its displacement can be calculated from the positions of the previous and next frames.

[0088] Displacement formula: The displacement vector v i of the target is expressed as: v i =p i,t -p i,t-1 =(x i,t -x i,t-1 , y i,t -y i,t-1 ); Speed formula: The speed u i of the target is the displacement per unit time: where Δt is the time interval; Based on the displacement information of the previous and next frames, predict the position of the occluded part of the target in the current frame The expression is: The formula is based on the displacement information of the previous two frames to infer the position of the occluded part of the target in the current frame.

[0089] When the target is partially occluded in the current frame I t , its position and shape of the occluded part can be inferred by its movement trajectory in I t-1 and I t+1 for motion compensation, specifically:

[0090] Within a small time interval Δt, assuming that the brightness of the target remains unchanged, the optical flow equation is used to approximate the motion trajectory, and the expression is: I x u + I y v + I t = 0; I x , I y are the brightness gradients of the image in the x and y directions, u, v are the horizontal and vertical components of the optical flow, representing the velocity components of the target, and I t is the change in the image brightness in the time direction. Based on the optical flow method, the motion compensation formula of the target is: where, Δv i =(u i , v i ) is the motion component obtained by optical flow calculation, used to compensate the position of the occluded part of the target; The shape of the target can be inferred from the continuous change between frames. Assuming that the shape of the target remains consistent in a short time, the shape S of the occluded part in the current framei It can be estimated by the shapes of the front and rear frames, and the expression is: S i,t-1 and S i,t+1 are the shapes of the target in the front and rear frames, indicating that the average value of the shapes of the front and rear frames is used to infer the shape of the occluded part in the current frame.

[0091] Through the time information of consecutive frames, the motion trajectories of each target can be tracked to determine its position in each frame. Using trajectory tracking methods (such as the Kalman filter) can predict the position and shape of the occluded target, specifically:

[0092] The trajectory tracking formula (Kalman filter prediction model) is: K is the Kalman gain matrix, z i,t is the measured target position, and H is the measurement matrix.

[0093] Based on the motion compensation information of the front and rear frames, the time series analysis method is used to distinguish dense and overlapping targets: using a time series prediction model (such as the ARIMA model) to predict the future position of the target and distinguish overlapping targets. Predicting the next position of the target through the time series model further helps to distinguish different overlapping targets.

[0094] In this application, through the time series compensation method, the information of the front and rear frames can be used to infer the position and shape of the occluded target in the current frame. The specific steps include compensating through the positions and motion trajectories of the front and rear frames, calculating the motion components of the target by the optical flow method, combining trajectory tracking methods such as the Kalman filter to infer the motion of the occluded target, and using the time series analysis model to distinguish and predict dense and overlapping targets. Through these methods, the accuracy of target recognition of the deep learning model in a dense ship overlapping environment can be effectively improved.

[0095] In this embodiment, after obtaining images in the remote sensing images where vessels are densely navigating and there are overlapping targets, first, feature extraction is performed on these images, and based on the feature data, a dense target test dataset and an occlusion test dataset are constructed. Then, an instance segmentation model is used to run on the dense target test dataset to generate a pixel-level segmentation mask of the target, and the average boundary coincidence rate between the segmentation mask and the true boundary is calculated to evaluate the consistency of target segmentation. Next, in the occlusion test dataset, detection boxes and segmentation masks of the targets are generated, the deviation between the detection boxes and the true regions is measured, and the abnormal changes in the deviation are analyzed to evaluate the occlusion resistance of the model. On this basis, the segmentation consistency and the occlusion resistance are comprehensively analyzed to evaluate the accuracy of target recognition of the deep learning model in the dense ship overlapping environment. If the model recognition is inaccurate, then based on the time series compensation method, the motion compensation technology of the previous and subsequent frames is used to infer the position and shape of the occluded part in the current frame, and combined with the time information of consecutive frames, the dense and overlapping targets are further distinguished and recognized.

[0096] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0097] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0098] It should be understood that the term "and / or" in this text is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this text generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.

[0099] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this text can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0100] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application.

Claims

1. A method for identifying marine targets in multimodal remote sensing images based on AI, characterized by: The steps include: S1: acquiring an image of a plurality of ships sailing densely and with overlapping targets from a remote sensing image, wherein the image of overlapping targets includes ships of different densities and different types of overlapping situations; S2: Extract features from the target overlapping images, and construct a dense target test dataset and an occlusion test dataset based on the extracted target overlapping image feature data; S3: Use the deep learning model to run on the dense object test dataset to generate pixel-level segmentation masks for each object; For overlapping targets, the average boundary overlap ratio between the segmentation mask of each target and its true boundary is calculated. After analyzing the change of the average boundary overlap ratio, the consistency of the target segmentation is evaluated. After analyzing the change of the average boundary overlap rate of the target overlapping area, the average boundary overlap rate fluctuation index is generated. The method for obtaining the average boundary overlap rate fluctuation index is: Construct a time series of average boundary coincidence rates in target overlap regions ; and set the sliding window size to w, that is, the window contains w consecutive IoU values; for each time point i in the sequence, from The initial sliding window contains the next w IoU values ; The IoU value in each window is calculated by weighted average; the weight of each window is set to , the weighted average IoU in the sliding window is: ;in, is the weighted average in the sliding window starting from the i-th IoU, represents the kth IoU value in the sliding window, is the weight of the kth IoU value, and calculates the global IoU mean , the expression is: ; Weighted average for each sliding window , calculate its mean IoU with the global IoU Deviation , the expression is: ; Calculate the average boundary overlap rate fluctuation index, the expression is: ; In the formula, is the total number of sliding window calculations, is the average boundary coincidence rate fluctuation index; S4: Run the deep learning model in the occlusion test dataset to generate the detection box and segmentation mask of each target; measure the deviation between the occluded target detection box and the real target area, and analyze the abnormal changes of the deviation within a fixed time period to evaluate the anti-occlusion ability of the deep learning model; S5: After comprehensively analyzing the consistency of target segmentation and the anti-occlusion ability of the deep learning model, the accuracy of the deep learning model in target recognition in a dense ship overlapping environment is evaluated; S6: When the deep learning model cannot accurately identify targets in a dense and overlapping ship environment, the time series compensation method is used to infer the position and shape of the occluded part in the current frame through the motion compensation technology of the previous and next frames. By combining the time information of consecutive frames, the dense and overlapping targets can be further distinguished and identified.

2. The method for identifying marine targets in multimodal remote sensing images based on AI according to claim 1, characterized in that: In S4, the abnormal change of the deviation within a fixed time period is analyzed to generate a mask deviation abnormality index. The method for obtaining the mask deviation abnormality index is: Collect the mask deviation data obtained within the W time period and establish the corresponding time series set ;in is the mask deviation value at the nth time point; for the mask deviation sequence, its mean is calculated. For a single-dimensional mask deviation, the covariance matrix is ​​a scalar; for a multi-dimensional mask deviation, the covariance matrix is ​​a d×d matrix, where d is the dimension of the deviation vector, and the covariance matrix The calculation formula is: ; T is the matrix transpose, and for the mask deviation value at the i-th time point , calculate its distance from the mean μ , the expression is: ; is the difference between the mask deviation value and the mean at the i-th time point, is the inverse of the covariance matrix, is a vector Transpose; the distance between the mask deviation value obtained at the i-th time point and the mean μ The mask deviation anomaly index is calculated by comparing the data points with the standard distance under normal conditions preset in the historical data and taking the data points greater than or equal to the standard distance as abnormal points. The proportion of abnormal points to all data points is calculated.

3. The method for identifying marine targets using multimodal remote sensing images based on AI according to claim 2, characterized in that: In S5, after comprehensively analyzing the consistency of target segmentation and the anti-occlusion ability of the deep learning model, the accuracy of the deep learning model in target recognition in a dense ship overlapping environment is evaluated. Specifically: The average boundary overlap rate fluctuation index and the mask deviation anomaly index are converted into the first eigenvector, and the first eigenvector is used as the input of the machine learning model. The machine learning model uses each group of first eigenvectors to predict the accuracy value label of target recognition in a dense ship overlapping environment as the prediction target, and takes minimizing the sum of prediction errors of the accuracy value labels of target recognition in all dense ship overlapping environments as the training target. The machine learning model is trained until the sum of prediction errors converges, and the model training is stopped. The accuracy value of target recognition in a dense ship overlapping environment is determined according to the model output results, wherein the machine learning model is a polynomial regression model.

4. The method for identifying marine targets using multimodal remote sensing images based on AI according to claim 3 is characterized by: The acquired accuracy value of target recognition in the dense ship overlapping environment is compared with the gradient standard threshold, the gradient standard threshold includes a first standard threshold and a second standard threshold, and the first standard threshold is less than the second standard threshold, and the accuracy value of target recognition in the dense ship overlapping environment is compared with the first standard threshold and the second standard threshold respectively; If the accuracy value of target recognition in the dense ship overlapping environment is greater than the second standard threshold, it means that the deep learning model has high accuracy in target recognition in the dense ship overlapping environment, and a high-accuracy recognition signal is generated and marked as accurate recognition; If the accuracy value of target recognition in the dense ship overlapping environment is greater than or equal to the first standard threshold and less than or equal to the second standard threshold, it means that the accuracy of target recognition of the deep learning model in the dense ship overlapping environment is average, and a medium accuracy recognition signal is generated and marked as medium accuracy recognition; If the accuracy value of target recognition in a dense ship overlapping environment is less than the first standard threshold, it means that the accuracy of target recognition of the deep learning model in a dense ship overlapping environment is low. At this time, a low-accuracy recognition signal is generated and marked as an inaccurate recognition.

5. The method for identifying marine targets using multimodal remote sensing images based on AI according to claim 4, characterized in that: In S6, when the deep learning model cannot accurately identify the target in a dense ship overlapping environment, the position and shape of the blocked part in the current frame are inferred through the motion compensation technology of the previous and next frames based on the time series compensation method, specifically: When the deep learning model cannot accurately identify the target in the dense ship overlapping environment, that is, the accuracy value of the target recognition in the dense ship overlapping environment is less than the first standard threshold, for the subsequent fixed time period, a continuous frame image sequence is constructed. ;in is the current frame, and They are the images of the previous frame and the next frame respectively. There are multiple overlapping and dense objects in each frame. The position of the target in the previous frame: Let the i-th target be in the previous frame The position in is: ; The position of the target in the current frame: The target is in the current frame The position in is: ; The target's position in the next frame: The target is in the next frame The position in is: ; Calculate its displacement through the positions of the previous and next frames. The displacement formula is: the displacement vector of the target It is expressed as: ; Speed ​​formula: target speed is the displacement per unit time: ; Where Δt is the time interval; Based on the displacement information of the previous and next frames, the position of the occluded part of the target in the current frame is predicted , the expression is: ; The formula is based on the displacement information of the previous two frames to infer the position of the obscured part of the target in the current frame.

6. The method for identifying marine targets in multimodal remote sensing images based on AI according to claim 5, characterized in that: When the target is in the current frame When it is partially blocked, and The motion trajectory in is used to perform motion compensation, thereby inferring the position and shape of the blocked part, specifically: In the time interval Δt, the motion trajectory is approximated by the optical flow equation, which is expressed as: ; , is the brightness gradient of the image in the x and y directions, , are the horizontal and vertical components of the optical flow, representing the velocity component of the target, is the change of image brightness in the time direction; based on the optical flow method, the motion compensation formula of the target is: ;in, It is the motion component calculated by optical flow, which is used to compensate for the position of the occluded part of the target; the shape of the occluded part in the current frame The shape of the previous and next frames is estimated, and the expression is: ; Indicates that the shape of the occluded part in the current frame is estimated by using the average value of the shapes of the previous and next frames. and is the shape of the object in the previous and next frames.

Citation Information

Patent Citations

  • Object tracking method based on core

    CN101251928A

  • A remote sensing image ship integrated recognition method based on deep learning

    CN109583425A

  • Multi-armored target online tracking method and device and storage medium

    CN117437567A