Moving target segmentation method and system based on satellite video

By adopting shape prior extraction and semantic affinity constraint processing methods in satellite videos, combined with spatiotemporal semantic relationship modeling, the problem of insufficient segmentation efficiency and accuracy of motion targets in the prior art is solved, and adaptability and efficient segmentation effects to complex backgrounds and lighting changes are achieved.

CN120070507AActive Publication Date: 2025-05-30TECH & ENG CENT FOR SPACE UTILIZATION CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510525509.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The prior art cannot efficiently and accurately realize the segmentation of moving targets in satellite videos, and there are problems such as complex calculation processes, insufficient refinement of the target segmentation edges, and insufficient utilization of timing information mining.

Method used

A motion target segmentation method based on satellite video is adopted, through shape prior extraction and semantic affinity constraint processing, combined with spatiotemporal semantic relationship modeling, the edge feature expression ability is significantly enhanced and the robustness of semantic learning is improved.

Benefits of technology

It realizes accurate segmentation of moving targets in satellite videos, can effectively adapt to complex backgrounds and lighting changes, and improves the efficiency and accuracy of moving target segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070507A_ABST
    Figure CN120070507A_ABST
Patent Text Reader

Abstract

The invention provides a moving target segmentation method and system based on a satellite video, and the method comprises the steps: firstly, determining a reference frame image based on a to-be-segmented query frame image, and determining a corresponding real mask image according to the reference frame image; then, shape prior extraction is carried out according to the real mask image, and an edge mask image is determined; and finally, performing feature extraction, space-time semantic relationship modeling and semantic affinity constraint processing on the query frame image, the reference frame image, the real mask image and the edge mask image through the moving target segmentation model, and outputting to obtain a target mask image for representing a moving target segmentation result. Thus, the method can realize efficient interactive learning of intra-frame and inter-frame semantic information of the satellite video, improve the robustness and accuracy of semantic learning of a moving target segmentation model, and significantly enhance the edge feature expression ability, thereby effectively improving the efficiency and accuracy of satellite video moving target segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of satellite video processing and analysis, and in particular to a method and system for segmenting moving targets based on satellite video. Background Art

[0002] Video satellites are a new type of earth observation satellite. Compared with traditional earth observation satellites, video satellites can continuously observe a specific area, thereby obtaining remote sensing images with high spatial resolution and high spectral resolution, and can also obtain remote sensing images with high temporal resolution. The intelligent processing and analysis of satellite video can automatically extract and analyze the information in the scene of interest, providing important technical support for practical application scenarios such as disaster monitoring, ocean monitoring, and ecosystem disturbance monitoring, and playing an important role in them. Among them, satellite video target segmentation can achieve the fine segmentation of the target of interest in a time-continuous scene. It can not only locate the position of the target, but also obtain additional fine description information to provide support for more refined analysis and applications.

[0003] For satellite video, the related technologies of moving target segmentation are mainly divided into two categories: unsupervised and semi-supervised. Specifically, the method based on optical flow combined with unsupervised learning can capture the motion characteristics of the target in the time dimension through optical flow information. However, this type of method relies highly on the accuracy of optical flow estimation. When complex scenes (such as occlusion, illumination change, or rapid target movement) appear in satellite video, the optical flow estimation is likely to fail, resulting in a decline in segmentation performance. The supervised learning method based on static image segmentation, such as OSVOS (One-Shot Video Object Segmentation), this type of method relies on the annotation of the first frame for model fine-tuning, and only processes subsequent frames frame by frame using static features, ignoring the temporal information of the video, and it is difficult to effectively handle target appearance changes or background dynamic interference. In addition, these methods generally have problems of large computational overhead and insufficient real-time performance, which limit their application on satellite payloads with limited resources.

[0004] It can be seen that the related methods for segmenting moving targets based on satellite video have problems such as complex calculation processes, insufficient refinement of target segmentation edges, and insufficient mining and utilization of temporal information, resulting in the inability to efficiently and accurately achieve moving target segmentation. Summary of the Invention

[0005] The technical problem to be solved by the present invention is the problem of being unable to efficiently and accurately achieve moving target segmentation based on satellite video.

[0006] To solve the above technical problem, the present invention provides a method and system for segmenting moving targets based on satellite video, and specifically adopts the following technical solutions: In a first aspect, the present invention provides a method for segmenting moving objects based on satellite video, including: First, determine a reference frame image based on the query frame image to be segmented in the satellite video, and determine the corresponding ground truth mask image according to the reference frame image. Wherein, the reference frame image is a frame image in the satellite video that has been object-segmented, and the ground truth mask image is used to represent the position and shape of the moving objects in the reference frame image. Then, perform shape prior extraction according to the ground truth mask image to determine an edge mask image, and the edge mask image is used to represent the edge shape of the moving objects. Secondly, perform feature extraction, spatio-temporal semantic relationship modeling, and semantic affinity constraint processing on the query frame image, the reference frame image, the ground truth mask image, and the edge mask image through a moving object segmentation model, and output a target mask image, which is used to represent the moving object segmentation result.

[0007] This method significantly enhances the edge feature expression ability through a shape prior fusion method. At the same time, aiming at the interference of static objects on the semantic learning of the moving object segmentation model, this method adopts foreground-background semantic affinity constraints to improve the robustness and accuracy of the semantic learning of the moving object segmentation model. In this way, through feature extraction, spatio-temporal semantic relationship modeling, and semantic affinity constraint processing by the moving object segmentation model, accurate segmentation of moving objects in satellite video can be achieved, and it can effectively adapt to complex backgrounds and illumination changes, thereby effectively improving the efficiency and accuracy of moving object segmentation in satellite video.

[0008] Combined with the first aspect, in an alternative implementation, the above-mentioned moving object segmentation model includes: a feature extraction module, a spatio-temporal semantic relationship modeling module, and a semantic affinity constraint segmentation decoder head module. Among them, the feature extraction module can be used to extract the features for representing the object spatial information and edge characteristics in the query frame image, the reference frame image, the ground truth mask image, and the edge mask image respectively, to obtain the corresponding first extracted feature, second extracted feature, third extracted feature, and fourth extracted feature, and splice the first extracted feature, second extracted feature, third extracted feature, and fourth extracted feature to determine the target spliced feature. The spatio-temporal semantic relationship modeling module can be used to add position encoding, perform spatio-temporal semantic relationship modeling encoding, and perform target prediction decoding according to the target spliced feature to determine the encoded feature and the decoded feature. The semantic affinity constraint segmentation decoder head module can be used to perform mask prediction according to the encoded feature and the decoded feature to determine the target mask image.

[0009] In this implementation, the moving object segmentation model can accurately extract features that characterize the spatial information and edge characteristics of the query frame image, reference frame image, ground truth mask image, and edge mask image through the feature extraction module. The spatio-temporal semantic relationship modeling module constructs spatio-temporal semantic relationships, enabling the moving object segmentation model to effectively process the spatio-temporal information of moving objects, solving the problems of diversity and uncertainty in moving object segmentation, and thus improving the accuracy of moving object segmentation. The semantic affinity constraint segmentation head module can achieve robust representation learning of semantic consistency, effectively improving the robustness and accuracy of the moving object segmentation model for semantic learning of moving objects.

[0010] Combined with the first aspect, in an alternative implementation, the above spatio-temporal semantic relationship modeling module includes: a position encoding module, a Transformer encoder, and a Transformer decoder. Among them, the position encoding module can be used to add position encoding to the target concatenated features to obtain position-encoded concatenated features. The Transformer encoder can be used to perform spatio-temporal semantic relationship modeling encoding based on the position-encoded concatenated features through a multi-head self-attention module and a fully connected feed-forward neural network, determine and output encoded features. The Transformer decoder can be used to perform target prediction decoding based on the encoded features through a multi-head self-attention module, a multi-head cross-attention module, and a fully connected feed-forward neural network, determine and output decoded features.

[0011] In this implementation, the position encoding is introduced through the position encoding module to provide position information, enabling the Transformer encoder to effectively capture the spatio-temporal dependence relationships between pixels in the input frame through spatio-temporal semantic relationship modeling encoding. It allows the moving object segmentation model to model the dynamic changes between objects within multiple time steps, improving the accuracy of object detection and tracking. The Transformer encoder can also learn and establish the spatio-temporal correspondence relationships of moving objects between the query frame image and the reference frame image through spatio-temporal semantic relationship modeling encoding. The Transformer decoder can predict the spatial position information of moving objects in each frame image through target prediction decoding, thereby enhancing the spatio-temporal reasoning ability of the moving object segmentation model. In this way, not only can the transformation of moving objects between different time steps be accurately captured, but also by learning the robust representation of the target object, the moving object segmentation model can improve the accuracy of moving object segmentation when facing the confusion problem of similar objects.

[0012] In combination with the first aspect, in an alternative implementation, the above-mentioned semantic affinity constraint segmentation decoder head module includes: a target affinity module and a target mask prediction module. Among them, the target affinity module can be used to determine and output a self-attention feature map based on the decoded features and the encoded features. The target mask prediction module can be used to determine a target mask image through upsampling and convolution processing based on the self-attention feature map and the encoded features.

[0013] In this implementation, through the target affinity module and the target mask prediction module, the learning process of the moving object segmentation model can be explicitly constrained, and the moving object segmentation model can more accurately distinguish moving objects from static objects, optimizing the semantic consistency between the background and the foreground.

[0014] In combination with the first aspect, in an alternative implementation, the above-mentioned semantic affinity constraint segmentation decoder head module further includes: a semantic affinity prediction module. Among them, the semantic affinity prediction module can be used to generate a semantic affinity map through convolution processing, batch normalization processing, and Sigmoid activation processing based on the self-attention feature map; wherein, the semantic affinity map is a heat map used to characterize the semantic correlation between the foreground and the background, and the semantic affinity map is used to train the moving object segmentation model.

[0015] In this implementation, during the training process of the moving object segmentation model, the semantic affinity prediction module can generate a semantic affinity map, providing strong supervision for the semantic relationship between pixel points to the moving object segmentation model, thereby optimizing the accuracy and robustness of the moving object segmentation model in segmenting moving objects.

[0016] In combination with the first aspect, in an alternative implementation, the expression of the loss function of the above-mentioned moving object segmentation model is: ; ; ; ; Among them, represents the loss function of the moving object segmentation model, represents the classification loss of the pixels in the frame image, represents the target mask prediction loss, represents the affinity loss, represents the set of pixels in the target mask image, represents the moving object 's predicted mask image, represents the moving object 's ground truth mask image, represents the number of moving objects, Indicates the size of the semantic affinity graph, Indicates the feature points of the semantic affinity graph, Indicates the predicted probability of the semantic affinity graph, Indicates the true value of the target probability distribution. Indicates the th moving target, Indicates the th target mask image corresponding to the moving target, Indicates the value of the target mask image at pixel point ; Indicates the value of the true value mask image at pixel point ;

[0017] Combined with the first aspect, in an alternative implementation, the above-mentioned shape prior extraction based on the true mask image to determine the edge mask image includes: First, preprocess the true mask image to obtain the preprocessed true mask image. Then, perform mask binarization on the preprocessed true mask image to obtain the binarized mask image. Finally, perform edge detection on the binarized mask image to obtain the edge mask image.

[0018] Combined with the first aspect, in an alternative implementation, the above-mentioned edge detection of the binarized mask image to obtain the edge mask image includes: First, determine the horizontal gradient and vertical gradient in the image according to the binarized mask image. Then, determine the gradient magnitude according to the horizontal gradient and vertical gradient. Next, perform binarization on the gradient magnitude to determine the binarized gradient magnitude. Finally, generate the edge mask image according to the binarized gradient magnitude.

[0019] Combined with the first aspect, in an alternative implementation, the expressions of the above-mentioned horizontal gradient and vertical gradient are: ; ; ; ; where represents the horizontal gradient, represents the vertical gradient, represents the binarized mask image, represents the Sobel operator in the horizontal direction, represents the Sobel operator in the vertical direction, represents the convolution operation, represents the abscissa of the pixel point in the binarized mask image, represents the ordinate of the pixel point in the binarized mask image. The expression of the gradient magnitude is: ; Among them, represents the gradient magnitude. The expression of the gradient magnitude after binarization is: ; Among them, represents the gradient magnitude after binarization, represents the standard deviation.

[0020] In a second aspect, the present invention provides a moving target segmentation system based on satellite video, including: an acquisition module, a shape prior extraction module, and a moving target segmentation module. Among them, the acquisition module can be used to determine a reference frame image based on the query frame image to be segmented in the satellite video, and determine the corresponding true mask image according to the reference frame image; the reference frame image is a frame image in the satellite video that has been target-segmented, and the true mask image is used to characterize the position and shape of the moving target in the reference frame image. The shape prior extraction module can be used to perform shape prior extraction according to the true mask image to determine an edge mask image, and the edge mask image is used to characterize the edge shape of the moving target. The moving target segmentation module can be used to perform feature extraction, spatio-temporal semantic relationship modeling, and semantic affinity constraint processing on the query frame image, the reference frame image, the true mask image, and the edge mask image through a moving target segmentation model, and output a target mask image, and the target mask image is used to characterize the moving target segmentation result.

[0021] In a third aspect, the present invention provides an electronic device, including: a memory, one or more processors; the memory is coupled to the processor; among them, computer program code is stored in the memory, and the computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device is caused to execute the method provided in the first aspect and any one of its alternative implementation manners as described above.

[0022] In a fourth aspect, the present invention provides a computer-readable storage medium, including computer instructions. When the computer instructions are run on an electronic device, the electronic device is caused to execute the method provided in the first aspect and any one of its alternative implementation manners as described above.

[0023] It can be understood that the beneficial effects that can be achieved by the moving target segmentation system based on satellite video provided in the second aspect, the electronic device in the third aspect, and the computer-readable storage medium in the fourth aspect can refer to the beneficial effects in the first aspect and any one of its alternative implementation manners, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a schematic diagram of the principle of the moving target segmentation method based on satellite video provided in the embodiments of the present application; Figure 2 Schematic flowchart of the moving object segmentation method based on satellite video provided by an embodiment of the present application; Figure 3 Schematic structural diagram of the moving object segmentation model provided by an embodiment of the present application; Figure 4 Schematic diagram of an example of experimental data provided by an embodiment of the present application; Figure 5 Comparison chart of the ablation experiment moving object segmentation results provided by an embodiment of the present application; Figure 6 Schematic structural diagram of the moving object segmentation system based on satellite video provided by an embodiment of the present application. Detailed implementation manners

[0025] The embodiments will be described in detail below, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following embodiments do not represent all implementation manners consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application described in detail in the claims.

[0026] Video satellites are a new type of Earth observation satellites. Compared with traditional Earth observation satellites, video satellites can continuously observe a specific area, thereby obtaining remote sensing images with high spatial resolution and high spectral resolution, and can also obtain remote sensing images with high temporal resolution. In this way, the timeliness of traditional remote sensing applications is improved, such as disaster monitoring, ocean monitoring, and ecosystem disturbance monitoring, etc. At the same time, it makes some applications that traditional remote sensing cannot perform well become a reality, such as traffic condition monitoring, etc.

[0027] The intelligent processing and analysis of satellite video can automatically extract and analyze the information in the scene of interest, which will provide important technical support for the above actual application scenarios and play an important role in them. Satellite video object detection, segmentation, and tracking are several typical tasks. Among them, satellite video object segmentation can achieve the fine segmentation of the target of interest in a time-continuous scene. Compared with detection and tracking, object segmentation can not only locate the position of the target, but also obtain additional fine description information, which provides support for subsequent more refined analysis and applications.

[0028] Different from camera imaging in natural scenes, satellite video imaging is achieved by looking down at the Earth from high altitude. Therefore, the resolution of satellite video data is usually about 1 meter, lower than that of general videos and high-resolution remote sensing data. In satellite videos, typical remote sensing targets such as vehicles, ships, and airplanes usually have small appearances, ordinary shapes and texture features, and low contrast. In addition, satellite videos have a long imaging distance and a high frame rate. The changes of targets between adjacent frames in satellite videos are small, and the background changes slowly, which results in a large spatio-temporal redundancy. Satellite video images present large-scale scenes, and the proportion of objects of interest is very small, and many details are not significant enough. At the same time, satellite video imaging has a single perspective, the background changes slowly, and the perspective and background do not change greatly, resulting in information redundancy. Especially for important rigid targets such as airplanes and ships, the perspective is single and there is no deformation. At the same time, temporal information is a unique and important attribute of satellite video data. In video scenes, people usually focus on targets that change dynamically over time, such as target tracking and real-time monitoring tasks. Therefore, in areas where rigid targets gather and dock, such as airports and ports, the appearances of many targets are extremely similar, and most of the static targets in the background are not the dynamic targets we really care about.

[0029] For satellite videos, the related methods of moving target segmentation are mainly divided into two categories: unsupervised and semi-supervised. Traditional unsupervised satellite video target segmentation methods mainly use the idea of foreground extraction. Commonly used methods for moving foreground detection and extraction include optical flow models, background subtraction models, frame difference models, and visual background extraction models. However, these traditional foreground extraction models have high requirements for video quality and assume that the background is stationary. Due to the movement of the satellite video shooting platform and the change of perspective, the background usually undergoes slow movement changes. Therefore, these foreground extraction models usually extract many background shadows and false positives. In addition, the above algorithms require a lot of manual intervention and have a single feature expression, resulting in a more complex processing process.

[0030] With the rapid development of deep learning and computer technology, due to the powerful feature representation and learning ability of convolutional neural networks (CNNs), a series of methods based on deep learning and CNNs have also been widely applied to satellite video moving target segmentation. These satellite video moving target segmentation methods mainly rely on different degrees of prior information as input, and some of these algorithms process each frame of the video independently without involving the temporal information between frames. For example, the OSVOS (One-Shot Video Object Segmentation) method uses a fully convolutional network trained offline and performs online training on images to achieve single-frame video moving target segmentation. However, since the video frames are only segmented separately for each frame without using the temporal information of the video, the moving target segmentation effect is poor.

[0031] In the related art, there is a method for segmenting moving objects in satellite videos by performing frame-by-frame optical flow estimation on satellite videos to learn and utilize the appearance and motion information of the objects. Although this method has certain advantages in dealing with the combination of appearance and motion information, it has a strong dependence on motion information and relies on optical flow estimation to extract motion features. Optical flow calculation may not be accurate enough for objects in some fast-moving or complex-background scenarios, especially when the video quality is low or the background changes dynamically, the optical flow estimation is easily affected by noise and errors. At the same time, the two-stream network and the multi-scale progressive fusion module need to process features in multiple stages, which increases the computational overhead, may lead to a relatively large computational complexity and memory occupancy of the model, and limits its applicability in real-time applications. For long video sequences, the frame-by-frame processing of the model may face high temporal dependence problems. As the length of the video increases, the transmission of frame-by-frame information and the accumulation of semantic information may lead to a decline in the model's ability to model time information, especially for complex scenarios and objects that change over a long time.

[0032] It can be seen that the related methods for segmenting moving objects based on satellite videos have problems such as complex computational processes, insufficient refinement of the edges of object segmentation, and insufficient exploration and utilization of temporal information, resulting in the inability to efficiently and accurately achieve the segmentation of moving objects.

[0033] To solve the above problems, an embodiment of the present application provides a method and system for segmenting moving objects based on satellite videos. Figure 1 The following is a schematic diagram of the principle of the method for segmenting moving objects based on satellite videos provided by an embodiment of the present application. As Figure 1 shown, first, feature extraction is performed on the query frame image, the reference frame image, the ground truth mask image, and the edge mask image. Then, based on the extracted features, a moving object segmentation model performs spatio-temporal semantic relationship modeling to achieve efficient interactive learning of semantic information within and between satellite video frames. Finally, based on the feature results enhanced by the spatio-temporal relationship, target mask prediction is performed to output the target mask image, so as to segment the moving objects in the query frame image.

[0034] This method fully combines the non-deformable characteristics of rigid moving objects and adopts a shape prior fusion method to significantly enhance the edge feature expression ability. At the same time, for stationary objects that, although not belonging to the defined pure background, interfere with the semantic learning of the moving object segmentation model, this method adopts foreground-background semantic affinity constraints to improve the robustness and accuracy of the semantic learning of the moving object segmentation model. In this way, through feature extraction, spatio-temporal semantic relationship modeling, and semantic affinity constraint processing by the moving object segmentation model, accurate segmentation of moving objects in satellite videos can be achieved, and it can effectively adapt to complex backgrounds and illumination changes, thereby effectively improving the efficiency and accuracy of segmenting moving objects in satellite videos.

[0035] The following will introduce the solution provided by the embodiments of the present application in conjunction with the accompanying drawings.

[0036] Specifically, Figure 2 is a schematic flowchart of the moving target segmentation method based on satellite video provided by the embodiments of the present application. As Figure 2 shown, the moving target segmentation method based on satellite video provided by the embodiments of the present application includes the following steps S101 - S104: S101. Determine a reference frame image based on the query frame image to be segmented in the satellite video, and determine the corresponding ground truth mask image according to the reference frame image.

[0037] In the embodiments of the present application, the query frame image is a video frame image that needs to be segmented for moving targets in the satellite video (sequence) to be processed. Based on historical information and reference data (i.e., the reference frame image and the ground truth mask image), the moving target area in the query frame image can be predicted to achieve moving target segmentation.

[0038] Exemplarily, moving targets may include, for example: wide - body aircraft, narrow - body aircraft, rear - engine aircraft, four - engine aircraft, business jets, speedboats, yachts, cruise ships, cargo ships, warships, other ships, large cars, small cars, and trains, etc.

[0039] The reference frame image is a frame image in the satellite video that has been segmented for targets. Specifically, the reference frame image and the query frame image are in the same satellite video. The reference frame image is a known frame image that has accurately completed moving target segmentation, that is, the reference frame image can provide historical moving target segmentation information and prior knowledge. In this way, the reference frame image can accurately provide the appearance information of the moving target for moving target segmentation.

[0040] The ground truth mask image corresponding to the reference frame image can be used to characterize the position and shape of the moving target in the reference frame image. Through the ground truth mask image, the network moving target segmentation model can be supervised to learn the spatial distribution of the moving target.

[0041] In one implementation, based on the query frame image, the reference frame image and the ground truth mask image corresponding to the reference frame image can be determined from the satellite video through historical segmentation results and prior knowledge.

[0042] S102. Extract shape priors according to the ground truth mask image to determine an edge mask image.

[0043] Further, in the embodiments of the present application, shape prior extraction can be performed based on the real mask image to determine the edge mask image. Specifically, the edge mask image can be used to characterize the edge shape of the moving target, that is, it can provide the prior features of the moving target. In this way, the boundary information of the moving target can be strengthened through the edge mask image, avoiding the problem of unclear boundaries caused by low contrast between the moving target and the background or motion blur, and thus the accuracy of moving target segmentation can be improved.

[0044] In some embodiments, the above S102 may specifically include the following steps S1021 - S1023: S1021. Preprocess the real mask image to obtain the preprocessed real mask image.

[0045] Specifically, the real mask image can be preprocessed such as normalized processing and noise elimination to improve the consistency and integrity of the preprocessed real mask image.

[0046] S1022. Perform mask binarization processing on the preprocessed real mask image to obtain the binarized mask image.

[0047] Specifically, the mask annotations of multiple categories of moving targets in the preprocessed real mask image can be used to extract the target regions, set the target regions to value 1, and the background to value 0 to generate the binarized mask image. Exemplarily, based on the preset threshold T, the binary mask can be calculated to determine the binarized mask image: ; Wherein, represents the binarized mask image, represents the abscissa of the pixel point in the binarized mask image, represents the ordinate of the pixel point in the binarized mask image, represents the preset gray threshold, represents the pixel point 's pixel value.

[0048] S1023. Perform edge detection on the binarized mask image to obtain the edge mask image.

[0049] Finally, based on the binarized mask image obtained in S1022, edge detection can be performed to extract the moving target boundary, generating a target contour boundary with high continuity and refinement, that is, obtaining the edge mask image.

[0050] In some embodiments, the gradient calculation method can be used for edge detection. Specifically, S1023 may specifically include the following steps S10231 - S10234: S10231. Determine the horizontal gradient and vertical gradient in the image according to the binarized mask image.

[0051] Specifically, the Sobel operator can be used to calculate the horizontal gradient and vertical gradient. Exemplarily, the expressions for the horizontal gradient and vertical gradient are: ; ; ; ; Among them, represents the horizontal gradient, represents the vertical gradient, represents the binarized mask image, represents the Sobel operator in the horizontal direction, represents the Sobel operator in the vertical direction, represents the convolution operation.

[0052] S10232. Determine the gradient magnitude according to the horizontal gradient and vertical gradient.

[0053] Exemplarily, the expression for the gradient magnitude is: ; Among them, represents the gradient magnitude. The gradient magnitude can characterize the intensity of the pixel value change in the binarized mask image. The regions with large magnitudes in the binarized mask image usually correspond to the boundaries of moving objects. To determine the direction of the moving object boundary, the gradient direction can also be calculated. Exemplarily, the expression for the gradient direction is: .

[0054] S10233. Perform binarization processing on the gradient magnitude to determine the binarized gradient magnitude.

[0055] Furthermore, the gradient magnitude can be binarized to retain the boundary points and remove noise. Exemplarily, the expression for the binarized gradient magnitude is: ; Among them, represents the binarized gradient magnitude, represents the standard deviation.

[0056] S10234. Generate an edge mask image according to the binarized gradient magnitude.

[0057] Finally, based on the determined gradient magnitude after binarization in S10233, an edge mask image can be accurately generated to provide boundary information of the enhanced moving object.

[0058] S103. Feature extraction, spatio-temporal semantic relationship modeling, and semantic affinity constraint processing are performed on the query frame image, reference frame image, true mask image, and edge mask image through the moving object segmentation model, and the target mask image is output.

[0059] In the embodiment of the present application, through the feature learning network model that fuses shape priors, that is, the moving object segmentation model, based on the query frame image, reference frame image, true mask image, and edge mask image, the target mask image can be determined. The target mask image includes the mask of the segmented moving object, and this target mask image is used to represent the moving object segmentation result, thereby realizing the accurate segmentation of the moving object.

[0060] Specifically, the moving object segmentation model can perform feature extraction on the query frame image, reference frame image, true mask image, and edge mask image. In this way, by fusing the shape prior features of the true mask image and the edge mask image, the refinement degree of the moving object segmentation edge can be improved, so as to improve the moving object segmentation accuracy and the robustness to complex background interference.

[0061] Then, spatio-temporal semantic relationship modeling can be performed based on the extracted features. In this way, through the efficient interaction of intra-frame and inter-frame semantic information, the accuracy of moving object segmentation can be improved. Not only can the temporal characteristics of satellite videos be used to dynamically model moving objects, but also the semantic features of moving objects can be accurately captured in static frame images, effectively reducing the limitations in the insufficient utilization of spatio-temporal information, and further improving the applicability of moving object segmentation in satellite videos with complex backgrounds and small-sized objects.

[0062] Furthermore, semantic affinity constraint processing is performed based on the modeled features, and the target mask image is output. In this way, through semantic affinity constraint processing, foreground-background semantic affinity constraints can be realized, explicitly modeling the semantic relationship between stationary objects and moving objects, solving the problem of stationary object interference in semantic learning, and thus improving the robustness and generalization ability of semantic learning.

[0063] By using the moving object segmentation method based on satellite video provided in the embodiments of the present application, the edge feature expression ability is significantly enhanced through the shape prior fusion method. At the same time, in view of the interference of static objects on the semantic learning of the moving object segmentation model, the method adopts foreground-background semantic affinity constraints to improve the robustness and accuracy of the semantic learning of the moving object segmentation model. In this way, through feature extraction, spatio-temporal semantic relationship modeling, and semantic affinity constraint processing by the moving object segmentation model, the accurate segmentation of moving objects in satellite video can be achieved, and it can effectively adapt to complex backgrounds and illumination changes, thereby effectively improving the efficiency and accuracy of moving object segmentation in satellite video.

[0064] In some embodiments, Figure 3 is a schematic structural diagram of the moving object segmentation model provided in the embodiments of the present application, as Figure 3 shown, the above-mentioned moving object segmentation model 300 includes: a feature extraction module 310, a spatio-temporal semantic relationship modeling module 320, and a semantic affinity constraint segmentation decoding head module 330.

[0065] Among them, the feature extraction module 310 can be used to extract features for characterizing the spatial information and edge characteristics of objects in the query frame image, reference frame image, ground truth mask image, and edge mask image respectively, to obtain the corresponding first extracted feature, second extracted feature, third extracted feature, and fourth extracted feature. And splice the first extracted feature, second extracted feature, third extracted feature, and fourth extracted feature to determine the target spliced feature.

[0066] Specifically, the feature extraction module 310 can extract features respectively through multiple convolutional layers based on the query frame image, reference frame image, ground truth mask image, and edge mask image. The extracted features characterize the spatial information and edge characteristics of objects, that is, the first extracted feature, second extracted feature, third extracted feature, and fourth extracted feature. Further, the extracted features can be spliced to fuse the representation information from different inputs.

[0067] Exemplarily, the feature extraction module 310 may adopt a three-path input layer structure. Among them, the first path input layer may adopt a regular (e.g., 3X3) convolutional layer for feature extraction of the query frame image. The second path input layer may adopt three convolutional layers, which are respectively used to extract the foreground of the reference frame image, the real mask image, and the background of the real mask image, and then add the output features of the three convolutional layers as the final output feature. The third path input layer may adopt a convolutional layer for feature extraction of the edge mask image. Then, feature splicing is performed to obtain the spliced feature. Further, the first four stages of the ResNet network can also be used as a feature extractor to perform feature extraction on the spliced feature again. Finally, the input query frame image, reference frame image, real mask image, and edge mask image are mapped to a feature map , where represent height, width, and number of channels respectively, and is the number of reference frame images.

[0068] Next, the spatial channel number of the feature map can also be reduced from to (e.g., 256) through 1×1 convolution for downsampling to achieve lightweight processing of the feature map and generate a new feature map . Then, the spatio-temporal channel number of the new feature map can also be smoothed to one dimension, i.e., from two-dimensional to one-dimensional, to adjust the shape and size of the feature map so as to adapt to the form required by the input of the Transformer and obtain the final target spliced feature . The spatio-temporal semantic relationship modeling module 320 can be used to add position encoding, perform spatio-temporal semantic relationship modeling encoding, and perform target prediction decoding based on the target spliced feature to determine the encoded feature and the decoded feature.

[0069] In the embodiments of the present application, the spatio-temporal semantic relationship modeling module 320 adopts positional encoding. By incorporating spatio-temporal position information, the encoder can effectively capture the spatio-temporal dependence relationships among the pixels in the input frame through spatio-temporal semantic relationship modeling encoding. This enables the moving object segmentation model 300 to model the dynamic changes among objects within multiple time steps, improving the accuracy of object detection and tracking. In addition, through spatio-temporal semantic relationship modeling encoding, the encoder can also learn and establish the spatio-temporal correspondence relationships of moving objects between the query frame image and the reference frame image. In this way, the moving object segmentation model 300 can learn to understand the relative changes of moving objects in different frame images and effectively model the structural features of moving objects in specific frames. The decoder can predict the spatial position information of moving objects in each frame image through target prediction decoding, thereby enhancing the spatio-temporal reasoning ability of the moving object segmentation model 300. In this way, not only can the transformation of moving objects between different time steps be accurately captured, but also by learning the robust representation of the target object, the moving object segmentation model 300 can still quickly respond and maintain high efficiency when facing the confusion problem of similar objects. This mechanism is particularly applicable to complex moving object tracking scenarios and can significantly improve the processing speed and accuracy of the moving object segmentation model 300 in dynamic environments. Therefore, by combining spatio-temporal modeling with positional encoding, the spatio-temporal semantic relationship modeling module 320 enables the moving object segmentation model 300 to effectively process the spatio-temporal information of moving objects, solves the problems of diversity and uncertainty in moving object segmentation, and thus can improve the accuracy of moving object segmentation.

[0070] Specifically, in some embodiments, as Figure 3 shown, the spatio-temporal semantic relationship modeling module 320 specifically includes: a positional encoding module 321, a Transformer encoder 322, and a Transformer decoder 323.

[0071] Among them, the positional encoding module 321 can be used to add positional encoding to the target concatenated feature to obtain the positionally encoded concatenated feature.

[0072] In this embodiment, since the self-attention mechanism does not have the ability to capture sequence position information, positional encoding (PE) can be introduced to provide position information.

[0073] In one implementation, the position encoding module 321 can add position encoding to the target concatenated feature using the sine position encoding method. The sine position encoding is based on a fixed trigonometric function pattern, making the encodings for different positions unique and maintaining a certain inductive ability under different sequence lengths of the target concatenated feature. Exemplarily, the expression for the position encoding concatenated feature can be: ; ; ; where, represents the position encoding concatenated feature, represents the target concatenated feature, represents the sine position encoding, represents the two-dimensional position index of the elements in the target concatenated feature, represents the index of the feature dimension (using sine for even indices and cosine for odd indices), represents the dimension of the position encoding (i.e., the input feature dimension of the Transformer encoder).

[0074] The Transformer encoder 322 can be used to perform spatio-temporal semantic relationship modeling and encoding on the position encoding concatenated feature through the multi-head self-attention module and the fully connected feed-forward neural network, and determine and output the encoded feature.

[0075] In this embodiment, through the Transformer encoder 322, the semantic information within and between frames can be learned to perform spatio-temporal semantic relationship modeling and encoding, forming a high-dimensional feature representation, that is, determining and outputting the encoded feature.

[0076] Specifically, the Transformer encoder 322 can be used to perform spatio-temporal relationship modeling on the position encoding concatenated feature and output the encoded feature . The encoded feature is refined by the multi-head self-attention module and the fully connected feed-forward neural network to capture the pixel-level semantic relationships within the frame and the temporal correlations between frames for the motion target segmentation model 300.

[0077] Among them, the encoded feature has stronger global representation ability. Specifically: In the spatial dimension, the encoded feature learns the long-range dependence relationships between pixels in the position encoding concatenated feature through the multi-head self-attention mechanism, strengthens the structural integrity of the foreground target, and reduces the influence of local noise. In the temporal dimension, the encoded feature Integrates the motion information between frames, realizes the modeling of the position and shape changes of moving objects, and improves the spatio-temporal consistency description of dynamic moving objects. In terms of feature representation, the encoded features After multiple non-linear transformations, it can aggregate relevant information more closely, making the semantic representation clearer and more robust.

[0078] Exemplarily, the Transformer encoder 322 can be composed of 8 encoding layers, and each encoding layer includes standard structural modules: the multi-head self-attention module and the fully connected feed-forward network. By adopting 8 different multi-head self-attention modules, the encoder can parallelly capture the spatio-temporal dependencies between pixels in the input sequence within multiple subspaces, thereby enhancing the understanding of the target motion pattern.

[0079] The Transformer decoder 323 can be used to perform target prediction decoding based on the encoded features through the multi-head self-attention module, the multi-head cross-attention module, and the fully connected feed-forward neural network, and determine and output the decoded features.

[0080] In this embodiment, the spatial position information of the moving object in each frame of image (such as the query frame image) can be predicted through the Transformer decoder 323. The Transformer decoder 323 receives the encoded features output by the Transformer encoder 322 Determine and output the decoded features through the multi-head self-attention module, the multi-head cross-attention module, and the fully connected feed-forward neural network , and the decoded features Can be used for the generation of the target mask image.

[0081] In the training stage of the moving object segmentation model 300, the Transformer decoder 323 can also determine and output the decoded features according to the encoded features And the training target query Through the multi-head self-attention module, the multi-head cross-attention module, and the fully connected feed-forward neural network . Among them, the training target query Represents the feature query vector of the moving object to be predicted, and can be used to extract the corresponding moving object features.

[0082] Specifically, the decoded features output by the Transformer decoder 323 Is a high-dimensional representation of the spatial position information and semantic features of the moving object in each frame of image. This decoded feature Fuses the global spatio-temporal information output by the encoder and the training target query during the training stage The extracted target-specific information can accurately depict the spatial contour of the moving target and its changing trend over time after multiple layers of decoding.

[0083] Among them, the decoding features have the following characteristics: Target area prediction: The decoding features can be further upsampled and passed through a convolutional layer to generate the final target mask image, achieving pixel-level moving target area segmentation. Target shape optimization: The decoding features combined with context information can be used to optimize the edge details of the moving target, improving the accuracy of the moving target segmentation result. Temporal consistency enhancement: The decoding features inherit the spatio-temporal correlation modeling ability of the Transformer encoder, enabling the moving target segmentation result to maintain coherence between sequential frames and reducing the impact of temporal jitter and drift.

[0084] Exemplarily, the Transformer decoder 323 can be composed of 8 decoder layers, each layer including a multi-head self-attention module, a multi-head cross-attention module, and a fully connected feed-forward network. Through the multi-head cross-attention module, the target features obtained in the Transformer encoder stage can be effectively and deeply fused with the query information of the Transformer decoder, further optimizing the spatio-temporal feature modeling of the moving target. In this way, the Transformer decoder 323 can better integrate and utilize the feature information extracted by the Transformer encoder when dealing with complex scenarios, thereby improving the positioning accuracy of the moving target.

[0085] It can be seen that through the interaction and optimization of multiple decoding layers, the Transformer decoder 323 can ensure the efficient prediction and stable output of the spatial position information of the moving target. Through the spatio-temporal relationship modeling ability of the Transformer and the complex multi-head cross-attention mechanism in the Transformer decoder, it can effectively handle problems in the moving target segmentation task, including multi-target confusion, target occlusion, and changes in the motion trajectory, etc., thus significantly improving the accuracy of the moving target segmentation model 300 in segmenting moving targets in a dynamic environment.

[0086] The semantic affinity constraint segmentation head module 330 can be used to perform mask prediction based on the encoded features and the decoding features to determine the target mask image.

[0087] In this embodiment, the semantic affinity constraint segmentation decoder module 330 can achieve robust representation learning of semantic consistency. During the moving object segmentation process, the ground truth annotation can provide information about the class to which each pixel belongs, but it is difficult for the moving object segmentation model 300 to learn complete context information from a single pixel-level feature. Especially in the video scenario of earth observation satellites, the target size is small, and the ratio of foreground and background pixels is seriously imbalanced, which further increases the learning difficulty of the moving object segmentation model 300.

[0088] In this context, stationary background targets (such as ground objects or environmental features) usually have similar appearances and classes to moving objects. Since stationary targets do not move and usually belong to the same class or model as the target of interest, they are easily misjudged as the background. In the case of misclassifying these stationary targets as the background, it will lead to significant biases in the semantic information learning of moving objects, seriously affecting the accurate segmentation and recognition of moving objects. The embodiment of this application adopts a semantic affinity constraint mechanism, that is, the semantic affinity constraint segmentation decoder module 330 explicitly constrains the learning process of the moving object segmentation model 300, so that the moving object segmentation model 300 can more accurately distinguish moving objects from stationary objects and optimize the semantic consistency between the background and the foreground.

[0089] The semantic affinity constraint segmentation decoder module 330 can explicitly regulate the training of the moving object segmentation model 300, strengthen the semantic differences between the target and the background, reduce the risk of misjudging stationary targets as the background, and thus improve the segmentation performance of the moving object segmentation model 300 for moving objects in complex scenarios. The semantic affinity constraint segmentation decoder module 330 can effectively solve the semantic confusion problem in moving object segmentation. Especially in the case where the appearances of stationary objects and moving objects are similar, it can effectively improve the robustness and accuracy of the semantic learning of the moving object segmentation model 300 for moving objects.

[0090] In the embodiment of this application, the semantic affinity constraint segmentation decoder module 330 is composed of a dual-branch parallel processing architecture, including a target mask prediction branch (i.e., the target mask prediction module) and a semantic affinity prediction branch (i.e., the semantic affinity prediction module), and realizes the segmentation of moving objects in satellite videos through multi-scale feature fusion and cross-modal attention mechanisms.

[0091] Specifically, in some embodiments, as Figure 3 shown, the semantic affinity constraint segmentation decoder module 330 includes: a target affinity module 331 and a target mask prediction module 332.

[0092] Among them, the target affinity module 331 can be used to determine and output a self-attention feature map according to the decoded feature and the encoded feature.

[0093] Specifically, the target affinity module 331 can use the encoded features as the query Q, the decoded features as the key K and the value V, and determine and output the self-attention feature map based on the Scaled Dot-Product Attention (SDPA) mechanism.

[0094] The target mask prediction module 332 can be used to determine the target mask image according to the self-attention feature map and the encoded features through upsampling and convolutional processing.

[0095] Exemplarily, the target mask prediction module 332 can include: multiple (e.g., 3) feature collaboration blocks, 2 convolutional layers (1×1, or 3×1), and Softmax activation to perform upsampling and convolutional processing on the self-attention feature map and the encoded features to determine the target mask image.

[0096] Among them, the feature collaboration block can adopt a 3-level cascaded processing unit structure, which can realize fusing multi-scale features through skip connections, reducing the computational complexity with depthwise separable convolutions, and gradually generating 256-channel intermediate features.

[0097] In some embodiments, as Figure 3 shown, the semantic affinity constraint segmentation head module 330 further includes: a semantic affinity prediction module 333.

[0098] The semantic affinity prediction module 333 can be used to generate a semantic affinity map according to the self-attention feature map through convolutional processing, batch normalization processing, and Sigmoid activation processing. Among them, the semantic affinity map is a heat map used to characterize the semantic correlation between the foreground and the background, and the semantic affinity map is used to train the moving object segmentation model.

[0099] In the embodiments of the present application, during the training process of the moving object segmentation model 300, the semantic affinity prediction module 333 can generate a semantic affinity map to provide strong supervision on the semantic relationship between pixel points for the moving object segmentation model 300, thereby optimizing the accuracy and robustness of the moving object segmentation model 300 for segmenting moving objects.

[0100] Exemplarily, the semantic affinity prediction module 333 can include: a feature adaptation module and a prediction layer. Among them, the feature adaptation module can decompose the self-attention feature map using fully separable convolutions with kernels of and through a cascaded operation, and generate adapted features through channel compression. The prediction layer can include: a 1×1 convolutional layer, a batch normalization layer, and Sigmoid activation. The prediction layer can be used to generate a semantic affinity map according to the adapted features Generate and output a semantic affinity graph.

[0101] In some embodiments, to construct an ideal semantic affinity graph, the sample frame images of the input satellite video first have corresponding ground truth labels . The sample frame images are processed by a feature extraction module, and high-dimensional feature representations are extracted through a spatio-temporal semantic relationship modeling module. Subsequently, these high-dimensional feature representations can be input into a multi-layer perceptron (MLP) and a segmentation head, and finally a segmentation feature map with a size of is output.

[0102] To be consistent with the segmentation feature map, the ground truth labels can be downsampled to the same resolution as the segmentation feature map, denoted as . Then, is one-hot encoded, and is converted into a tensor with a dimension of , where is the total number of segmentation categories. The tensor with this dimension of is reshaped and reorganized into a matrix form of , where represents the total number of pixels in the segmentation feature map.

[0103] During the construction of the semantic affinity graph, by calculating ( is the label tensor after shape adjustment), a symmetric matrix A of can be obtained. This matrix A can intuitively represent the category affinity relationship between pixels, that is, the ideal semantic affinity graph. Through this semantic affinity graph, strong supervision on the semantic relationship between pixel points can be provided for the moving object segmentation model, thereby optimizing the accuracy and robustness of moving object segmentation.

[0104] In some embodiments, the expression of the loss function for training the moving object segmentation model 300 is: ; ; ; ; where, represents the loss function of the moving object segmentation model, represents the classification loss of the pixels in the frame image, represents the target mask prediction loss, represents the affinity loss, represents the set of pixels in the target mask image, represents the moving object The predicted mask image, represents the moving target The ground truth mask image, represents the number of moving targets, represents the size of the semantic affinity map, represents the feature points of the semantic affinity map, represents the predicted probability of the semantic affinity map, represents the ground truth of the target probability distribution, represents the th moving target, represents the th target mask image corresponding to the moving target, represents the value of the target mask image at the pixel point , represents the value of the ground truth mask image at the pixel point .

[0105] In some embodiments, to verify the performance and effectiveness of the moving target segmentation method based on satellite video provided in the above embodiments of the present application (hereinafter referred to as: the technical solution of the present invention), a verification analysis is carried out.

[0106] Experimental dataset: The dataset in the satellite video multi-task benchmark dataset SAT-MTB (Satellite Video Multiple Tasks Benchmark) is selected as the experimental dataset. The SAT-MTB dataset supports multiple tasks such as object detection, object tracking, and object segmentation. The experimental dataset contains moving targets. Specifically, the experimental dataset contains 30 videos, with a frame rate of 10 frames per second and an average duration of about 22 seconds, for a total of 5226 frames of images.

[0107] Exemplarily, Figure 4 is a schematic diagram of an experimental data sample provided in an embodiment of the present application. As shown in (a) in Figure 4 , it is a schematic diagram of an experimental data sample with an airplane as the moving target. As shown in (b) in Figure 4 , it is a schematic diagram of an experimental data sample with a ship as the moving target.

[0108] Evaluation method: In this technical solution, J-Mean, F-Mean, and J&F Mean evaluation metrics are used to evaluate the similarity between the moving target segmentation result and the true label.

[0109] Specifically, J-Mean refers to the average of the intersection over union of the segmentation result and the true label, with a value range of 0-1. J-Mean calculates the IoU for each sample and takes the average of the IoUs of all samples, which can be used to measure the regional overlap between the moving target segmentation result and the true label.

[0110] F-Mean is the average value of F1-score based on the accuracy of the segmentation boundary, which can be used to evaluate the matching degree between the moving object segmentation result and the real label boundary, and its value range is 0-1.

[0111] The specific calculation method is as follows: Among them, represents the proportion captured by the moving object segmentation model in the pixels corresponding to the real moving object, represents the pixels predicted as moving objects by the moving object segmentation model, which are consistent with the real label, represents the moving object pixels missed by the moving object segmentation model, that is, missed detection, represents the proportion of the pixels predicted as moving objects by the moving object segmentation model that are truly moving objects, represents the background pixels mispredicted as moving objects by the moving object segmentation model, that is, false alarm, represents the result of overall balanced consideration of Recall and Precision.

[0112] J&F Mean is the arithmetic mean of J-Mean and F-Mean, which can be used to comprehensively evaluate the performance of the segmentation region and the segmentation boundary. Considering the performance of region overlap (J-Mean) and boundary matching (F-Mean) comprehensively, it is a more comprehensive segmentation quality evaluation index. The specific calculation method is as follows: ; Among them, represents J&F Mean, represents J-Mean, represents F-Mean.

[0113] The technical solution of the present invention is compared with the related technology one RVOS method and the related technology two OSVOS method in the satellite video moving object segmentation accuracy experiment. The experimental results are shown in Table 1.

[0114] Table 1 Experimental results of satellite video moving object segmentation accuracy comparison Finally, the J&F Mean values of different methods are shown in Table 2.

[0115] Table 2 Comparison results of J&F Mean values The of the technical solution of the present invention on the SAT-MTB dataset is 0.409, is 0.554, The value is 0.471, which is superior to the existing two technical solutions in terms of precision indicators. It can be seen that by fusing the significant features of moving ships, higher-quality pseudo-label samples can be obtained. Based on the target detection network resistant to noise interference, during the training and learning process of the moving target segmentation model, the interference brought by the boundary noise in the pseudo-label samples of ships can be resisted, and more accurate and detailed feature descriptions of moving ships can be realized. Finally, the target localization accuracy is further improved through candidate region consistency regression.

[0116] Moreover, the technical solution of the present invention is different from the related technology one RVOS method in that the technical solution of the present invention adopts an end-to-end manner to construct the significant fusion features of satellite video sequence frames, extract the moving foreground, etc., which are only used to obtain pseudo-label samples in the training stage. After the moving target segmentation model completes network training, the moving target segmentation model can directly perform rapid inference on the original satellite video to be tested, with higher inference efficiency. Therefore, in the statistical comparison experiment of the moving ship detection efficiency, the model inference duration of the technical solution of the present invention, the related technology one RVOS method, and the related technology two OSVOS method on the SAT-MTB test set is compared to quantitatively analyze the detection efficiency of moving ships in satellite videos. The experimental results show that while achieving better moving ship detection effects, the technical solution of the present invention has shorter model inference time, and the detection speed is not only far ahead of the related technology one RVOS method but also superior to the related technology two OSVOS method.

[0117] Furthermore, in some embodiments, ablation experiment analysis is carried out to verify the functions of each step in the technical solution of the present invention. By gradually adding and introducing spatio-temporal semantic affinity constraints (i.e., the above S103), shape priors (i.e., the above S102), and the combined method of the two (i.e., the technical solution of the present invention), the influence of each step on the accuracy of moving target segmentation in satellite videos is analyzed. Figure 5 This is a comparison diagram of the moving target segmentation results of the ablation experiment provided by the embodiments of the present application. As Figure 5 shown in (1) therein, it is a schematic diagram of the segmentation result obtained by using the baseline method to segment the moving targets in the satellite video. As Figure 5 shown in (2) therein, it is a schematic diagram of the segmentation result obtained by using the baseline method and spatio-temporal semantic affinity constraints to segment the moving targets in the satellite video. As Figure 5 shown in (3) therein, it is a schematic diagram of the segmentation result obtained by using the baseline method and shape priors to segment the moving targets in the satellite video. As Figure 5 shown in (4) therein, it is a schematic diagram of the segmentation result obtained by using the baseline method, spatio-temporal semantic affinity constraints, and shape priors (i.e., the technical solution of the present invention) to segment the moving targets in the satellite video. As Figure 5 shown in (5) therein, it is a schematic diagram of the ground truth label of the segmentation result of the moving targets in the satellite video. Table 3 shows the results of the ablation experiment on moving target segmentation in satellite videos.

[0118] Table 3 Ablation Experiment Results of Moving Object Segmentation in Satellite Videos As shown in Table 3, Table 3 presents the ablation experiment results of moving object segmentation on eight typical satellite video sequences (including ship and aircraft categories). The evaluation metrics are J-Mean and F-Mean, which are used to measure the regional overlap and boundary contour similarity between the segmented region and the ground truth annotation. Starting from the baseline method, the ablation experiment respectively introduces spatio-temporal semantic affinity constraints, shape priors, and a combined strategy of both, and comparatively analyzes their performance differences in different object types and background scenes. It can be seen from the results that after introducing the affinity constraints, the J-Mean value of the moving object segmentation model has improved on most sequences, especially significantly in videos such as Air1 and Air2 with small foreground objects and complex backgrounds, indicating that the affinity mechanism can effectively enhance cross-frame semantic consistency and foreground recognition ability. After introducing the shape priors, good results are also achieved on objects with stable structures and clear boundaries (such as aircraft sequences), especially in sequences such as Air1, Air4, and Air5. The moving object segmentation model can not only accurately segment the foreground but also maintain smooth and consistent edges. When the spatio-temporal semantic affinity constraints and shape priors are jointly introduced, that is, when the technical solution of the present invention is adopted, the J-Mean and F-Mean of the moving object segmentation model reach the optimal on most videos, verifying the complementary enhancement effect of the technical solution of the present invention in moving object contour retention and semantic information aggregation. Generally speaking, the technical solution of the present invention can more comprehensively improve the perception and segmentation quality of the moving object segmentation model for moving foreground objects, taking into account semantic consistency and structural constraints, and showing good generalization ability and robustness for moving object segmentation.

[0119] The embodiment of the present application also provides a moving object segmentation system based on satellite videos. Specifically, Figure 6 is the structural schematic diagram of the moving object segmentation system based on satellite videos provided by the embodiment of the present application. As Figure 6 shown, the moving object segmentation system 600 based on satellite videos includes: an acquisition module 601, a shape prior extraction module 602, and a moving object segmentation module 603.

[0120] Among them, the acquisition module 601 can be used to determine a reference frame image based on the query frame image to be segmented in the satellite video, and determine the corresponding ground truth mask image according to the reference frame image; the reference frame image is a frame image in the satellite video that has been object-segmented, and the ground truth mask image is used to represent the position and shape of the moving object in the reference frame image.

[0121] The shape prior extraction module 602 can be used to extract the shape prior according to the real mask image and determine the edge mask image, which is used to characterize the edge shape of the moving target.

[0122] The moving target segmentation module 603 can be used to perform feature extraction, spatio-temporal semantic relationship modeling, and semantic affinity constraint processing on the query frame image, reference frame image, real mask image, and edge mask image through the moving target segmentation model, and output the target mask image, which is used to characterize the moving target segmentation result.

[0123] By using the moving target segmentation system based on satellite video provided in the embodiments of the present application, through the shape prior fusion method, the edge feature expression ability is significantly enhanced. At the same time, aiming at the interference of static targets on the semantic learning of the moving target segmentation model, the system adopts foreground-background semantic affinity constraints to improve the robustness and accuracy of the semantic learning of the moving target segmentation model. In this way, through feature extraction, spatio-temporal semantic relationship modeling, and semantic affinity constraint processing by the moving target segmentation model, the accurate segmentation of moving targets in satellite video can be achieved, and it can effectively adapt to complex backgrounds and illumination changes, thereby effectively improving the efficiency and accuracy of moving target segmentation in satellite video.

[0124] The embodiments of the present invention further provide an electronic device, which may include: a display screen, a memory, and one or more processors. The display screen, memory, and processor are coupled. The memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device can execute each method or step executed in the embodiments of the above-mentioned moving target segmentation method based on satellite video. Of course, the electronic device includes but is not limited to the above-mentioned display screen, memory, and one or more processors.

[0125] The embodiments of the present invention further provide a computer-readable storage medium for storing computer instructions for running the above-mentioned moving target segmentation method based on satellite video.

[0126] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above-mentioned division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0127] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0128] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0129] For the similar parts between the embodiments provided in this application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of this application and do not constitute a limitation on the protection scope of this application. For those skilled in the art, any other embodiments extended based on the solution of this application without creative efforts belong to the protection scope of this application.

Claims

1. A moving target segmentation method based on satellite video, characterized in that: include: Determine a reference frame image based on a query frame image to be segmented in a satellite video, and determine a corresponding true mask image according to the reference frame image; The reference frame image is a frame image in the satellite video that has been segmented by a target, and the real mask image is used to characterize the position and shape of the moving target in the reference frame image; Performing shape priori extraction according to the real mask image to determine an edge mask image, wherein the edge mask image is used to characterize the edge shape of the moving target; The query frame image, the reference frame image, the real mask image and the edge mask image are subjected to feature extraction, spatiotemporal semantic relationship modeling and semantic affinity constraint processing through a moving target segmentation model, and a target mask image is outputted. The target mask image is used to characterize the moving target segmentation result.

2. The method according to claim 1, characterized in that The moving target segmentation model includes: a feature extraction module, a spatiotemporal semantic relationship modeling module and a semantic affinity constraint segmentation decoding head module; wherein, The feature extraction module is used to extract features for characterizing the spatial information and edge characteristics of the object from the query frame image, the reference frame image, the real mask image and the edge mask image, respectively, to obtain corresponding first extracted features, second extracted features, third extracted features and fourth extracted features, and to splice the first extracted features, the second extracted features, the third extracted features and the fourth extracted features to determine a target splicing feature; The spatiotemporal semantic relationship modeling module is used to perform position coding addition, spatiotemporal semantic relationship modeling coding and target prediction decoding according to the target splicing features, and determine coding features and decoding features; The semantic affinity constraint segmentation decoding head module is used to perform mask prediction according to the encoding features and the decoding features to determine the target mask image.

3. The method according to claim 2, characterized in that The spatiotemporal semantic relationship modeling module includes: a position encoding module, a Transformer encoder and a Transformer decoder; wherein, The position coding module is used to add position coding to the target splicing feature to obtain a position coding splicing feature; The Transformer encoder is used to perform spatiotemporal semantic relationship modeling encoding through a multi-head self-attention module and a fully connected feedforward neural network according to the position coding splicing features, and determine and output the coding features; The Transformer decoder is used to perform target prediction decoding according to the encoding features through a multi-head self-attention module, a multi-head cross-attention module and a fully connected feedforward neural network to determine and output the decoding features.

4. The method according to claim 3, characterized in that The semantic affinity constraint segmentation decoding head module includes: a target affinity module and a target mask prediction module; wherein, The target affinity module is used to determine and output a self-attention feature map according to the decoding feature and the encoding feature; The target mask prediction module is used to determine the target mask image through upsampling and convolution processing according to the self-attention feature map and the encoding feature.

5. The method according to claim 4, characterized in that The semantic affinity constraint segmentation decoding head module further includes: a semantic affinity prediction module; The semantic affinity prediction module is used to generate a semantic affinity map according to the self-attention feature map through convolution processing, batch normalization processing and Sigmoid activation processing; wherein the semantic affinity map is a heat map used to characterize the semantic correlation between foreground and background, and the semantic affinity map is used to train the moving target segmentation model.

6. The method according to claim 5, characterized in that The loss function of the moving target segmentation model is expressed as: ; ; ; ; in, represents the loss function of the moving target segmentation model, represents the classification loss of pixels in the frame image, represents the target mask prediction loss, represents affinity loss, represents the set of pixels in the target mask image, Indicates sports goal The predicted mask image, Indicates sports goal The true value mask image, Indicates the number of moving targets, represents the size of the semantic affinity graph, Represents the feature points of the semantic affinity graph, represents the predicted probability of the semantic affinity graph, represents the true value of the target probability distribution, Indicates sports goals, Indicates The target mask image corresponding to the moving target, Indicates the target mask image at pixel point The value at Represents the true value mask image at pixel point The value at .

7. The method according to claim 1, characterized in that The step of performing shape priori extraction according to the real mask image to determine the edge mask image includes: Preprocessing the real mask image to obtain a preprocessed real mask image; Performing mask binarization processing on the preprocessed real mask image to obtain a binary mask image; Perform edge detection on the binary mask image to obtain the edge mask image.

8. The method according to claim 7, characterized in that The performing edge detection on the binary mask image to obtain the edge mask image includes: Determining a horizontal gradient and a vertical gradient in an image according to the binary mask image; Determining a gradient amplitude according to the horizontal gradient and the vertical gradient; Binarizing the gradient amplitude to determine the binarized gradient amplitude; The edge mask image is generated according to the binarized gradient amplitude.

9. The method according to claim 8, characterized in that The expressions of the horizontal gradient and the vertical gradient are: ; ; ; ; in, represents the horizontal gradient, represents the vertical gradient, represents the binary mask image, represents the Sobel operator in the horizontal direction, represents the Sobel operator in the vertical direction, represents the convolution operation, Represents the horizontal coordinate of the pixel in the binary mask image, Represents the ordinate of the pixel in the binary mask image; The expression of the gradient amplitude is: ; in, represents the gradient amplitude; The expression of the gradient amplitude after binarization is: ; in, represents the gradient amplitude after binarization, Represents standard deviation.

10. A moving target segmentation system based on satellite video, characterized in that: include: Acquisition module, shape prior extraction module, moving target segmentation module; among them, The acquisition module is used to determine a reference frame image based on a query frame image to be segmented in a satellite video, and determine a corresponding real mask image according to the reference frame image; the reference frame image is a frame image in the satellite video that has been segmented by a target, and the real mask image is used to characterize the position and shape of a moving target in the reference frame image; The shape prior extraction module is used to perform shape prior extraction based on the real mask image to determine an edge mask image, wherein the edge mask image is used to characterize the edge shape of the moving target; The moving target segmentation module is used to perform feature extraction, spatiotemporal semantic relationship modeling and semantic affinity constraint processing on the query frame image, the reference frame image, the real mask image and the edge mask image through a moving target segmentation model, and output a target mask image, which is used to represent the moving target segmentation result.

Citation Information

Patent Citations

  • Timing sequence remote sensing image segmentation method and device, electronic equipment and storage medium

    CN114387275A

  • Adaptive video target segmentation method for processing multiple prior knowledge

    CN114494297A

  • Target segmentation method and system based on spiking neural network

    CN117253039A

  • Video object segmentation by reference-guided mask propagation

    US20190311202A1

  • Graph-based video instance segmentation

    US20230118401A1

Cited By

  • Method and device for extracting region of interest of rice and electronic equipment

    CN121415056A

  • Method and device for extracting rice region of interest, and electronic equipment

    CN121415056B

  • Load action and target behavior identification method based on space time sequence image

    CN121686116A

  • Moving target detection method, electronic equipment and storage medium

    CN122223047A

  • Rat swing action segmentation method, system, equipment and medium

    CN122244958A