A method for simultaneous detection and segmentation of space targets

By using the ResNet-FPN network in the instance segmentation technology to extract and fuse multi-layer feature maps, and using multi-classified Focal Loss loss function, the object detection and segmentation network is built, which solves the problems of easy omission of small objects, image degradation and uneven category samples, and achieves efficient detection and segmentation of spatial targets.

CN114898092BActive Publication Date: 2025-05-06NANJING UNIV OF SCI & TECH

Patent Information

Application Number
CN202210396418.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-15
Publication Date
2025-05-06
Estimated Expiration
2042-04-15

AI Technical Summary

Technical Problem

Small objects in the existing instance segmentation technology are prone to omission, image degradation, and uneven category samples.

Method used

By constructing the training sample set, multi-layer feature maps are extracted and fused using the ResNet-FPN network, the object detection and segmentation network is built, the region suggestion model and detection segmentation model are used for training, and the multi-classified Focal Loss loss function is used for category marking prediction.

Benefits of technology

It improves the accuracy of detection of small objects, enhances the ability to express image features, solves the problem of uneven category samples, and achieves stable detection and segmentation of spatial targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114898092B_ABST
    Figure CN114898092B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for simultaneous detection and segmentation of space targets. ResNet-FPN is used to extract multiple layers of feature maps of different scales, and the multiple layers of feature maps of different scales are further fused. The features of all layers are fused on the feature maps of different scales, and the edge shape and other information of the shallow network and the semantic information of the deep network are retained as much as possible. Finally, the obtained feature expression ability is stronger, and the effect of dealing with problems such as omission of small objects, geometric transformation, and image degradation is more robust. Multi-classification loss FocalLoss is designed as the loss function in component classification detection to avoid the problem of uneven category samples when mining difficult samples. Without losing the inference rate, the detection and segmentation effects of space targets can be kept stable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method for synchronous detection and segmentation of space targets. Background Art

[0002] In the field of computer vision, object detection and semantic segmentation are both basic tasks. Object detection aims to identify object instances in an image and detect their locations, while semantic segmentation aims to classify each pixel in the image. The instance segmentation task is a combination of object detection and semantic segmentation. Its purpose is to first detect the object in the image and then assign a category label to each pixel of the object. Instance segmentation can distinguish different instances with the same foreground semantic category, which is the biggest difference between it and semantic segmentation. Because it plays an important role in application technology support in the fields of geographic information systems, medical imaging, autonomous driving, and robots, instance segmentation has become an important part of image segmentation and has very important research significance.

[0003] Instance segmentation can draw on the advantages of both target detection and semantic segmentation, and simultaneously complete the detection and segmentation tasks. At the same time, this task is also extremely challenging, so it has attracted widespread attention from scholars at home and abroad. In recent years, the representative ones are: EXie et al. proposed the PolarMask method in 2019, which does not require the generation of a detection frame, and compared with FCOS, 4 rays are scattered to 36 rays, and instance segmentation and target detection are expressed in the same modeling method; H Chen et al. proposed the BlendMask method in 2020, which combines the top-down and bottom-up ideas, adds a bottom-level module on the basis of FCOS to extract low-level detail features, and generates corresponding attention at the top level, and finally proposes a blender module to better fuse the two features; Although the above methods can obtain detection and segmentation results, for real data sets containing complex scenes, such as the spatial targets in the present invention, there are problems such as small objects are easily missed, image degradation, and uneven category samples. In response to the above problems, the present invention further fuses the multi-scale feature maps output by the feature pyramid network and adopts a multi-classification Focal Loss loss function to provide a method for synchronous detection and segmentation of spatial targets. Summary of the invention

[0004] The purpose of the present invention is to provide a method for synchronous detection and segmentation of space targets, so as to solve the problems of easy omission of small objects, image degradation and uneven category samples existing in the existing instance segmentation technology.

[0005] To achieve the above object, the present invention provides a method for synchronous detection and segmentation of space targets, comprising the following steps:

[0006] Constructing a training sample set including a plurality of training samples; the training samples include a spatial image and an image tag corresponding to the spatial image; the image tag includes a detection tag, a segmentation tag and a category tag in the spatial image, and the category tag includes category tags of multiple target components;

[0007] Generate an initial set of regions of interest for each training sample;

[0008] The ResNet-FPN network is used to extract the multi-layer feature map of each training sample, and the features of all other feature maps are fused on each layer of feature map to obtain the multi-layer fused feature map of each training sample;

[0009] Extracting a feature matrix of each initial region of interest of the training sample from the multi-layer fusion feature map of the training sample;

[0010] Constructing a target detection and segmentation network; the target detection and segmentation network includes a region proposal model and a detection and segmentation model; the region proposal model is used to determine whether the corresponding initial region of interest belongs to the background according to the feature matrix of each initial region of interest, and to screen and optimize the initial regions of interest in the initial region of interest set; the detection and segmentation model is used to perform category label prediction, detection label prediction and segmentation label prediction for various target components in the spatial image; when predicting the category label of the region of interest in the region proposal model, a binary classification loss is used as its loss function, and when predicting the category label in the detection and segmentation model, a multi-classification loss Focal Loss is used as its loss function;

[0011] Using the multi-layer fusion feature map of each training sample, the initial region of interest set, the feature matrix of each initial region of interest and the training sample set, the region proposal model and the detection and segmentation model are trained respectively to obtain a trained target detection and segmentation network;

[0012] Use the trained target detection and segmentation network to detect and segment the spatial targets in the image to be detected.

[0013] Optionally, the ResNet-FPN network is used to extract a multi-layer feature map of each training sample, and the features of all other feature maps are fused on each layer of the feature map to obtain a multi-layer fused feature map of each training sample, specifically including:

[0014] The spatial image of each training sample is subjected to feature extraction through the ResNet-FPN network to obtain a multi-layer feature map;

[0015] Each layer of feature map is fused with the features of all other feature maps to obtain a multi-layer fused feature map; wherein each layer of feature map is fused with the down-sampling of all its shallow feature maps and the up-sampling of all its deep feature maps.

[0016] Optionally, the using of the multi-layer fusion feature map of each training sample, the initial region of interest set, the feature matrix of each initial region of interest and the training sample set to respectively train the region proposal model and the detection segmentation model specifically includes:

[0017] Iteratively training a first classification detection branch in the region proposal model according to the initial region of interest set, the feature matrix of each initial region of interest and the training sample set to obtain a trained region proposal model; in the process of iteratively training the first classification detection branch, taking each feature matrix of each initial region of interest as an input of the first classification detection branch, and taking the category label and detection label of the corresponding region of interest in the spatial image as the target output of the first classification detection branch;

[0018] Optimizing the initial set of regions of interest of each training sample using the trained region proposal model to obtain an optimized set of regions of interest of each training sample;

[0019] According to the set of optimized regions of interest of each training sample, an adaptive ROIAlign algorithm is used to extract multiple optimized feature matrices corresponding to each optimized region of interest from the multi-layer fusion feature map of the corresponding training sample; the sizes of each optimized feature matrix are the same;

[0020] Add and fuse multiple optimization feature matrices corresponding to each optimized region of interest to obtain a fused optimization feature matrix corresponding to each optimized region of interest;

[0021] According to the fused optimized feature matrix set of each training sample and the image label of each training sample, the second classification detection branch and the segmentation branch in the detection and segmentation model are iteratively trained in turn to obtain a trained detection and segmentation model; in the process of iterative training of the second classification detection branch, the fused optimized feature matrix set of the training samples is used as the input of the second classification detection branch, and the category label and detection label in the training sample are used as the target output of the second classification detection branch; in the process of iterative training of the segmentation branch, the fused optimized feature matrix set of the training samples is used as the input of the segmentation branch, and the segmentation label in the training sample is used as the output of the segmentation branch; the fused optimized feature matrix set of each training sample includes the fused optimized feature matrices corresponding to all optimized regions of interest of each training sample.

[0022] Optionally, the iterative training of the first classification detection branch in the region proposal model according to the initial region of interest set, the feature matrix of each initial region of interest and the training sample set specifically includes:

[0023] Inputting the feature matrix of each initial region of interest into the first classification detection branch in sequence to obtain a first category prediction result and a first detection prediction result;

[0024] Calculating a first category prediction loss of the first classification detection branch according to the first category prediction result and the category label of the corresponding region of interest;

[0025] Calculating a first detection prediction loss of the first classification detection branch according to the first detection prediction result and the detection mark of the corresponding region of interest;

[0026] According to the first category prediction loss and the first detection prediction loss, the parameters of the first classification detection branch are updated through a back propagation algorithm until an iteration termination condition is met.

[0027] Optionally, the iterative training of the second classification detection branch and the segmentation branch in the detection and segmentation model in sequence according to the fusion optimization feature matrix set of each training sample and the image label of the corresponding training sample specifically includes:

[0028] Inputting the fusion optimized feature matrix set of each training sample into the second classification detection branch in sequence to obtain a second category prediction result and a second detection prediction result;

[0029] Inputting the fusion optimized feature matrix set of each training sample into the segmentation branch in sequence to obtain a segmentation prediction result;

[0030] Calculating a second category prediction loss of the second classification detection branch according to the second category prediction result and the category label in the corresponding training sample;

[0031] Calculating a second detection prediction loss of the second classification detection branch according to the second detection prediction result and the detection mark in the corresponding training sample;

[0032] Calculating the segmentation prediction loss of the segmentation branch according to the segmentation prediction result and the segmentation mark in the corresponding training sample;

[0033] According to the second category prediction loss and the second detection prediction loss, the parameters of the second classification detection branch are updated through a back propagation algorithm; according to the segmentation prediction loss, the parameters of the segmentation branch are updated through a back propagation algorithm until an iteration termination condition is met.

[0034] Optionally, the optimizing the initial region of interest set of each training sample by using the trained region proposal model specifically includes:

[0035] Using the first classification detection branch in the trained region proposal model, determining whether each initial region of interest in the initial region of interest set belongs to the background;

[0036] The initial region of interest identified as background by the first classification detection branch is removed from the initial region of interest set by using the region of interest optimization branch in the region proposal model, and the region of interest is adjusted according to the bounding box offset of the first detection prediction mark to obtain an optimized region of interest set for each training sample.

[0037] Optionally, generating an initial set of regions of interest for each training sample specifically includes:

[0038] The spatial images of each training sample are resized to a uniform size, and the size of the deepest feature map is w×h, where w is the number of columns of feature pixels and h is the number of rows of feature pixels.

[0039] Generate n×m regions of interest according to a preset size n and a preset aspect ratio m for each feature pixel point in the deepest feature map of the spatial image, and obtain an initial region of interest set; the initial region of interest set includes w×h×n×m regions of interest;

[0040] Deleting the regions of interest whose areas are smaller than a preset threshold and beyond the boundary in the initial region of interest set from the initial region of interest set;

[0041] The initial regions of interest of about 2k are retained by the non-maximum suppression algorithm; the size of each initial region of interest is (y1, x1, y2, x2), where x1, y1 are the coordinates of the upper left corner of the initial region of interest, and x2, y2 are the coordinates of the lower right corner of the initial region of interest;

[0042] The coordinates of the upper left corner and the lower right corner of each initial region of interest are normalized to between 0 and 1.

[0043] Optionally, the detection and segmentation model adopts a multi-task learning framework, and the second classification detection branch and the segmentation branch share a ResNet skeleton network and jointly update the parameters of the ResNet skeleton network.

[0044] According to the specific invention content provided by the present invention, the present invention discloses the following technical effects:

[0045] The present invention provides a method for synchronous detection and segmentation of space targets, comprising constructing a training sample set; generating an initial region of interest set of each training sample in the training sample set; extracting a multi-layer feature map of each training sample using a ResNet-FPN network, and fusing the features of all other layers on each layer of the feature map to obtain a multi-layer fused feature map; extracting multiple feature matrices from the multi-layer fused feature map of the corresponding training sample; constructing a target detection and segmentation network including a region proposal model and a detection and segmentation model, determining whether an initial region of interest belongs to a background through the region proposal model, and screening and optimizing the initial region of interest in the initial region of interest set; performing category label prediction, detection label prediction and segmentation label prediction on multiple target components in a space image through a detection and segmentation model; and adopting a multi-classification loss Focal Loss is used as its loss function; the region proposal model and the detection and segmentation model are iteratively trained respectively; the trained target detection and segmentation network is used to detect and segment the spatial targets in the image to be detected; the present invention adopts ResNet-FPN to extract multiple layers of feature maps of different scales, and further fuses the multiple layers of feature maps of different scales, fuses the features of all layers on the feature maps of different scales, and retains the edge shape and other information of the shallow network and the semantic information of the deep network as much as possible. Finally, the feature expression ability obtained is stronger, and the effect of dealing with problems such as omission of small objects, geometric transformation, and image degradation is more robust; multi-classification loss Focal Loss is designed as the loss function for component classification detection to avoid the problem of uneven category samples when mining difficult samples. Without losing the inference rate, the detection and segmentation effects of spatial targets can be kept stable.

[0046] In addition, before training the detection and segmentation model, multiple feature matrices of the same scale for each region of interest are extracted through the adaptive ROIAlign algorithm, and the final feature matrix of each region of interest is obtained by superposition; the characteristic of this algorithm is that the sizes of the optimized feature matrices obtained are the same. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0048] Figure 1 A flowchart of a method for synchronous detection and segmentation of space targets provided in Example 1 of the present invention;

[0049] Figure 2This is a flowchart of step S3 in the method provided in Example 1 of the present invention;

[0050] Figure 3 This is a flowchart of step S6 in the method provided in Example 1 of the present invention;

[0051] Figure 4 This is a flowchart of step S61 in the method provided in Example 1 of the present invention;

[0052] Figure 5 This is a flowchart of step S65 in the method provided in Example 1 of the present invention;

[0053] Figure 6 This is a schematic diagram of the target detection and segmentation network structure in the method provided in Example 1 of the present invention. DETAILED DESCRIPTION

[0054] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0055] The purpose of the present invention is to provide a method for synchronous detection and segmentation of space targets, so as to solve the problems of easy omission of small objects, image degradation and uneven category samples existing in the existing instance segmentation technology.

[0056] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0057] Embodiment 1:

[0058] like Figure 1 As shown, the present invention provides a method for synchronous detection and segmentation of space targets, comprising the following steps:

[0059] S1. Construct a training sample set including several training samples; the training samples include a spatial image and image tags corresponding to the spatial image; the image tags include detection tags, segmentation tags and category tags in the spatial image, and the category tags include category tags of multiple target components.

[0060] To simulate the real scene of shooting a small satellite in a space environment, the small satellite model is placed on a model rotator and rotated counterclockwise at a speed of 0.5° per second to ensure that the entire background environment is completely dark. Only one light source is used to illuminate the small satellite from the side. The monocular camera is fixed on a guide rail that can move forward and backward. The camera is facing the small satellite model, and the guide rail is controlled to capture images of the small satellite at various angles at different distances. Four images are taken per second, and only one is added to the training set. The labeling software labelme is used to label each target component with category labels and segmentation labels. The detection label does not need to be manually labeled, and can be generated from the segmentation label through code.

[0061] S2. Generate an initial set of regions of interest for each training sample, specifically including:

[0062] S21. Unify the size of the spatial images of each training sample to 640×640, and the size of the deepest feature map is 28×28.

[0063] S22, generating n×m regions of interest according to a preset size n and a preset aspect ratio m for each pixel point in the spatial image, and obtaining an initial region of interest set; the initial region of interest set includes 28×28×n×m regions of interest. In this embodiment, each feature point of the deepest feature map of the spatial image is taken as the center point, and 9 regions of interest are generated according to three sizes {32, 64, 128} and three aspect ratios {1:2, 1:1, 2:1}.

[0064] S23: Delete the regions of interest whose areas are smaller than a preset threshold and exceed the boundary in the initial region of interest set from the initial region of interest set.

[0065] S24. Retain about 2k initial regions of interest by a non-maximum suppression algorithm; the size of each initial region of interest is (y1, x1, y2, x2), where x1, y1 are the coordinates of the upper left corner of the initial region of interest, and x2, y2 are the coordinates of the lower right corner of the initial region of interest.

[0066] S25, normalizing the values ​​of the coordinates of the upper left corner and the lower right corner of each initial region of interest to between 0 and 1.

[0067] S3. Use the ResNet-FPN network to extract the multi-layer feature map of each training sample, and fuse the features of all other feature maps on each layer of feature map to obtain the multi-layer fused feature map of each training sample.

[0068] like Figure 2 As shown, step S3 specifically includes:

[0069] S31. Perform feature extraction on the spatial image of each training sample through the ResNet-FPN network to obtain a multi-layer feature map; in each of two adjacent layers of feature maps, the scale of the shallow feature map is upsampled by twice the scale of the deep feature map.

[0070] In this embodiment, the spatial image is passed through the ResNet-FPN network to obtain P 2 , P 3 , P 4 , P 5 Feature maps of four scales from large to small, P i The scale is equivalent to P i+1 Double upsampling.

[0071] S32, fuse the features of each layer of feature maps with the features of all other feature maps to obtain a multi-layer fused feature map; wherein each layer of feature map is fused with the down-sampling of all its shallow feature maps and the up-sampling of all its deep feature maps. 2 Need to merge P 3 Double upsampling, P 4 Four times upsampling and P 5 Eight times upsampling, P 3 Need to merge P 2 The double downsampling, P 4 The double upsampling and P 5 Through the above network structure, four scale feature maps P' can be obtained that fully retain the shallow edge shape information and deep semantic information. 2 , P' 3 , P' 4 , P' 5 .

[0072] S4. Extracting a feature matrix of each initial region of interest of the training sample from the multi-layer fusion feature map of the training sample.

[0073] S5. Construct a target detection and segmentation network; the target detection and segmentation network includes a region proposal model and a detection segmentation model; the region proposal model is used to determine whether the corresponding initial region of interest belongs to the background according to the feature matrix of each initial region of interest, and to screen and optimize the initial regions of interest in the initial region of interest set; the detection segmentation model is used to perform category label prediction, detection label prediction and segmentation label prediction for multiple target components in the spatial image.

[0074] S6. Using the multi-layer fusion feature map of each training sample, the initial region of interest set, the feature matrix of each initial region of interest and the training sample set, the region proposal model and the detection and segmentation model are trained respectively to obtain a trained target detection and segmentation network.

[0075] according to Figure 3 As shown, step S6 specifically includes:

[0076] S61. Iteratively train the first classification detection branch in the region proposal model according to the initial region of interest set, the feature matrix of each initial region of interest and the training sample set to obtain a trained region proposal model; in the process of iteratively training the first classification detection branch, use the feature matrix of each initial region of interest as the input of the first classification detection branch, and use the category label and detection label of the corresponding region of interest in the spatial image as the target output of the first classification detection branch.

[0077] according to Figure 4 As shown, step S61 specifically includes:

[0078] S611, input the feature matrix of each initial region of interest into the first classification detection branch in sequence to obtain the first category prediction result and the first detection prediction result. Each feature matrix passes through a 3*3 convolution and several fully connected layers, and the output logical value is used as the category prediction result; according to the category prediction result, it is determined whether each initial region of interest belongs to the foreground or the background.

[0079] S612: Calculate the first category prediction loss of the first classification detection branch according to the first category prediction result and the category label of the corresponding region of interest, as follows:

[0080]

[0081] Among them, y represents the actual category label, Represents the prediction result of the first category.

[0082] S613: Calculate the first detection prediction loss of the first classification detection branch according to the first detection prediction result and the detection mark of the corresponding region of interest. For the initial region of interest that does not belong to the background, its detection mark is predicted again, and the initial region of interest that belongs to the background only contributes to the first category prediction loss, and does not contribute to the first detection prediction loss. The first detection prediction loss L det as follows:

[0083]

[0084] Where x represents the difference between the first detection prediction result and the true detection mark.

[0085] S614. Update the parameters of the first classification detection branch through a back propagation algorithm according to the first category prediction loss and the first detection prediction loss until an iteration termination condition is met.

[0086] S62: Utilize the trained region proposal model to optimize the initial region of interest set of each training sample to obtain an optimized region of interest set of each training sample.

[0087] Step S62 specifically includes:

[0088] S621: Using the first classification detection branch in the trained region proposal model, determine whether each initial region of interest in the initial region of interest set belongs to the background.

[0089] S622. Utilize the region of interest optimization branch in the region proposal model to remove the initial region of interest identified as background by the first classification detection branch from the initial region of interest set, and adjust the region of interest according to the bounding box offset of the first detection prediction mark to obtain an optimized region of interest set for each training sample.

[0090] S63. Based on the set of optimized regions of interest of each training sample, an adaptive ROIAlign algorithm is used to extract multiple optimized feature matrices corresponding to each optimized region of interest from the multi-layer fusion feature map of the corresponding training sample. The characteristic of this algorithm is that the sizes of the obtained optimized feature matrices are the same.

[0091] S64, after passing through a fully connected layer, multiple optimized feature matrices corresponding to each optimized region of interest are added and fused to obtain a fused optimized feature matrix corresponding to each optimized region of interest.

[0092] S65. According to the fused optimized feature matrix set of each training sample and the image label of each training sample, the second classification detection branch and the segmentation branch in the detection and segmentation model are iteratively trained in turn to obtain a trained detection and segmentation model; in the process of iterative training of the second classification detection branch, the fused optimized feature matrix set of the training samples is used as the input of the second classification detection branch, and the category label and detection label in the training sample are used as the target output of the second classification detection branch; in the process of iterative training of the segmentation branch, the fused optimized feature matrix set of the training samples is used as the input of the segmentation branch, and the segmentation label in the training sample is used as the output of the segmentation branch; the fused optimized feature matrix set of each training sample includes the fused optimized feature matrices corresponding to all optimized regions of interest of each training sample.

[0093] according to Figure 5 As shown, step S65 specifically includes:

[0094] S651. Input the fusion optimization feature matrix set of each training sample into the second classification detection branch in sequence to obtain a second category prediction result and a second detection prediction result.

[0095] In this embodiment, the fusion optimized feature matrix set of the training samples passes through several fully connected layers to obtain the second category prediction result [number of regions of interest, number of categories] and the second detection prediction result [number of regions of interest, (y1', x1', y2', x2')].

[0096] S652, input the fused optimized feature matrix set of each training sample into the segmentation branch in sequence to obtain a segmentation prediction result. In this embodiment, the fused optimized feature matrix set of the training sample is passed through a deconvolution layer to obtain a segmentation prediction result [28, 28, number of categories].

[0097] S653: Calculate the second category prediction loss of the second classification detection branch according to the second category prediction result and the category label in the corresponding training sample; use the multi-classification Focal Loss suitable for the detection task as the classification loss L cls ,as follows:

[0098]

[0099] The parameter to be adjusted γ can be used to adjust the rate at which the weight of a simple sample is reduced. In this embodiment, the parameter takes an empirical value of 2.

[0100] S654: Calculate the second detection prediction loss L of the second classification detection branch according to the second detection prediction result and the detection mark in the corresponding training sample. det ; The loss function of the second detection prediction loss is consistent with the loss function of the first detection prediction loss.

[0101] S655: Calculate the segmentation prediction loss of the segmentation branch according to the segmentation prediction result and the segmentation mark in the corresponding training sample. The segmentation prediction loss L mask as follows:

[0102]

[0103] Among them, p is each pixel point on the segmentation prediction result, k is the category of the component, and y p represents the true segmentation label of the p-th pixel, Represents the predicted value of the p-th pixel on the prediction mask of class k.

[0104] S656. According to the second category prediction loss and the second detection prediction loss, the parameters of the second classification detection branch are updated through the back propagation algorithm; according to the segmentation prediction loss, the parameters of the segmentation branch are updated through the back propagation algorithm until the iteration termination condition is met; the detection and segmentation model adopts a multi-task learning framework, and the second classification detection branch and the segmentation branch share the ResNet skeleton network and jointly update the parameters of the ResNet skeleton network. The total loss of the detection and segmentation model is as follows:

[0105] L=L cls +L det +L mask (5)

[0106] S7. Use the trained target detection and segmentation network to detect and segment the spatial targets in the image to be detected. Input the image to be predicted. The input form can be local reading or real-time shooting. First, the region proposal model is used to filter and optimize the set of regions of interest, and then it is input into the detection and segmentation model, and the image containing the detection and segmentation results is output.

[0107] like Figure 6 As shown in FIG, a schematic diagram of the target detection and segmentation network structure constructed by the present invention is shown. Table 1 compares the detection accuracy of the method provided by the present invention with that of the classic target detection method Faster RCNN, and compares the segmentation average intersection-over-union ratio with that of the classic instance segmentation method Mask RCNN on various components of the spatial target data set. The results show that the detection and segmentation effects of the method provided by the present invention are better on this complex data set.

[0108] Table 1 Comparison of the detection accuracy and segmentation average intersection-over-union ratio of the proposed method on various components of the space target data set with existing advanced methods

[0109]

[0110]

[0111] The program part of the technology can be considered as a "product" or "manufactured product" in the form of executable code and / or related data, which is participated in or realized by computer-readable media. Tangible and permanent storage media can include any memory or storage used by any computer, processor, or similar device or related module. For example, various semiconductor memories, tape drives, disk drives or any similar devices that can provide storage functions for software.

[0112] All or part of the software may sometimes communicate over a network, such as the Internet or other communications network. Such communications can load software from one computer device or processor to another. For example: from a server or host computer of a video target detection device to a hardware platform of a computer environment, or other computer environment that implements the system, or a system with similar functions related to providing information required for target detection. Therefore, another medium capable of transmitting software elements can also be used as a physical connection between local devices, such as light waves, radio waves, electromagnetic waves, etc., transmitted through cables, optical cables or air. Physical media used to carry carriers, such as cables, wireless connections or optical cables and the like, can also be considered as media that carry software. As used herein, unless limited to tangible "storage" media, other terms referring to computer or machine "readable media" refer to media involved in the process of the processor executing any instructions.

[0113] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core ideas of the present invention. Those skilled in the art should understand that the modules or steps of the present invention can be implemented by a general-purpose computer device. Alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into integrated circuit modules, or multiple modules or steps can be made into a single integrated circuit module. The present invention is not limited to any specific combination of hardware and software.

[0114] Meanwhile, for those skilled in the art, according to the concept of the present invention, there will be changes in the specific implementation and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A method for synchronous detection and segmentation of space targets, characterized in that: The method comprises: Constructing a training sample set including a plurality of training samples; the training samples include a spatial image and an image tag corresponding to the spatial image; the image tag includes a detection tag, a segmentation tag and a category tag in the spatial image, and the category tag includes category tags of multiple target components; Generate an initial set of regions of interest for each training sample; The ResNet-FPN network is used to extract the multi-layer feature map of each training sample, and the features of all other feature maps are fused on each layer of feature map to obtain the multi-layer fused feature map of each training sample; Extracting a feature matrix of each initial region of interest of the training sample from the multi-layer fusion feature map of the training sample; Constructing a target detection and segmentation network; the target detection and segmentation network includes a region proposal model and a detection and segmentation model; the region proposal model is used to determine whether the corresponding initial region of interest belongs to the background according to the feature matrix of each initial region of interest, and to screen and optimize the initial regions of interest in the initial region of interest set; the detection and segmentation model is used to predict category labels, detection labels and segmentation labels for various target components in the spatial image; when predicting the category labels of the region of interest in the region proposal model, a binary classification loss is used as its loss function, and when predicting the category labels in the detection and segmentation model, a multi-classification loss FocalLoss is used as its loss function; Using the multi-layer fusion feature map of each training sample, the initial region of interest set, the feature matrix of each initial region of interest and the training sample set, the region proposal model and the detection and segmentation model are trained respectively to obtain a trained target detection and segmentation network; Use the trained target detection and segmentation network to detect and segment the spatial targets in the image to be detected; The region proposal model and the detection segmentation model are trained respectively by using the multi-layer fusion feature map of each training sample, the initial region of interest set, the feature matrix of each initial region of interest and the training sample set, specifically including: Iteratively training a first classification detection branch in the region proposal model according to the initial region of interest set, the feature matrix of each initial region of interest and the training sample set to obtain a trained region proposal model; in the process of iteratively training the first classification detection branch, taking each feature matrix of each initial region of interest as an input of the first classification detection branch, and taking the category label and detection label of the corresponding region of interest in the spatial image as the target output of the first classification detection branch; Optimizing the initial set of regions of interest of each training sample using the trained region proposal model to obtain an optimized set of regions of interest of each training sample; According to the set of optimized regions of interest of each training sample, an adaptive ROIAlign algorithm is used to extract multiple optimized feature matrices corresponding to each optimized region of interest from the multi-layer fusion feature map of the corresponding training sample; the sizes of each optimized feature matrix are the same; Add and fuse multiple optimization feature matrices corresponding to each optimized region of interest to obtain a fused optimization feature matrix corresponding to each optimized region of interest; According to the fused optimized feature matrix set of each training sample and the image label of each training sample, the second classification detection branch and the segmentation branch in the detection and segmentation model are iteratively trained in turn to obtain a trained detection and segmentation model; in the process of iterative training of the second classification detection branch, the fused optimized feature matrix set of the training samples is used as the input of the second classification detection branch, and the category label and detection label in the training sample are used as the target output of the second classification detection branch; in the process of iterative training of the segmentation branch, the fused optimized feature matrix set of the training samples is used as the input of the segmentation branch, and the segmentation label in the training sample is used as the output of the segmentation branch; the fused optimized feature matrix set of each training sample includes the fused optimized feature matrices corresponding to all optimized regions of interest of each training sample.

2. The method according to claim 1, characterized in that The ResNet-FPN network is used to extract the multi-layer feature map of each training sample, and the features of all other feature maps are fused on each layer of the feature map to obtain the multi-layer fused feature map of each training sample, which specifically includes: The spatial image of each training sample is subjected to feature extraction through the ResNet-FPN network to obtain a multi-layer feature map; Each layer of feature map is fused with the features of all other feature maps to obtain a multi-layer fused feature map; wherein each layer of feature map is fused with the down-sampling of all its shallow feature maps and the up-sampling of all its deep feature maps.

3. The method according to claim 1, characterized in that The iterative training of the first classification detection branch in the region proposal model according to the initial region of interest set, the feature matrix of each initial region of interest and the training sample set specifically includes: Inputting the feature matrix of each initial region of interest into the first classification detection branch in sequence to obtain a first category prediction result and a first detection prediction result; Calculating a first category prediction loss of the first classification detection branch according to the first category prediction result and the category label of the corresponding region of interest; Calculating a first detection prediction loss of the first classification detection branch according to the first detection prediction result and the detection mark of the corresponding region of interest; According to the first category prediction loss and the first detection prediction loss, the parameters of the first classification detection branch are updated through a back propagation algorithm until an iteration termination condition is met.

4. The method according to claim 1, characterized in that: The iterative training of the second classification detection branch and the segmentation branch in the detection and segmentation model in sequence according to the fusion optimization feature matrix set of each training sample and the image label of the corresponding training sample specifically includes: Inputting the fusion optimized feature matrix set of each training sample into the second classification detection branch in sequence to obtain a second category prediction result and a second detection prediction result; Inputting the fusion optimized feature matrix set of each training sample into the segmentation branch in sequence to obtain a segmentation prediction result; Calculating a second category prediction loss of the second classification detection branch according to the second category prediction result and the category label in the corresponding training sample; Calculating a second detection prediction loss of the second classification detection branch according to the second detection prediction result and the detection mark in the corresponding training sample; Calculating the segmentation prediction loss of the segmentation branch according to the segmentation prediction result and the segmentation mark in the corresponding training sample; According to the second category prediction loss and the second detection prediction loss, the parameters of the second classification detection branch are updated through a back propagation algorithm; according to the segmentation prediction loss, the parameters of the segmentation branch are updated through a back propagation algorithm until an iteration termination condition is met.

5. The method according to claim 3, characterized in that: The optimizing the initial region of interest set of each training sample by using the trained region proposal model specifically includes: Using the first classification detection branch in the trained region proposal model, determining whether each initial region of interest in the initial region of interest set belongs to the background; The initial region of interest identified as background by the first classification detection branch is removed from the initial region of interest set by using the region of interest optimization branch in the region proposal model, and the region of interest is adjusted according to the bounding box offset of the first detection prediction loss to obtain an optimized region of interest set for each training sample.

6. The method according to claim 1, characterized in that The generating of the initial set of regions of interest for each training sample specifically includes: The spatial images of each training sample are resized to a uniform size, and the size of the deepest feature map is w×h, where w is the number of columns of feature pixels and h is the number of rows of feature pixels. Generate n×m regions of interest according to a preset size n and a preset aspect ratio m for each feature pixel point in the deepest feature map of the spatial image, and obtain an initial region of interest set; the initial region of interest set includes w×h×n×m regions of interest; Deleting the regions of interest whose areas are smaller than a preset threshold and beyond the boundary in the initial region of interest set from the initial region of interest set; 2,000 initial regions of interest are retained by a non-maximum suppression algorithm; the size of each initial region of interest is (y1, x1, y2, x2), where x1, y1 are the coordinates of the upper left corner of the initial region of interest, and x2, y2 are the coordinates of the lower right corner of the initial region of interest; The coordinates of the upper left corner and the lower right corner of each initial region of interest are normalized to between 0 and 1.

7. The method according to claim 1, characterized in that The detection and segmentation model adopts a multi-task learning framework, and the second classification detection branch and the segmentation branch share a ResNet skeleton network and jointly update the parameters of the ResNet skeleton network.

Citation Information

Patent Citations

  • Instance segmentation method fusing hole convolution and edge information

    CN110348445A

  • Small target detection method based on multi-scale images and weighted fusion loss

    CN111461110A

Cited By

  • Region-of-interest (ROI)-based image enhancement using a residual network

    US12573013B2

  • Region-of-interest (ROI)-based image enhancement using a residual network

    US20240005458A1