Example segmentation method based on edge attention enhancement
By introducing edge attention mechanism and adaptive query decoder into the instance segmentation method, the performance bottleneck of the prior art when processing local overlap but not completely occlusion in chromosomal data is solved, and higher segmentation accuracy and better edge processing capabilities are achieved.
Patent Information
- Application Number
- CN202510431423.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art faces performance bottlenecks when dealing with local overlap but not completely obstructed in chromosomal data, resulting in inaccurate segmentation results.
An instance segmentation method based on edge attention enhancement is adopted to extract features through a backbone network of feature pyramid structure, and an adaptive query decoder is used to generate object query embedding vectors, and a BoundaryLoss function is used for model training to improve the edge clarity of segmentation results.
Improve the accuracy and edge clarity of instance segmentation, especially in complex scenarios, effectively removing the mask of repeated segmentation, improving the quality of the final segmentation result.
Smart Images

Figure CN119942131A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology, and specifically is an instance segmentation method based on edge attention enhancement. Background Art
[0002] Chromosome karyotype analysis, also often referred to as chromosome examination in clinical practice, is a key cytogenetic task and is hailed as the gold standard for genetic disease diagnosis. This analysis identifies genetic abnormalities by detecting metaphase chromosome images. The analysis process first involves preparing a complete set of metaphase chromosome microscopic images from a cell sample; then segmenting the chromosome instances, and classifying and pairing these chromosomes by type to form 23 pairs, arranged into a karyotype diagram; finally, these karyotype images are analyzed to detect whether there are genetic abnormalities, thereby obtaining a diagnosis.
[0003] However, the placement structures of chromosome instances under microscope images may vary greatly. It is extremely time-consuming and laborious to accurately determine the components of the same chromosome and precisely segment a single chromosome, even for experienced analysts. In order to alleviate this burden and improve efficiency, artificial intelligence-assisted karyotype analysis can not only significantly reduce the workload of clinicians and improve inspection efficiency, but also greatly shorten the analysis cycle, win precious treatment time for patients, and show extremely high clinical value and social significance.
[0004] In the prior art, instance segmentation is mainly performed through MaskRCNN, yolov5 and other methods to form a two-stage instance segmentation model. For the two-stage instance segmentation model, since its design goal is the instance segmentation task on a general data set, its internal labels are usually large targets, such as people, animals, large objects, etc. The target detection algorithm used is mainly for larger rigid objects, and there are fewer cases of local overlap but incomplete occlusion in the data set. The chromosome instance should be regarded as a tiny non-rigid long-tail target, and there are a large number of cases of local overlap but incomplete occlusion in the chromosome data set. The two-stage instance segmentation model faces certain performance bottlenecks when processing targets with local overlap but incomplete occlusion. This is mainly because the rectangular box used by the upstream network of the model cannot accurately describe the overlapping occlusion relationship between chromosome instances, resulting in the non-maximum suppression algorithm mistakenly identifying two or more overlapping targets as the same target and mistakenly filtering out one of the instances. This prior knowledge limitation limits the performance of the model in dealing with complex small target overlap situations. Summary of the invention
[0005] The purpose of the present invention is to provide an instance segmentation method based on edge attention enhancement to solve the problem raised in the background technology: since there are a large number of local overlaps but incomplete occlusions in the chromosome dataset, the algorithms in the prior art face certain performance bottlenecks when processing targets with local overlaps but incomplete occlusions.
[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: An instance segmentation method based on edge attention enhancement comprises the following steps: Step S1, use the backbone network of the feature pyramid structure to extract features from the image and output a set of two-dimensional feature maps and cls tokens; Step S2, using the adaptive query decoder to dynamically generate a set of vectors based on the cls token for the query input in the attention mechanism; specifically: Step S201, the adaptive query decoder decodes the cls token to obtain Token, denoted by z; Step S202, generating an object query vector E by transforming the token z; decoding the object query vector E to obtain a set of query vectors Q for the image segmentation task; Step S3, evaluating the predicted segmentation result of the instance segmentation model and guiding the instance segmentation model to converge; Step S4, sorting the prediction result set by the prediction confidence score given by the instance segmentation model, and calculating the intersection of the foreground areas of any two mask images in the set, and then comparing the calculated intersection area with the foreground areas of the two mask images themselves, to determine whether the ratio is higher than a preset threshold. If it is higher than the threshold, the prediction result with lower confidence in the mask image is removed from the set; if it is not higher than the threshold, it is retained; after the judgment is completed, the image segmentation is completed.
[0007] According to the above technical solution, in step S1, using the feature pyramid to extract features from the image includes the following steps: Step S101, firstly, the image is cut into image blocks of size P by a segmentation operation; Step S102, structurally expressing the edges in the image blocks cropped in step S101, representing each sub-image as a set of g = (U, Q, W), and using the set to express the boundary information in the image, where U (u, v) is the pixel coordinate; Q is the direction of the line, represented as a unit vector; The corresponding angle value that the ray needs to rotate in the clockwise direction; Step S103, the sub-graphs g after the structured expression outputted in step S102 are assembled into a large graph G, G={g}; and the large graph G is used as the input image x, and a feature extraction operation is performed using a backbone network with a feature pyramid structure to obtain an output y; Step S104, again using a convolution kernel to extract features from the output y of step S103 according to a certain step length, to obtain a feature map F n , the feature map F n A feature pyramid F is obtained by the collection of , which is used to effectively extract multi-scale features in multi-type pre-training tasks and improve the performance of the model in complex image tasks.
[0008] According to the above technical solution, in step S103, the feature vectorization operation is specifically performed as follows: in, is the vector of cls tokens, is the feature of the sub-image; is the position embedding code, Operates for multi-head edge self-attention mechanism; For the layernorm operation, is a multi-layer perceptron; the backbone network is a recurrent neural network structure with a total of m layers. The backbone network needs to perform multiple repeated calculations on the input. represents the initial state of the input vector when it enters the backbone network, is the state of the input vector after looping through m layers, for An intermediate state of layer calculation, Backbone Network The state of the layer.
[0009] According to the above technical solution, in step S104, the feature map F n Size: In the formula, represents the step length, Indicates the number of channels of the feature map, Indicates the height of the image. Indicates the width of the image; feature pyramid .
[0010] According to the above technical solution, in step S103, the input image x is subjected to feature extraction through the backbone network as follows: In the formula, represents the object query vector; Indicates that the input image x is obtained by feature extraction through the backbone network mark; , Indicates the corresponding angle value that the ray needs to rotate in the clockwise direction; , are learnable weights that can be converged by back-propagation, represents the linear rectification function.
[0011] According to the above technical solution, in step S202, the decoding calculation is as follows: In the formula, represent The dimension of V is obtained by multiplying the converted image feature vector y calculated in step S103 with the weight W. ; represents the embedding tag, Q represents the query vector, express The transpose of .
[0012] According to the above technical solution, in step S3, the model prediction segmentation result is evaluated and the model convergence is guided as follows: Step S301, using the BoundaryLoss function to pre-process the marked mask image; Step S302, construct a boundary loss function Dist(∂G,∂S), where ∂G is a representation of the boundary of the ground truth region G, ∂Sθ is the boundary of the segmentation region defined by the network output; the distance change between ∂S and ∂G; Among them, P is the boundary A point on Representatives along On the normal direction boundary The point where the two boundaries meet is the intersection point; the distance between the predicted boundary and the true boundary is represented by calculating the distance integral between the differential points on the boundary.
[0013] According to the above technical solution, in step S4, the ratio of the intersection of the foreground area between the two mask images to the foreground area of the mask image itself and the confidence are calculated to jointly determine whether the mask image is repeatedly segmented, specifically: For any two mask images, the intersection of the two is calculated using the following formula: In the formula, represents the intersection area, The n in is the intersection operation, , for any two mask images.
[0014] According to the above technical solution, the foreground area of each mask image is calculated Intersection Area Ratio , as shown below: In the formula, For each mask image, its own foreground area Intersection Area The ratio of is the intersection area, is the foreground area of each mask image.
[0015] According to the above technical solution, the set is judged by setting a threshold θ and a new matrix Q is generated: in is the unique mapping k = map (i, j) of i and j, θ represents the manually set threshold. In order to prevent the two mask images from being discarded due to complete overlap, the algorithm finally evaluates whether each mask is dominated by other masks with higher scores. If there are both (i, j) and (j, i) mask pairs in , the mask with higher confidence will be retained.
[0016] Compared with the prior art, the present invention has the following beneficial effects: In the present invention, feature pyramid is used for feature extraction, which can effectively capture the features of objects at different scales, thereby improving the accuracy of segmentation; object query embedding vectors are generated by Adaptive Query Decoder, which provides powerful prior information for the model and enhances the model's understanding and segmentation capabilities of objects. The BoundaryLoss function introduces attention to edge information in model training, which can promote the model to process object boundaries more accurately, thereby improving the edge clarity of the segmentation results. By calculating the ratio and confidence of the foreground area intersection between mask images and their own foreground area, the mask of repeated segmentation can be effectively removed, the quality of the final segmentation result can be improved, and redundant information can be reduced.
[0017] The method of the present invention can achieve higher segmentation accuracy and better segmentation effect in instance segmentation tasks, especially the ability to process object edges in complex scenes is more prominent. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a schematic diagram of pyramid feature extraction of the present invention; Figure 2A schematic diagram of expressing the subgraph encoding of the present invention as an edge structure; Figure 3 This is a schematic diagram of the structure of the adaptive query decoder of the present invention; Figure 4 It is a schematic diagram of the backbone network structure of the present invention; wherein the backbone network encodes the input image and outputs a feature map and a cls token; Figure 5 This is a comparison diagram before and after chromosome segmentation of the present invention, where a is the source code model reasoning result (after combining with the mask filtering algorithm), b is the improved model reasoning result (after combining with the mask filtering algorithm), c is the original model prediction result, and d is the effect after processing by the mask filtering algorithm. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] Embodiment 1 like Figure 1 As shown, an instance segmentation method based on edge attention enhancement includes the following steps: Step S1, use the backbone network of the feature pyramid structure to extract features from the image, and output a set (usually 4) of two-dimensional feature maps of different sizes and cls tokens. The two-dimensional feature maps are used for the detection and segmentation of subsequent instance segmentation targets, and the cls tokens are used for the subsequent generation of dynamic query vectors. Step S2: Use the Adaptive Query Decoder to dynamically generate a set of fixed-length vectors based on the input cls tokens. The vector contains the instance information of the input image and is named the object query vector, which is used as the query input in the attention mechanism. Step S201, the adaptive query decoder decodes the cls token, and the specific process is as follows: let the input image be x, and extract features through the backbone network of the feature pyramid structure to obtain the [cls] token, which is represented by z; Step S202: Transform the token z through FFN (feed forward network) to generate an object query vector E. Next, use the Transformer module to decode the tag embedding E to obtain a set of query vectors Q for the image segmentation task.
[0021] Step S3, using the BoundaryLoss function to evaluate the model prediction segmentation results and guide the model convergence; Step S4, mask filtering algorithm, sorts the prediction result set by the prediction confidence score given by the model, and calculates the ratio of the intersection of the foreground area between the two mask images in the set to their own foreground area; determines whether the ratio is higher than a preset threshold (the threshold is usually set to a decimal between 0 and 1, that is, if the ratio of the intersection of the two mask images to their own area is greater than the threshold, it is considered as repeated segmentation), and removes the prediction results with lower confidence from the set in the mask image pairs (pairwise maskresult) that are higher than the threshold.
[0022] In the present invention, feature pyramid is used for feature extraction, which can effectively capture the features of objects at different scales, thereby improving the accuracy of segmentation; object query embedding vectors are generated by Adaptive Query Decoder, which provides powerful prior information for the model and enhances the model's understanding and segmentation capabilities of objects. The BoundaryLoss function introduces attention to edge information in model training, which can promote the model to process object boundaries more accurately, thereby improving the edge clarity of the segmentation results. By calculating the ratio and confidence of the foreground area intersection between mask images and their own foreground area, the mask of repeated segmentation can be effectively removed, the quality of the final segmentation result can be improved, and redundant information can be reduced.
[0023] The method of the present invention can achieve higher segmentation accuracy and better segmentation effect in instance segmentation tasks, especially the ability to process object edges in complex scenes is more prominent.
[0024] Embodiment 2 This embodiment provides a specific implementation method: an instance segmentation method based on edge attention enhancement, comprising the following steps: Step S1, use the feature pyramid vision transformer (Pyramid Vision Transformer, hereinafter referred to as PVT) to extract features.
[0025] Step S101, assuming that the input image is H×W×3, the image is first cut into image blocks of a specified size P by a segmentation operation. In this embodiment, P is specified as 16, and the image blocks are segmented into the number of , a sub-image of size P×P×3 ; Step S102, encode the sub-image into an edge structured expression, represent each sub-image as a set of g = (U, Q, W), and use the set to express the boundary information in the image, where U (u, v) is the pixel coordinate, θ is the direction of the line, The corresponding angle value that the ray needs to rotate in the clockwise direction.
[0026] like Figure 2 As shown, the new sub-area image is divided into three areas by lines, and each area is called a widget (a widget is a basic module used to construct a graphical UI, just like Lego. Various complex interfaces can be formed by combining and assembling three widgets).
[0027] Step S103, perform feature vectorization operation on the input sub-image, and the feature vectorization operation is obtained by iterative calculation through the following four formulas: in, is the vector of cls tokens, is the feature of the sub-image; is the position embedding code, Operates for multi-head edge self-attention mechanism; For the layernorm operation, is a multi-layer perceptron; the backbone network is a recurrent neural network structure with a total of m layers. The backbone network needs to perform multiple repeated calculations on the input. represents the initial state of the input vector when it enters the backbone network, is the state of the input vector after looping through m layers, for An intermediate state of layer calculation, Backbone Network The state of the layer.
[0028] It contains two operators: gather and slice. The gather operation replaces each block area (widget) with the weighted average of the pixel values contained in the area. The slice operation calculates the distance map and feature map in all slices, and fits the target result by continuously iterating the output results of the two operators to achieve the expected result.
[0029] Step S104, again use the convolution kernel to extract features from y in step S102 according to the step lengths of 4, 8, 16 and 32. Taking the first layer as an example, the feature map F 1 Size: Among them, C 1 is the number of channels of the first layer feature map; the feature maps F are obtained in sequence 2 、F 3 and F 4; The sizes of the feature maps obtained in this way are 1 / 4, 1 / 8, 1 / 16 and 1 / 32 respectively relative to the input image; This series of operations produces a feature pyramid F = {F 1 , F 2 , F 3 , ... F n In this way, multi-scale features can be effectively extracted while performing multiple types of pre-training tasks, thereby improving the performance of the model in complex image tasks.
[0030] The boundary attention mechanism gradually optimizes the local boundary variable field associated with each pixel by densely applying local attention operations around the pixel. This enables the model to extract boundary information from the input image and generate a series of overlapping geometric primitives. These geometric primitives are used to generate unsigned distance functions of image boundaries, boundary-aware smoothed channel values, and soft local attention distributions associated with each pixel. The specific implementation is to aggregate the information of neighboring pixels to form a higher-dimensional representation so that the network can learn meaningful hidden states. Such operations help integrate local information into a more global perspective, thereby helping the network better understand the overall structure and characteristics of the image. These high-dimensional representations are then divided into smaller blocks for local operations and analysis. The slicing operation allows the network to perform more fine-grained processing in local areas, so that the subtle structures and features in the image can be more accurately captured.
[0031] Step S2 proposes to use an adaptive query decoder (AQD) to generate an object query embedding vector as a priori input to the model (e.g. Figure 4 The core goal of this module is to no longer rely on a set of preset fixed query vectors, but to give the network the ability to generate adaptive query vectors. The benefit of this is that the network can dynamically generate query vectors based on the changes in each specific input, thereby better coping with changes in data distribution and unseen scenarios.
[0032] like Figure 3 As shown in the figure, the adaptive query decoder consists of a forward pass neural network and multiple Transformer decoding blocks. The input on the left is the cls token output from the backbone network (such as PVT). These tokens contain the global feature information of the image. As the initial input of the adaptive query decoder, the first layer is a forward pass neural network (Feed-Forward Network, hereinafter referred to as FFN).
[0033] Step S201, the adaptive query decoder decodes the cls token, and the input image x is extracted through the backbone network to obtain Mark, represented by z.
[0034] In the formula, represents the object query vector; Indicates that the input image x is obtained by feature extraction through the backbone network mark; , Indicates the corresponding angle value that the ray needs to rotate in the clockwise direction; , are learnable weights that can be converged by back-propagation, represents the linear rectification function.
[0035] Transform the tag z through FFN to generate tag embedding Next, we use the transformer module to embed the tokens Decode to obtain a set of query vectors Q for image segmentation tasks. The decoding calculation is as follows: Where Q represents the query vector, represent The dimension of V is obtained by multiplying the converted image feature vector y calculated in step S103 with the weight W. ; represents the embedding tag, Q represents the query vector, express The transpose of . Taking this approach enables the model to formulate unique queries for different input images, thus achieving true adaptability and enhancing its segmentation capabilities. It is worth noting that during the back-propagation stage inside AQD, a stop gradient strategy is implemented. This decision was made based on what was observed in the experiments, but by The query generation process performed by the markers may interfere with the stable convergence of the backbone network during back-propagation. Therefore, in order to improve the final segmentation result, the gradient propagation is stopped before the gradient is returned to the backbone network.
[0036] Step S3: During training, the BoundaryLoss function is used to evaluate the model prediction segmentation results and guide the model convergence. BoundaryLoss preprocesses the labeled mask image, and for each foreground element (non-zero element) in the original binary mask image, calculates the shortest distance from the current foreground element to the background element (any zero-value element), thereby calculating the input distance map to calculate the model prediction results.
[0037] Step S301, BoundaryLoss preprocesses the marked mask image; First, the image is binarized, marking the target area as 1 and the background area as 0.
[0038] Then, for each pixel, calculate its Euclidean distance to all pixels in the foreground area, and select the minimum value as the distance value of the pixel.
[0039] Finally, in the obtained distance map, the distance value inside the target area is 0, and the distance value of the background area is positive.
[0040] Step S302, construct a boundary loss function Dist(∂G,∂S), which takes the form of a distance metric in the region boundary space in Ω, where ∂G is a representation of the boundary of the ground truth region G (such as the set sum of all points on the boundary), and ∂Sθ is the boundary of the segmented region defined by the network output. Consider the following representation of asymmetric L2distance in shape space, which evaluates the change in distance between two adjacent boundaries ∂S and ∂G; Among them, P is the boundary A point on Representatives along On the normal direction boundary The point where the two boundaries meet is the intersection point; the distance between the predicted boundary and the true boundary is represented by calculating the L2 distance integral between the differential points on the boundary.
[0041] Step S4, mask filtering algorithm. The one-stage deep learning chromosome instance segmentation algorithm is designed to output a mask set that is much larger than the number of detected instances, so the model output mask must have redundant results. In addition, there may be multiple high-confidence repeated instance masks in the mask set at the same time, making it difficult to effectively filter the final result only by means of a threshold. In response to this behavior of the model, this embodiment proposes a mask filtering algorithm, which calculates the ratio of the intersection of the foreground area between the two mask images and their own foreground area and the confidence to jointly determine whether the mask image is repeatedly segmented and remove it from the prediction result. The result is as follows: Figure 5 shown.
[0042] Figure 5 (a), (b), (c), (d) respectively show the comparison of the processing results of the mask filtering algorithm in the case of very few repeated segmentation instances and more repeated segmentation instances. Figure 5 There is a duplicate instance in the instance segmentation result in (a). After processing in step S4, the result is as follows: Figure 5 As shown in (b), the existing instances of repeated segmentation are removed by precise filtering.
[0043] Figure 5(c) shows that there are a lot of chromosome repetitions. After processing in step S4, the results are as follows: Figure 5 As shown in (d), instances of repeated segmentation are removed by precise filtering.
[0044] The algorithm is implemented as follows. For any two mask images, the intersection of the two is calculated using the following formula: In the formula, represents the intersection area, The n in is the intersection operation, , for any two mask images.
[0045] Calculate the foreground area of each mask image Intersection Area Ratio ,in The unique mapping k = map (i, j) for i and j is as follows: In the formula, For each mask image, its own foreground area Intersection Area The ratio of is the intersection area, is the foreground area of each mask image.
[0046] By setting the threshold, the collection is judged and a new matrix is generated: To prevent the situation where two masks are discarded due to complete overlap, the algorithm finally evaluates whether each mask is dominated by other masks with higher scores. If there are both (i, j) and (j, i) mask pairs in , the mask with higher confidence will be retained.
[0047] The present invention utilizes a backbone network with a feature pyramid structure for feature extraction, which can extract rich feature information of multiple scales from images and better capture the features of target objects of different sizes. The adaptive query decoder dynamically generates query vectors according to cls tokens, which can adaptively generate appropriate query inputs for different image contents, making the attention mechanism more targeted, focusing on target objects more accurately, and improving the accuracy of instance segmentation.
[0048] Evaluating the model's predicted segmentation results and guiding the model's convergence helps the model continuously adjust parameters during the training process and gradually learn more accurate segmentation patterns, thereby improving the model's overall performance and segmentation effect.
[0049] By sorting the prediction confidence scores and calculating the mask image foreground area ratio to filter the results, we can effectively remove duplicate or inaccurate predictions and retain segmentation results with high confidence and uniqueness, making the final instance segmentation results more accurate and improving the quality of the segmentation results.
[0050] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0051] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An instance segmentation method based on edge attention enhancement, comprising an instance segmentation model, characterized in that: Image segmentation using the instance segmentation model includes the following steps: Step S1, use the backbone network of the feature pyramid structure to extract features from the image and output a set of two-dimensional feature maps and cls tokens; Step S2, using the adaptive query decoder to dynamically generate a set of vectors based on the cls token for the query input in the attention mechanism; specifically: Step S201, the adaptive query decoder decodes the cls token to obtain Token, denoted by z; Step S202, generating an object query vector E by transforming the token z; decoding the object query vector E to obtain a set of query vectors Q for the image segmentation task; Step S3, evaluating the predicted segmentation result of the instance segmentation model and guiding the instance segmentation model to converge; Step S4, sorting the prediction result set by the prediction confidence score given by the instance segmentation model, and calculating the intersection of the foreground areas of any two mask images in the set, and then comparing the calculated intersection area with the foreground areas of the two mask images themselves, to determine whether the ratio is higher than a preset threshold. If it is higher than the threshold, the prediction result with lower confidence in the mask image is removed from the set; if it is not higher than the threshold, it is retained; after the judgment is completed, the image segmentation is completed.
2. The instance segmentation method based on edge attention enhancement according to claim 1, characterized in that: In step S1, extracting features from an image using a feature pyramid includes the following steps: Step S101, firstly, the image is cut into image blocks of size P by a segmentation operation; Step S102, structurally expressing the edges in the image blocks cropped in step S101, representing each sub-image as a set of g = (U, Q, W), and using the set to express the boundary information in the image, where U (u, v) is the pixel coordinate; Q is the direction of the line, represented as a unit vector; The corresponding angle value that the ray needs to rotate in the clockwise direction; Step S103, the sub-graphs g after the structured expression outputted in step S102 are assembled into a large graph G. ; And take the large image G as the input image x, use the backbone network of the feature pyramid structure to perform feature extraction operations, and obtain the output y; Step S104, again using a convolution kernel to extract features from the output y of step S103 according to a certain step length, to obtain a feature map , the feature map A feature pyramid F is obtained by the collection of , which is used to effectively extract multi-scale features in multi-type pre-training tasks and improve the performance of the model in complex image tasks.
3. The instance segmentation method based on edge attention enhancement according to claim 2, characterized in that: In step S103, the feature vectorization operation is specifically performed as follows: in, is the vector of cls tokens, is the feature of the sub-image; is the position embedding code, Operates for multi-head edge self-attention mechanism; For the layernorm operation, is a multi-layer perceptron; The backbone network is a recurrent neural network structure with a total of m layers. The backbone network needs to perform multiple repeated calculations on the input. represents the initial state of the input vector when it enters the backbone network, is the state of the input vector after looping through m layers, for An intermediate state of layer calculation, Backbone Network The state of the layer.
4. The instance segmentation method based on edge attention enhancement according to claim 3, characterized in that: In step S104, the feature map Size: In the formula, represents the step length, Indicates the number of channels of the feature map, Indicates the height of the image. Indicates the width of the image; Feature Pyramid .
5. The instance segmentation method based on edge attention enhancement according to claim 4, characterized in that: In step S103, the input image x is subjected to feature extraction through the backbone network as follows: In the formula, represents the object query vector; Indicates that the input image x is obtained by feature extraction through the backbone network mark; Indicates the corresponding angle value that the ray needs to rotate in the clockwise direction; are learnable weights that can be converged by back-propagation, represents the linear rectification function.
6. The instance segmentation method based on edge attention enhancement according to claim 5, characterized in that: In step S202, the decoding calculation is as follows: In the formula, represent The dimension V is obtained by multiplying the converted image feature vector y calculated in step S103 with the weight W, specifically: represents the embedding tag, Q represents the query vector, express The transpose of .
7. The instance segmentation method based on edge attention enhancement according to claim 6, characterized in that: In step S3, the model prediction segmentation results are evaluated and the model convergence is guided as follows: Step S301, using the BoundaryLoss function to pre-process the marked mask image; Step S302, construct a boundary loss function Dist(∂G,∂S), where ∂G is a representation of the boundary of the ground truth region G, ∂Sθ is the boundary of the segmentation region defined by the network output; the distance change between ∂S and ∂G; Among them, P is the boundary A point on Representatives along On the normal direction boundary The point where the two boundaries meet is the intersection point; the distance between the predicted boundary and the true boundary is represented by calculating the distance integral between the differential points on the boundary.
8. The instance segmentation method based on edge attention enhancement according to claim 7, characterized in that: In step S4, determining whether the mask image is repeatedly segmented is specifically as follows: For any two mask images, the intersection of the two is calculated using the following formula: In the formula, represents the intersection area, In n For the intersection operation, , for any two mask images.
9. The instance segmentation method based on edge attention enhancement according to claim 8, characterized in that: Calculate the foreground area of each mask image Intersection Area Ratio , as shown below: In the formula, For each mask image, its own foreground area Intersection Area The ratio of is the intersection area, is the foreground area of each mask image.
10. The instance segmentation method based on edge attention enhancement according to claim 9, characterized in that: By setting the threshold θ, the set is judged and a new matrix Q is generated: in is the unique mapping k = map (i, j) of i and j, θ represents the manually set threshold. In order to prevent the two mask images from being discarded due to complete overlap, the algorithm finally evaluates whether each mask is dominated by other masks with higher scores. If there are both (i, j) and (j, i) mask pairs in , the mask with higher confidence will be retained.
Citation Information
Patent Citations
Character region boundary detection method and device, equipment and storage medium
CN112560857A
Remote sensing scene classification method and system based on homogeneity and heterogeneity Transformers
CN114091514A
Scene image semantic segmentation method based on transform structure
CN117315241A
High-precision three-dimensional medical image semantic segmentation system and method based on improved SAM model
CN118864855A