Image instance segmentation method, device, electronic device and storage medium

Through the Transformer network and preset query vector in the deep learning model, the problem of manual intervention and low segmentation accuracy in image instance segmentation is solved, and end-to-end efficient image instance segmentation is achieved.

CN114596436BActive Publication Date: 2025-05-06HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210136807.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-15
Publication Date
2025-05-06
Estimated Expiration
2042-02-15

AI Technical Summary

Technical Problem

In the prior art, in image instance segmentation, there are problems in which positive and negative samples need to be manually allocated, relying on non-maximum suppression algorithms for post-processing, detection results depend on segmentation results, and fixed-scale feature maps, resulting in poor segmentation accuracy of large-scale objects.

Method used

By obtaining the pending image and multiple preset query vectors, input image features and query vectors into the Transformer network of the deep learning model, generate instance features and dynamic parameters, and dynamically configure the Transformer network to realize image instance segmentation.

Benefits of technology

End-to-end image instance segmentation is realized, manual intervention is reduced, segmentation accuracy and efficiency is improved, and is suitable for mixed task scenarios of object detection and instance segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114596436B_ABST
    Figure CN114596436B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide an image instance segmentation method, an apparatus, an electronic device, and a storage medium to implement image instance segmentation. The query vector and the Transformer network can be used to differentially extract and aggregate the features of the image to be processed. The number of preset query vectors is greater than the number of instances in the image to be processed, so that each object instance in the image to be processed can obtain its unique features. Based on this feature, the category and segmentation mask of the object can be predicted by an augmented prediction network. The parameters in the dynamic Transformer are output by the prediction network, and the parameters are dynamically and adaptively configured, so that a segmentation mask can be generated for each object instance, thereby implementing end-to-end image instance segmentation, which can be directly applied to mixed task scenarios of target detection and instance segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to an image instance segmentation method, device, electronic device and storage medium. Background Art

[0002] With the development of computer technology, image instance segmentation based on computer vision technology has become possible. Especially after the emergence of deep learning models, computer vision technology based on deep learning models has made rapid progress.

[0003] Instances refer to objects that need to be detected in an image. For example, when the detection scene is to detect vehicles in an image, the instances are vehicles; for example, when the detection scene is to detect pedestrians and buildings, the instances are pedestrians and buildings. Image instance segmentation refers to the use of computer vision technology to divide the image area where the instance is located from the image. It provides the premise for target detection, event recognition, trajectory prediction, etc. based on the image. Image instance segmentation is increasingly used in production and life. How to achieve image instance segmentation has become a technical problem that needs to be solved urgently. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide an image instance segmentation method, device, electronic device and storage medium to achieve image instance segmentation. The specific technical solution is as follows:

[0005] In a first aspect, an embodiment of the present application provides an image instance segmentation method, comprising:

[0006] Acquire an image to be processed and a plurality of preset query vectors, wherein the number of the preset query vectors is greater than the number of instances in the image to be processed;

[0007] Inputting the image to be processed into a first neural network of a pre-trained deep learning model to obtain a first image feature of the image to be processed;

[0008] Inputting the first image feature and each of the preset query vectors into a first Transformer network of the deep learning model to obtain a second image feature and a plurality of instance features, wherein the number of the instance features is the same as the number of the preset query vectors;

[0009] Inputting the instance features into the prediction network of the deep learning model to obtain classification results and dynamic parameters of each instance feature;

[0010] Using each of the dynamic parameters respectively to configure parameters of the dynamic Transformer network of the deep learning model to obtain a plurality of second Transformer networks, wherein the number of the second Transformer networks is the same as the number of the preset query vectors;

[0011] Based on the second image features and each of the second Transformer networks, an instance segmentation result of the image to be processed is obtained.

[0012] In a possible implementation, the first Transformer network of the deep learning model includes an encoder and a decoder;

[0013] The step of inputting the first image feature and each of the preset query vectors into the first Transformer network of the deep learning model to obtain a second image feature and a plurality of instance features includes:

[0014] Encoding the first image feature using the encoder of the first Transformer network to obtain a second image feature;

[0015] The second image features and each of the preset query vectors are input into a decoder of the first Transformer network to obtain a plurality of instance features.

[0016] In a possible implementation, the encoder includes a plurality of encoder layers; and encoding the first image feature using the encoder of the first Transformer network to obtain the second image feature includes:

[0017] For each feature element in the first image feature, using the first encoder layer of the encoder to calculate a weight coefficient of each feature element in the first image feature for the feature element, and performing weighted summation on the feature element and other feature elements in the first image feature according to the weight coefficient to obtain an encoded feature element of the feature element, wherein the image feature output by the first encoder layer includes the encoded feature elements of each feature element in the first image feature;

[0018] For each feature element in the image feature output by the i-1th encoder, the i-th encoder layer of the encoder is used to calculate the weight coefficient of each feature element in the image feature output by the i-1th encoder for the feature element, and the feature element and other feature elements in the image feature output by the i-1th encoder are weighted and summed according to the weight coefficient to obtain the encoded feature element of the feature element; wherein i belongs to 2 to M, M is the number of encoder layers in the encoder, the image feature output by the i-th encoder layer includes the encoded feature elements of each feature element in the image feature output by the i-1th encoder, and the second image feature is the image feature output by the Mth encoder layer.

[0019] In a possible implementation, the decoder includes a plurality of decoder layers; the second image feature and each of the preset query vectors are input into the decoder of the first Transformer network to obtain a plurality of instance features, including:

[0020] For each of the preset query vectors, inputting the preset query vector and the second image feature into a first decoder layer of the decoder to obtain a fused query feature of the preset query vector output by the first decoder layer;

[0021] The fused query features of the preset query vector output by the i-1th decoder layer and the second image features are input into the i-th decoder layer to obtain the fused query features of the preset query vector output by the i-th decoder layer; wherein i belongs to 2 to N, N is the number of decoder layers in the decoder, and the instance features of the preset query vector are the fused query features of the preset query vector output by the N-th decoder layer.

[0022] In a possible implementation, the step of inputting the fused query feature of the preset query vector output by the i-1th decoder layer and the second image feature into the i-th decoder layer to obtain the fused query feature of the preset query vector output by the i-th decoder layer includes:

[0023] The fused query features of the preset query vector output by the i-1th decoder layer and the second image features are input into the i-th decoder layer, the i-th decoder layer is used to determine the weight coefficient of the input fused query features relative to each feature in the second image features, and the weight coefficient is used to perform weighted summation on each feature in the second image features to obtain the fused query features of the preset query vector output by the i-th decoder layer.

[0024] In a possible implementation, obtaining an instance segmentation result of the image to be processed based on the second image feature and each of the second Transformer networks includes:

[0025] Inputting the second image feature into a second neural network of the deep learning model to obtain a third image feature of the image to be processed, wherein the second neural network is used to perform feature conversion on the second image feature;

[0026] The third image features of the image to be processed are respectively input into each of the second Transformer networks to obtain an instance segmentation result of the image to be processed.

[0027] In one possible implementation, the process of pre-training a deep learning model includes:

[0028] Acquire a sample image and a plurality of sample query vectors, wherein the number of the sample query vectors is greater than the number of instances in the sample image;

[0029] Inputting the sample image and each of the sample query vectors into a deep learning model to obtain a plurality of predicted segmentation results of the sample image, wherein the number of the predicted segmentation results is the same as the number of the sample query vectors;

[0030] Determining a segmentation error of each instance in the sample image based on each predicted segmentation result of the sample image and each instance segmentation true value of the sample image;

[0031] Adjusting the parameters of the deep learning model and the value of the sample query vector according to each of the segmentation errors;

[0032] Other sample images are selected to continue training the deep learning model until a preset training end condition is met, thereby obtaining a pre-trained deep learning model and the plurality of preset query vectors.

[0033] In a possible implementation, determining the segmentation error of each instance in the sample image based on each predicted segmentation result of the sample image and each instance segmentation true value of the sample image includes:

[0034] For each predicted segmentation result of the sample image, determining a multidimensional vector of the predicted segmentation result;

[0035] For each instance segmentation true value of the sample image, determining a multidimensional vector of the instance segmentation true value;

[0036] For each instance segmentation true value of the sample image, the multidimensional vector of the instance segmentation true value is matched with the multidimensional vector of each predicted segmentation result to obtain a predicted segmentation result that matches the instance segmentation true value;

[0037] For each instance segmentation true value of the sample image, the error between the multidimensional vector of the instance segmentation true value and the multidimensional vector of the predicted segmentation result matching the instance segmentation true value is calculated to obtain the segmentation error of the instance corresponding to the instance segmentation true value.

[0038] In a possible implementation manner, determining a multidimensional vector of each predicted segmentation result of the sample image includes:

[0039] For each predicted segmentation result of the sample image, determining a key point of an instance in the predicted segmentation result;

[0040] The vectors from each pixel point to the key points of the instance in the predicted segmentation result are calculated respectively to obtain a multi-dimensional vector of the predicted segmentation result.

[0041] In a second aspect, an embodiment of the present application provides an image instance segmentation device, comprising:

[0042] A module for acquiring an image to be processed, used for acquiring an image to be processed and a plurality of preset query vectors, wherein the number of the preset query vectors is greater than the number of instances in the image to be processed;

[0043] A first image feature extraction module, used for inputting the image to be processed into a first neural network of a pre-trained deep learning model to obtain a first image feature of the image to be processed;

[0044] an instance feature extraction module, configured to input the first image feature and each of the preset query vectors into a first Transformer network of the deep learning model to obtain a second image feature and a plurality of instance features, wherein the number of the instance features is the same as the number of the preset query vectors;

[0045] A dynamic parameter determination module, used to input the instance features into the prediction network of the deep learning model to obtain the classification results and dynamic parameters of each instance feature;

[0046] A dynamic parameter configuration module, used to configure parameters of the dynamic Transformer network of the deep learning model using each of the dynamic parameters to obtain a plurality of second Transformer networks, wherein the number of the second Transformer networks is the same as the number of the preset query vectors;

[0047] An instance segmentation result determination module is used to obtain an instance segmentation result of the image to be processed based on the second image features and each of the second Transformer networks.

[0048] In a possible implementation, the first Transformer network of the deep learning model includes an encoder and a decoder;

[0049] The instance feature extraction module comprises:

[0050] A second image feature extraction submodule, configured to encode the first image feature using the encoder of the first Transformer network to obtain a second image feature;

[0051] The instance feature extraction submodule is used to input the second image feature and each of the preset query vectors into the decoder of the first Transformer network to obtain multiple instance features.

[0052] In one possible implementation, the encoder includes a plurality of encoder layers;

[0053] The second image feature extraction submodule is specifically used for:

[0054] For each feature element in the first image feature, using the first encoder layer of the encoder to calculate a weight coefficient of each feature element in the first image feature for the feature element, and performing weighted summation on the feature element and other feature elements in the first image feature according to the weight coefficient to obtain an encoded feature element of the feature element, wherein the image feature output by the first encoder layer includes the encoded feature elements of each feature element in the first image feature;

[0055] For each feature element in the image feature output by the i-1th encoder, the i-th encoder layer of the encoder is used to calculate the weight coefficient of each feature element in the image feature output by the i-1th encoder for the feature element, and the feature element and other feature elements in the image feature output by the i-1th encoder are weighted and summed according to the weight coefficient to obtain the encoded feature element of the feature element; wherein i belongs to 2 to M, M is the number of encoder layers in the encoder, the image feature output by the i-th encoder layer includes the encoded feature elements of each feature element in the image feature output by the i-1th encoder, and the second image feature is the image feature output by the Mth encoder layer.

[0056] In a possible implementation, the decoder includes a plurality of decoder layers; the instance feature extraction submodule is specifically configured to:

[0057] For each of the preset query vectors, inputting the preset query vector and the second image feature into a first decoder layer of the decoder to obtain a fused query feature of the preset query vector output by the first decoder layer;

[0058] The fused query features of the preset query vector output by the i-1th decoder layer and the second image features are input into the i-th decoder layer to obtain the fused query features of the preset query vector output by the i-th decoder layer; wherein i belongs to 2 to N, N is the number of decoder layers in the decoder, and the instance features of the preset query vector are the fused query features of the preset query vector output by the N-th decoder layer.

[0059] In a possible implementation manner, the instance feature extraction submodule is specifically used to:

[0060] The fused query features of the preset query vector output by the i-1th decoder layer and the second image features are input into the i-th decoder layer, the i-th decoder layer is used to determine the weight coefficient of the input fused query features relative to each feature in the second image features, and the weight coefficient is used to perform weighted summation on each feature in the second image features to obtain the fused query features of the preset query vector output by the i-th decoder layer.

[0061] In a possible implementation, the instance segmentation result determination module is specifically configured to:

[0062] Inputting the second image feature into a second neural network of the deep learning model to obtain a third image feature of the image to be processed, wherein the second neural network is used to perform feature conversion on the second image feature;

[0063] The third image features of the image to be processed are respectively input into each of the second Transformer networks to obtain an instance segmentation result of the image to be processed.

[0064] In a possible implementation, the device further includes:

[0065] A sample data acquisition module, used to acquire a sample image and a plurality of sample query vectors, wherein the number of the sample query vectors is greater than the number of instances in the sample image;

[0066] a segmentation result prediction module, used for inputting the sample image and each of the sample query vectors into a deep learning model to obtain a plurality of predicted segmentation results of the sample image, wherein the number of the predicted segmentation results is the same as the number of the sample query vectors;

[0067] A segmentation error determination module, configured to determine a segmentation error of each instance in the sample image based on each predicted segmentation result of the sample image and each instance segmentation true value of the sample image;

[0068] A query vector adjustment module, used to adjust the parameters of the deep learning model and the value of the sample query vector according to each of the segmentation errors;

[0069] The training end determination module is used to select other sample images to continue training the deep learning model until the preset training end conditions are met, thereby obtaining the pre-trained deep learning model and the multiple preset query vectors.

[0070] In a possible implementation manner, the segmentation error determination module includes:

[0071] A first multi-dimensional vector determination submodule, configured to determine, for each predicted segmentation result of the sample image, a multi-dimensional vector of the predicted segmentation result;

[0072] A second multi-dimensional vector determination submodule, configured to determine, for each instance segmentation true value of the sample image, a multi-dimensional vector of the instance segmentation true value;

[0073] A multi-dimensional vector matching submodule is used to match the multi-dimensional vector of each instance segmentation true value of the sample image with the multi-dimensional vector of each predicted segmentation result, so as to obtain a predicted segmentation result matching the instance segmentation true value;

[0074] The segmentation error calculation submodule is used to calculate, for each instance segmentation true value of the sample image, the error between the multidimensional vector of the instance segmentation true value and the multidimensional vector of the predicted segmentation result matching the instance segmentation true value, to obtain the segmentation error of the instance corresponding to the instance segmentation true value.

[0075] In a possible implementation manner, the first multidimensional vector determination submodule is specifically configured to:

[0076] For each predicted segmentation result of the sample image, determining a key point of an instance in the predicted segmentation result;

[0077] The vectors from each pixel point to the key points of the instance in the predicted segmentation result are calculated respectively to obtain a multi-dimensional vector of the predicted segmentation result.

[0078] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory;

[0079] The memory is used to store computer programs;

[0080] The processor is used to implement any image instance segmentation method described in the present application when executing the program stored in the memory.

[0081] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the image instance segmentation method described in any one of the present application is implemented.

[0082] In a fifth aspect, an embodiment of the present application provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute any of the image instance segmentation methods described in the present application.

[0083] Beneficial effects of the embodiments of the present application:

[0084] The image instance segmentation method, device, electronic device and storage medium provided in the embodiments of the present application can realize image instance segmentation, and can use query vectors and Transformer networks to extract and aggregate the features of the image to be processed in a differentiated manner. The number of preset query vectors is greater than the number of instances in the image to be processed, so that each object instance in the image to be processed can obtain its unique features; based on this feature, the category and segmentation mask of the object can be predicted by adding a prediction network, and the parameters in the dynamic Transformer are output by the prediction network, and the parameters are dynamically and adaptively configured, so that for each object instance, its segmentation mask can be generated accordingly, realizing end-to-end image instance segmentation, which can be directly applied to the mixed task scenario of target detection and instance segmentation. Of course, it is not necessary to achieve all the advantages described above at the same time to implement any product or method of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other embodiments can also be obtained based on these drawings.

[0086] Figure 1 A schematic diagram of a flow chart of an image instance segmentation method according to an embodiment of the present application;

[0087] Figure 2 A schematic diagram of a possible implementation of step S13 in an embodiment of the present application;

[0088] Figure 3 A schematic diagram of an encoder according to an embodiment of the present application;

[0089] Figure 4 A schematic diagram of a decoder according to an embodiment of the present application;

[0090] Figure 5A schematic diagram of a possible implementation of step S16 in an embodiment of the present application;

[0091] Figure 6 A schematic diagram of a process flow of a deep learning model training method according to an embodiment of the present application;

[0092] Figure 7 A schematic diagram of an image instance segmentation device according to an embodiment of the present application;

[0093] Figure 8 A schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0094] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field based on the present application belong to the scope of protection of the present application.

[0095] First, the terms in this application are explained:

[0096] Transformer network: a network structure based on the attention mechanism, which is different from the convolutional neural network structure, but is also a module for building neural networks. Generally speaking, a complete Transformer network consists of an encoder and a decoder. In special cases, there are also Transformer networks with only encoder or decoder structures.

[0097] Query vector: Query vector, generally used to represent the query on the result of the task.

[0098] End-to-end model: generally refers to a type of model that does not require additional post-processing and the network output can be directly used as the final result.

[0099] FPN (Feature Pyramid Networks): A feature fusion structure that integrates multi-level features.

[0100] NMS (non-maximum suppression) algorithm: a post-processing method for detection results that removes duplicate detection results based on the score of the detection box.

[0101] RoI (region of interest): In some detection and segmentation methods, the corresponding area of ​​the object is deducted, which is called RoI feature.

[0102] Image instance segmentation is increasingly being used in production and life. In the related image instance segmentation technologies: first, the image to be processed is input into a convolutional backbone network, such as ResNet (Residual Neural Network) and FPN network, to obtain features of multiple scales of the image to be processed; then, based on the multi-scale features, the detection frame of the instance in the image to be processed is densely predicted, and the detection frame of the instance in the image to be processed is obtained through the non-maximum suppression algorithm; according to the detection frame of the instance in the image to be processed, the features of the region of interest are extracted on the multi-scale features, and the extracted regional features are input into the segmentation branch network to obtain the segmentation mask of the corresponding object; the size of the segmentation mask obtained here is fixed, and the result is upsampled to the size of the detection frame of the object, and correspondingly pasted to the position of the object in the original image to be processed.

[0103] However, the above method has the following problems:

[0104] 1. During training, it is necessary to manually formulate the positive and negative sample allocation strategy.

[0105] 2. It is necessary to first obtain the detection result through the non-maximum suppression algorithm. The non-maximum suppression algorithm is an additional post-processing process and requires manual setting. The overall model is not end-to-end.

[0106] 3. First get the detection result, and then generate a segmentation mask within the detection box of the detection result. The segmentation result depends on the detection result.

[0107] 4. Generate masks on fixed-scale feature maps, resulting in poor segmentation accuracy for large-scale objects.

[0108] In view of this, embodiments of the present application provide an image instance segmentation method, an apparatus, an electronic device, and a storage medium to solve at least one of the above problems.

[0109] See also Figure 1 , Figure 1 : is a flow chart of an image instance segmentation method according to an embodiment of the present application, comprising:

[0110] S11, obtaining an image to be processed and a plurality of preset query vectors, wherein the number of the preset query vectors is greater than the number of instances in the image to be processed.

[0111] The image instance segmentation method of the embodiment of the present application can be implemented by an electronic device. Specifically, the electronic device can be a smart phone, smart glasses, a smart camera, a hard disk recorder or a personal computer.

[0112] The image to be processed is any image that needs to be segmented for instance, and the preset query vector is the query vector obtained during the deep learning model training process, which is the feature vector related to the instance learned by the deep learning module. The training process of the deep learning model will be described in the following embodiments.

[0113] S12: input the image to be processed into a first neural network of a pre-trained deep learning model to obtain a first image feature of the image to be processed.

[0114] The first neural network is a backbone network in the deep learning model, and is used to extract image features of the image to be processed, thereby obtaining the first image features of the image to be processed. The first neural network can be any type of feature extraction network, for example, ResNet, AlexNet, VGG (Visual Graphics Generator, target image generator) or Overfeat, etc., all within the scope of protection of this application.

[0115] S13, inputting the first image feature and each of the preset query vectors into the first Transformer network of the deep learning model to obtain a second image feature and multiple instance features, wherein the number of the instance features is the same as the number of the preset query vectors.

[0116] The first Transformer network of the deep learning model is used to perform feature transformation on the first image feature to obtain the second image feature of the image to be processed; the second image feature is respectively fused with each preset query vector using the first Transformer network to obtain the instance feature corresponding to each preset query vector.

[0117] S14, inputting the instance features into the prediction network of the deep learning model to obtain the classification results and dynamic parameters of each instance feature.

[0118] The specific structure of the prediction network can refer to the prediction network structure in the related art. In one example, the prediction network can be a stacked fully connected layer network. The prediction network of the deep learning model is used to analyze each instance feature separately to obtain the classification result and dynamic parameters of each instance feature. Among them, for any instance feature, the dynamic parameters of the instance feature are used to configure the parameters of the dynamic Transformer network to obtain the second Transformer network corresponding to the instance feature; the classification result of the instance feature represents: the type of instance segmentation result obtained by the second Transformer network corresponding to the instance feature, for example, pedestrian type or vehicle type, etc.

[0119] S15, respectively configuring parameters of the dynamic Transformer network of the deep learning model using each of the dynamic parameters to obtain a plurality of second Transformer networks, wherein the number of the second Transformer networks is the same as the number of the preset query vectors.

[0120] For each instance feature's dynamic parameters, the dynamic Transformer network is parameterized using the instance feature's dynamic parameters to obtain a second Transformer network corresponding to the instance feature, and finally multiple second Transformer networks are obtained.

[0121] S16: Based on the second image features and each of the second Transformer networks, obtain an instance segmentation result of the image to be processed.

[0122] In one example, the second Transformer network may include only an encoder. The segmentation mask of each instance is determined by the second Transformer network to obtain the instance segmentation result of the instance in the image to be processed. In one example, based on the second image feature, each second Transformer network will output an instance segmentation result and a confidence level. For instance segmentation results with a confidence level lower than a preset confidence threshold, they are judged as invalid instance segmentation results; for instance segmentation results with a confidence level not lower than a preset confidence threshold, they are judged as valid instance segmentation results; and each valid instance segmentation result is used as the instance segmentation result of the image to be processed.

[0123] The number of instance segmentation results finally obtained is the same as the number of instances in the image to be processed, which is less than the number of preset query vectors. Assuming that there are N preset query vectors and M instances in the image to be processed, N>M, N instance features will be obtained after the first Transformer network, where the preset query vectors correspond one-to-one to the instance features; and the number of second Transformer networks is also N, each second Transformer network can obtain an instance segmentation result and its confidence, a total of N instance segmentation results and N confidences are obtained, each instance segmentation result corresponds to a confidence, among which the confidence of M instance segmentation results is greater than the preset confidence threshold, which is considered to be a credible instance segmentation result, as the final instance segmentation result; and the confidence of the other (NM) instance segmentation results is less than or equal to the preset confidence threshold, which is considered to be an unreliable instance segmentation result and needs to be filtered out.

[0124] In an embodiment of the present application, query vectors and Transformer networks are used to differentially extract and aggregate features of the image to be processed. The number of preset query vectors is greater than the number of instances in the image to be processed, so that each object instance in the image to be processed can obtain its unique features; based on this feature, the category and segmentation mask of the object can be predicted by an augmented prediction network, and the parameters in the dynamic Transformer are output by the prediction network, and the parameters are dynamically and adaptively configured, so that a segmentation mask can be generated for each object instance, thereby realizing end-to-end image instance segmentation, which can be directly applied to hybrid task scenarios of target detection and instance segmentation.

[0125] In one possible implementation, the first Transformer network of the deep learning model includes an encoder and a decoder; see Figure 2 , the first image feature and each of the preset query vectors are input into the first Transformer network of the deep learning model to obtain a second image feature and a plurality of instance features, including:

[0126] S131, using the encoder of the first Transformer network to encode the first image feature to obtain a second image feature.

[0127] The first Transformer network includes an encoder and a decoder. In one example, the structure of the encoder can be as follows Figure 3 As shown, the encoder is composed of multiple encoder layers, each of which fuses input features through an attention mechanism to generate new output features. In one possible implementation, the encoder of the first Transformer network is used to encode the first image feature to obtain the second image feature, including:

[0128] Step 1: For each feature element in the first image feature, use the first encoder layer of the encoder to calculate the weight coefficient of each feature element in the first image feature for the feature element, and perform weighted summation on the feature element and other feature elements in the first image feature according to the weight coefficient to obtain the encoded feature element of the feature element, wherein the image feature output by the first encoder layer includes the encoded feature elements of each feature element in the first image feature.

[0129] Step 2: For each feature element in the image feature output by the i-1th encoder, use the i-th encoder layer of the encoder to calculate the weight coefficient of each feature element in the image feature output by the i-1th encoder for the feature element, and perform weighted summation on the feature element and other feature elements in the image feature output by the i-1th encoder according to the weight coefficient to obtain the encoded feature element of the feature element; wherein i belongs to 2 to M, M is the number of encoder layers in the encoder, the image feature output by the i-th encoder layer includes the encoded feature elements of each feature element in the image feature output by the i-1th encoder, and the second image feature is the image feature output by the Mth encoder layer.

[0130] For example, suppose N features are input. For a specific input feature, the relationship between it and all other features (including itself) is calculated in the encoder layer, and N weight coefficients are calculated based on this relationship. Based on this weight coefficient, the weighted sum of the N features is calculated to obtain the new feature after the fusion of this feature. In each second image feature, the same operation is performed on each input feature, and the shape of the final output feature remains consistent with the input feature. For example, the first image feature is represented as a 3D H*W*C, where H and W are the width and height of the first image feature, and C is the feature channel. The output feature value is still H*W*C. In one example, the input feature can be stretched to N*C, where N=H*W.

[0131] S132: Input the second image features and each of the preset query vectors into the decoder of the first Transformer network to obtain a plurality of instance features.

[0132] The preset query vector is extracted and fused with the information with high correlation in the second image feature through the decoder, so as to obtain the feature expression corresponding to each instance, that is, the instance feature. In an example, the structure of the decoder can be as follows Figure 4 As shown, a decoder is usually composed of multiple decoder layers. Each decoder layer has two inputs: query features and image features. In a possible implementation, the decoder includes multiple decoder layers; the second image features and each of the preset query vectors are input into the decoder of the first Transformer network to obtain multiple instance features, including:

[0133] Step 1: for each of the preset query vectors, the preset query vector and the second image feature are input into the first decoder layer of the decoder to obtain a fused query feature of the preset query vector output by the first decoder layer.

[0134] Step 2: Input the fused query features corresponding to the preset query vector output by the i-1th decoder layer and the second image features into the i-th decoder layer to obtain the fused query features of the preset query vector output by the i-th decoder layer; wherein i belongs to 2 to N, N is the number of decoder layers in the decoder, and the instance features of the preset query vector are the fused query features of the preset query vector output by the N-th decoder layer.

[0135] Assume that the number of input query features (specifically, the input of the first decoder layer is the preset query vector, and the input of the i-th decoder layer is the fused query feature output by the i-1-th decoder layer) is Q, and the number of image features in the second image feature is N. In a single decoder layer, for each input query feature, its relationship with the N image features is calculated, and N weight coefficients are calculated based on this relationship. Then, the weighted sum of the N image features is calculated based on the weight coefficients to obtain the fused and updated fused query feature. In a possible implementation, the step of inputting the fused query features of the preset query vector output by the i-1th decoder layer and the second image features into the i-th decoder layer to obtain the fused query features of the preset query vector output by the i-th decoder layer includes: inputting the fused query features of the preset query vector output by the i-1th decoder layer and the second image features into the i-th decoder layer, using the i-th decoder layer to determine the weight coefficient of the input fused query features relative to each feature in the second image features, and using the weight coefficient to perform weighted summation on each feature in the second image features to obtain the fused query features of the preset query vector output by the i-th decoder layer.

[0136] Each decoder layer still outputs Q fused query features. In the stack of decoder layers, the fused query features are continuously updated, while the image features remain unchanged. The Q fused query features output by the final decoder are considered to be all instance features in the image to be analyzed.

[0137] In order to facilitate the second Transformer network to obtain the segmentation mask of each instance, a neural network can be added based on the second image feature output by the first Transformer network. Figure 5 , obtaining the instance segmentation result of the image to be processed based on the second image features and each of the second Transformer networks, includes:

[0138] S161, inputting the second image feature into the second neural network of the deep learning model to obtain a third image feature of the image to be processed, wherein the second neural network is used to perform feature conversion on the second image feature.

[0139] S162, inputting the third image features of the image to be processed into each of the second Transformer networks respectively to obtain an instance segmentation result of the image to be processed.

[0140] The second neural network is used to perform feature transformation on the second image feature, so as to obtain a third image feature suitable for the second Transformer network. The network type of the second neural network can be customized according to actual conditions, for example, it can be ResNet, AlexNet, VGG (Visual Graphics Generator, target image generator) or Overfeat, etc., all within the scope of protection of this application.

[0141] The present application also provides a method for training a deep learning model. Figure 6 ,include:

[0142] S21, obtaining a sample image and a plurality of sample query vectors, wherein the number of the sample query vectors is greater than the number of instances in the sample image.

[0143] The sample images used in the current training are selected from the sample image set. The sample query vector can be randomly generated, and the number of sample query vectors must be greater than the number of instances in the sample image.

[0144] S22, inputting the sample image and each of the sample query vectors into a deep learning model to obtain a plurality of predicted segmentation results of the sample image, wherein the number of the predicted segmentation results is the same as the number of the sample query vectors.

[0145] The structure of the deep learning model can refer to the structure of the deep learning model in the above embodiment; the process of obtaining the predicted segmentation results of the sample image can refer to the process of obtaining the instance segmentation results of the image to be processed in the above embodiment, wherein the multiple predicted segmentation results of the sample image include the predicted segmentation results output by each second Transformer network, and are not limited to the predicted segmentation results whose confidence is not lower than the preset confidence threshold.

[0146] S23, determining the segmentation error of each instance in the sample image based on each predicted segmentation result of the sample image and each instance segmentation true value of the sample image.

[0147] The segmentation error of each instance in the sample image can be calculated according to each predicted segmentation result of the sample image and each instance segmentation true value through a preset loss function. The preset loss function can be customized according to actual conditions.

[0148] In a possible implementation, determining the segmentation error of each instance in the sample image based on each predicted segmentation result of the sample image and each instance segmentation true value of the sample image includes:

[0149] Step 1, for each predicted segmentation result of the sample image, determining a multidimensional vector of the predicted segmentation result;

[0150] In a possible implementation, for each predicted segmentation result of the sample image, determining the multidimensional vector of the predicted segmentation result includes: for each predicted segmentation result of the sample image, determining the key points of the instance in the predicted segmentation result; and respectively calculating the vectors from each pixel point to the key points of the instance in the predicted segmentation result to obtain the multidimensional vector of the predicted segmentation result.

[0151] In an example, the predicted segmentation result of the sample image can be represented as a two-dimensional binary image. Each pixel on the binary image corresponds to a coordinate position (hereinafter referred to as pos). For example, the coordinate of the upper left corner pixel is (0,0), and the coordinate of the lower right corner pixel is (H-1,W-1). Among them, H and W are the height and width of the feature map respectively. Each instance itself also has a position information (hereinafter referred to as pos q ), can be represented by the key points of the instance. Specifically, the key points of the instance can be the center point of the bounding box of the instance or the centroid of the instance. Using the location information of the instance, a unique location encoding map is calculated for each instance, thereby providing the prior information of the location of each instance segmentation mask. For example, for the feature map of H*W, a one-dimensional vector with a length of d is calculated for each point at each position in the feature map. The calculation formula of the one-dimensional vector at each point is as follows:

[0152]

[0153]

[0154] Wherein, i represents the number of bits in the one-dimensional vector, and d is the preset position coding feature length.

[0155] Step 2: for each instance segmentation true value of the sample image, determine a multi-dimensional vector of the instance segmentation true value.

[0156] The instance segmentation true value may be manually calibrated. The method for calculating the multidimensional vector of the instance segmentation true value may refer to the method for calculating the multidimensional vector of the predicted segmentation result, which will not be described in detail here.

[0157] Step three, for each instance segmentation true value of the sample image, the multidimensional vector of the instance segmentation true value is matched with the multidimensional vector of each predicted segmentation result to obtain the predicted segmentation result that matches the instance segmentation true value.

[0158] Step 4: For each instance segmentation true value of the sample image, calculate the error between the multidimensional vector of the instance segmentation true value and the multidimensional vector of the predicted segmentation result matching the instance segmentation true value to obtain the segmentation error of the instance corresponding to the instance segmentation true value.

[0159] Based on the multidimensional vector of each predicted segmentation result and the multidimensional vector of each instance segmentation true value, the matching degree between the instance segmentation true value and the predicted segmentation result is calculated. The matching degree calculation method can be customized according to the actual situation. For example, the classification loss plus the segmentation loss is used as the matching degree between the instance segmentation true value and the predicted segmentation result. By calculating the loss between all instance segmentation true values ​​and the predicted segmentation results, the matching matrix between the instance segmentation true value and the predicted segmentation result can be obtained. Assuming that there are Q predicted segmentation results and M instance segmentation true values, the size of the matching matrix is ​​Q*M. By solving the matching matrix globally (such as using the Hungarian algorithm), the true value that best matches each predicted segmentation result can be found on the basis of global optimality. This true value is the training true value corresponding to this predicted segmentation result during the training process. Because one true value corresponds to one predicted segmentation result, and the number of predicted segmentation results must be greater than the number of instance objects in the image, there will be a certain number of predicted segmentation results that do not match the true value, and these predicted segmentation results are considered to be negative samples.

[0160] S24, adjusting the parameters of the deep learning model and the value of the sample query vector according to each of the segmentation errors.

[0161] When the deep learning model starts training, the sample query vector can be randomly generated. During the training process of the deep learning model, the value of the sample query vector will be adjusted according to the segmentation error. After the training is completed, the value of the sample query vector is fixed, that is, the preset query vector in the above embodiment.

[0162] S25, selecting other sample images to continue training the deep learning model until a preset training end condition is met, thereby obtaining a pre-trained deep learning model and the plurality of preset query vectors.

[0163] The training end conditions can be customized according to the actual situation, for example, the loss of the deep learning model can converge, or the preset number of training times can be reached.

[0164] In the embodiment of the present application, the allocation of positive and negative samples is performed in this adaptive manner, which effectively reduces the involvement of artificial experience in model training, further improves the learning ability and performance of the network, and does not require manually set RoI operations and NMS operations, making deep learning model training simpler and more efficient.

[0165] The present application also provides an image instance segmentation device. Figure 7 ,include:

[0166] The to-be-processed image acquisition module 101 is used to acquire the to-be-processed image and a plurality of preset query vectors, wherein the number of the preset query vectors is greater than the number of instances in the to-be-processed image;

[0167] A first image feature extraction module 102, used for inputting the image to be processed into a first neural network of a pre-trained deep learning model to obtain a first image feature of the image to be processed;

[0168] An instance feature extraction module 103 is used to input the first image feature and each of the preset query vectors into a first Transformer network of the deep learning model to obtain a second image feature and a plurality of instance features, wherein the number of the instance features is the same as the number of the preset query vectors;

[0169] A dynamic parameter determination module 104 is used to input each of the instance features into the prediction network of the deep learning model to obtain the classification results and dynamic parameters of each of the instance features;

[0170] A dynamic parameter configuration module 105, configured to configure parameters of the dynamic Transformer network of the deep learning model using the dynamic parameters to obtain a plurality of second Transformer networks, wherein the number of the second Transformer networks is the same as the number of the preset query vectors;

[0171] The instance segmentation result determination module 106 is used to obtain the instance segmentation result of the image to be processed based on the second image features and each of the second Transformer networks.

[0172] In a possible implementation, the first Transformer network of the deep learning model includes an encoder and a decoder;

[0173] The instance feature extraction module comprises:

[0174] A second image feature extraction submodule, configured to encode the first image feature using the encoder of the first Transformer network to obtain a second image feature;

[0175] The instance feature extraction submodule is used to input the second image feature and each of the preset query vectors into the decoder of the first Transformer network to obtain multiple instance features.

[0176] In one possible implementation, the encoder includes a plurality of encoder layers;

[0177] The second image feature extraction submodule is specifically used for:

[0178] For each feature element in the first image feature, using the first encoder layer of the encoder to calculate a weight coefficient of each feature element in the first image feature for the feature element, and performing weighted summation on the feature element and other feature elements in the first image feature according to the weight coefficient to obtain an encoded feature element of the feature element, wherein the image feature output by the first encoder layer includes the encoded feature elements of each feature element in the first image feature;

[0179] For each feature element in the image feature output by the i-1th encoder, the i-th encoder layer of the encoder is used to calculate the weight coefficient of each feature element in the image feature output by the i-1th encoder for the feature element, and the feature element and other feature elements in the image feature output by the i-1th encoder are weighted and summed according to the weight coefficient to obtain the encoded feature element of the feature element; wherein i belongs to 2 to M, M is the number of encoder layers in the encoder, the image feature output by the i-th encoder layer includes the encoded feature elements of each feature element in the image feature output by the i-1th encoder, and the second image feature is the image feature output by the Mth encoder layer.

[0180] In a possible implementation, the decoder includes a plurality of decoder layers; the instance feature extraction submodule is specifically configured to:

[0181] For each of the preset query vectors, inputting the preset query vector and the second image feature into a first decoder layer of the decoder to obtain a fused query feature of the preset query vector output by the first decoder layer;

[0182] The fused query features of the preset query vector output by the i-1th decoder layer and the second image features are input into the i-th decoder layer to obtain the fused query features of the preset query vector output by the i-th decoder layer; wherein i belongs to 2 to N, N is the number of decoder layers in the decoder, and the instance features of the preset query vector are the fused query features of the preset query vector output by the N-th decoder layer.

[0183] In a possible implementation manner, the instance feature extraction submodule is specifically used to:

[0184] The fused query features of the preset query vector output by the i-1th decoder layer and the second image features are input into the i-th decoder layer, the i-th decoder layer is used to determine the weight coefficient of the input fused query features relative to each feature in the second image features, and the weight coefficient is used to perform weighted summation on each feature in the second image features to obtain the fused query features of the preset query vector output by the i-th decoder layer.

[0185] In a possible implementation, the instance segmentation result determination module is specifically configured to:

[0186] Inputting the second image feature into a second neural network of the deep learning model to obtain a third image feature of the image to be processed, wherein the second neural network is used to perform feature conversion on the second image feature;

[0187] The third image features of the image to be processed are respectively input into each of the second Transformer networks to obtain an instance segmentation result of the image to be processed.

[0188] In a possible implementation, the device further includes:

[0189] A sample data acquisition module, used to acquire a sample image and a plurality of sample query vectors, wherein the number of the sample query vectors is greater than the number of instances in the sample image;

[0190] a segmentation result prediction module, used for inputting the sample image and each of the sample query vectors into a deep learning model to obtain a plurality of predicted segmentation results of the sample image, wherein the number of the predicted segmentation results is the same as the number of the sample query vectors;

[0191] A segmentation error determination module, configured to determine a segmentation error of each instance in the sample image based on each predicted segmentation result of the sample image and each instance segmentation true value of the sample image;

[0192] A query vector adjustment module, used to adjust the parameters of the deep learning model and the value of the sample query vector according to each of the segmentation errors;

[0193] The training end determination module is used to select other sample images to continue training the deep learning model until the preset training end conditions are met, thereby obtaining the pre-trained deep learning model and the multiple preset query vectors.

[0194] In a possible implementation manner, the segmentation error determination module includes:

[0195] A first multi-dimensional vector determination submodule, configured to determine, for each predicted segmentation result of the sample image, a multi-dimensional vector of the predicted segmentation result;

[0196] A second multi-dimensional vector determination submodule, configured to determine, for each instance segmentation true value of the sample image, a multi-dimensional vector of the instance segmentation true value;

[0197] A multi-dimensional vector matching submodule is used to match the multi-dimensional vector of each instance segmentation true value of the sample image with the multi-dimensional vector of each predicted segmentation result, so as to obtain a predicted segmentation result matching the instance segmentation true value;

[0198] The segmentation error calculation submodule is used to calculate, for each instance segmentation true value of the sample image, the error between the multidimensional vector of the instance segmentation true value and the multidimensional vector of the predicted segmentation result matching the instance segmentation true value, to obtain the segmentation error of the instance corresponding to the instance segmentation true value.

[0199] In a possible implementation manner, the first multidimensional vector determination submodule is specifically configured to:

[0200] For each predicted segmentation result of the sample image, determining a key point of an instance in the predicted segmentation result;

[0201] The vectors from each pixel point to the key points of the instance in the predicted segmentation result are calculated respectively to obtain a multi-dimensional vector of the predicted segmentation result.

[0202] The embodiment of the present application also provides an electronic device, and the embodiment of the present application provides an electronic device, including a processor and a memory;

[0203] The memory is used to store computer programs;

[0204] The processor is used to implement any image instance segmentation method described in the present application when executing the program stored in the memory.

[0205] In a possible implementation, Figure 8 As shown, the electronic device of the embodiment of the present application further includes a communication interface 202 and a communication bus 204 , wherein the processor 201 , the communication interface 202 , and the memory 203 communicate with each other via the communication bus 204 .

[0206] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0207] The communication interface is used for communication between the above electronic device and other devices.

[0208] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0209] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0210] In another embodiment provided in the present application, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, any method described in the present application is implemented.

[0211] In another embodiment provided in the present application, a computer program product including instructions is also provided, which, when executed on a computer, enables the computer to execute any of the methods described in the present application.

[0212] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive Solid State Disk (SSD)), etc.

[0213] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0214] Each embodiment in this specification is described in a related manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referenced to each other.

[0215] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are included in the protection scope of the present application.

Claims

1. A method for image instance segmentation, characterized in that: include: Acquire an image to be processed and a plurality of preset query vectors, wherein the number of the preset query vectors is greater than the number of instances in the image to be processed; Inputting the image to be processed into a first neural network of a pre-trained deep learning model to obtain a first image feature of the image to be processed; Inputting the first image feature and each of the preset query vectors into a first Transformer network of the deep learning model to obtain a second image feature and a plurality of instance features, wherein the number of the instance features is the same as the number of the preset query vectors; Inputting the instance features into the prediction network of the deep learning model to obtain classification results and dynamic parameters of each instance feature; Using each of the dynamic parameters respectively to configure parameters of the dynamic Transformer network of the deep learning model to obtain a plurality of second Transformer networks, wherein the number of the second Transformer networks is the same as the number of the preset query vectors; Based on the second image features and each of the second Transformer networks, obtaining an instance segmentation result of the image to be processed; Wherein, the first Transformer network of the deep learning model includes an encoder and a decoder; The step of inputting the first image feature and each of the preset query vectors into the first Transformer network of the deep learning model to obtain a second image feature and a plurality of instance features includes: Encoding the first image feature using the encoder of the first Transformer network to obtain a second image feature; The second image features and each of the preset query vectors are input into a decoder of the first Transformer network to obtain a plurality of instance features.

2. The method according to claim 1, characterized in that The encoder includes a plurality of encoder layers; the encoder of the first Transformer network is used to encode the first image feature to obtain a second image feature, including: For each feature element in the first image feature, using the first encoder layer of the encoder to calculate a weight coefficient of each feature element in the first image feature for the feature element, and performing weighted summation on the feature element and other feature elements in the first image feature according to the weight coefficient to obtain an encoded feature element of the feature element, wherein the image feature output by the first encoder layer includes the encoded feature elements of each feature element in the first image feature; For each feature element in the image feature output by the i-1th encoder, the i-th encoder layer of the encoder is used to calculate the weight coefficient of each feature element in the image feature output by the i-1th encoder for the feature element, and the feature element and other feature elements in the image feature output by the i-1th encoder are weighted and summed according to the weight coefficient to obtain the encoded feature element of the feature element; wherein i belongs to 2 to M, M is the number of encoder layers in the encoder, the image feature output by the i-th encoder layer includes the encoded feature elements of each feature element in the image feature output by the i-1th encoder, and the second image feature is the image feature output by the Mth encoder layer.

3. The method according to claim 2, characterized in that The decoder includes a plurality of decoder layers; the second image feature and each of the preset query vectors are input into the decoder of the first Transformer network to obtain a plurality of instance features, including: For each of the preset query vectors, inputting the preset query vector and the second image feature into a first decoder layer of the decoder to obtain a fused query feature of the preset query vector output by the first decoder layer; The fused query feature corresponding to the preset query vector output by the i-1th decoder layer and the second image feature are input into the i-th decoder layer to obtain the fused query feature of the preset query vector output by the i-th decoder layer; wherein i belongs to 2 to N, N is the number of decoder layers in the decoder, and the instance feature of the preset query vector is the fused query feature of the preset query vector output by the N-th decoder layer.

4. The method according to claim 3, characterized in that The step of inputting the fused query feature of the preset query vector output by the i-1th decoder layer and the second image feature into the i-th decoder layer to obtain the fused query feature of the preset query vector output by the i-th decoder layer comprises: The fused query features of the preset query vector output by the i-1th decoder layer and the second image features are input into the i-th decoder layer, the i-th decoder layer is used to determine the weight coefficient of the input fused query features relative to each feature in the second image features, and the weight coefficient is used to perform weighted summation on each feature in the second image features to obtain the fused query features of the preset query vector output by the i-th decoder layer.

5. The method according to claim 1, characterized in that The obtaining, based on the second image features and each of the second Transformer networks, an instance segmentation result of the image to be processed includes: Inputting the second image feature into a second neural network of the deep learning model to obtain a third image feature of the image to be processed, wherein the second neural network is used to perform feature conversion on the second image feature; The third image features of the image to be processed are respectively input into each of the second Transformer networks to obtain an instance segmentation result of the image to be processed.

6. The method according to claim 1, characterized in that The process of pre-training a deep learning model includes: Acquire a sample image and a plurality of sample query vectors, wherein the number of the sample query vectors is greater than the number of instances in the sample image; Inputting the sample image and each of the sample query vectors into a deep learning model to obtain a plurality of predicted segmentation results of the sample image, wherein the number of the predicted segmentation results is the same as the number of the sample query vectors; Determining a segmentation error of each instance in the sample image based on each predicted segmentation result of the sample image and each instance segmentation true value of the sample image; Adjusting the parameters of the deep learning model and the value of the sample query vector according to each of the segmentation errors; Other sample images are selected to continue training the deep learning model until a preset training end condition is met, thereby obtaining a pre-trained deep learning model and the plurality of preset query vectors.

7. The method according to claim 6, characterized in that The determining the segmentation error of each instance in the sample image based on each predicted segmentation result of the sample image and each instance segmentation true value of the sample image comprises: For each predicted segmentation result of the sample image, determining a multidimensional vector of the predicted segmentation result; For each instance segmentation true value of the sample image, determining a multidimensional vector of the instance segmentation true value; For each instance segmentation true value of the sample image, the multidimensional vector of the instance segmentation true value is matched with the multidimensional vector of each predicted segmentation result to obtain a predicted segmentation result that matches the instance segmentation true value; For each instance segmentation true value of the sample image, the error between the multidimensional vector of the instance segmentation true value and the multidimensional vector of the predicted segmentation result matching the instance segmentation true value is calculated to obtain the segmentation error of the instance corresponding to the instance segmentation true value.

8. The method according to claim 7, characterized in that The step of determining a multidimensional vector of each predicted segmentation result of the sample image comprises: For each predicted segmentation result of the sample image, determining a key point of an instance in the predicted segmentation result; The vectors from each pixel point to the key points of the instance in the predicted segmentation result are calculated respectively to obtain a multi-dimensional vector of the predicted segmentation result.

9. An image instance segmentation device, characterized in that: include: A module for acquiring an image to be processed, used for acquiring an image to be processed and a plurality of preset query vectors, wherein the number of the preset query vectors is greater than the number of instances in the image to be processed; A first image feature extraction module, used for inputting the image to be processed into a first neural network of a pre-trained deep learning model to obtain a first image feature of the image to be processed; an instance feature extraction module, configured to input the first image feature and each of the preset query vectors into a first Transformer network of the deep learning model to obtain a second image feature and a plurality of instance features, wherein the number of the instance features is the same as the number of the preset query vectors; A dynamic parameter determination module, used to input the instance features into the prediction network of the deep learning model to obtain the classification results and dynamic parameters of each instance feature; A dynamic parameter configuration module, used to configure parameters of the dynamic Transformer network of the deep learning model using each of the dynamic parameters to obtain a plurality of second Transformer networks, wherein the number of the second Transformer networks is the same as the number of the preset query vectors; An instance segmentation result determination module, used for obtaining an instance segmentation result of the image to be processed based on the second image features and each of the second Transformer networks; Wherein, the first Transformer network of the deep learning model includes an encoder and a decoder; The instance feature extraction module is specifically used to encode the first image feature using the encoder of the first Transformer network to obtain the second image feature; and input the second image feature and each of the preset query vectors into the decoder of the first Transformer network to obtain multiple instance features.

10. An electronic device, characterized in that: including a processor and a memory; The memory is used to store computer programs; A processor, for implementing any of the methods described in claims 1-8 when executing a program stored in a memory.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • End-to-end instance segmentation method based on instance query

    CN112927245A

  • Image matching method based on deep learning

    CN113283525A