Training method of image segmentation model and image segmentation method

By using the attention mechanism in the image segmentation model to calculate the attention features and indicating the foreground features of the target object based on the mask image, the problem of low training accuracy of the image segmentation model in the existing technology is solved, and a more efficient and accurate training effect is achieved.

CN120689350APending Publication Date: 2025-09-23TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410331325.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-21
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

During the training process of existing image segmentation models, due to the differences in the appearance of the target objects in different images, the prototype features determined by average pooling or clustering methods are generalized and have low accuracy, which in turn affects the model performance.

Method used

By obtaining image features and mask images of multiple groups of sample groups, the attention feature is calculated using the attention mechanism, and the image segmentation model is iteratively trained based on the feature. The mask image is used to indicate the foreground features of the target object, thereby improving training efficiency and accuracy.

Benefits of technology

The training efficiency and accuracy of the image segmentation model are improved, the number of training times is reduced, and the model's ability to recognize target objects is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689350A_ABST
    Figure CN120689350A_ABST
Patent Text Reader

Abstract

The invention provides an image segmentation model training method and an image segmentation method, and belongs to the technical field of artificial intelligence. Comprising the steps that multiple sample groups are acquired, and each sample group comprises image features and mask images of a first image and image features and mask images of a second image; for each sample group, processing the image feature of the second image and the mask image in the sample group through an image segmentation model to obtain a first attention feature; image features of the first image and prototype features of the target object in the second image are processed through an image segmentation model, a third mask image is obtained, the prototype features are partial features in the first attention features, and the partial features correspond to the mask image of the second image; and iteratively training the image segmentation model based on the third mask images and the first mask images of the plurality of sample groups. The method improves the training efficiency of the image segmentation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to an image segmentation model training method and an image segmentation method. Background Art

[0002] With the development of artificial intelligence (AI) technology, it is being applied in more fields. Image segmentation is also a key application area of ​​AI. Segmenting an image containing a target object is accomplished using a matching technique based on prototype features. This technique trains an image segmentation model using the prototype features of the target object in a first image and a second image. The image segmentation model is then used to perform image segmentation. The first image is the image containing the target object, and the second image is a different image containing the target object.

[0003] In the related art, during the training process of the image segmentation model, the prototype features of the target object in the second image are determined by the average pooling or clustering method, and then the image segmentation model segments the target object in the first image based on the prototype features. However, due to the differences in the appearance of the target object in different images, the prototype features determined by the average pooling or clustering method are generalized, that is, the accuracy is low, which leads to poor performance of the trained image segmentation model. Summary of the Invention

[0004] The present invention provides an image segmentation model training method and an image segmentation method, which improves the training efficiency of the image segmentation model. The technical solution is as follows:

[0005] In one aspect, a method for training an image segmentation model is provided, the method comprising:

[0006] Acquire multiple sample groups, each sample group including image features and a mask image of a first image and image features and a mask image of a second image, the first image and the second image respectively including a target object, and foregrounds of the mask image of the first image and the mask image of the second image respectively corresponding to the target object;

[0007] For each sample group, processing the image features of the second image and the mask image in the sample group using an image segmentation model to obtain a first attention feature, wherein the image segmentation model is used to segment the target object from the first image, and the first attention feature is a fusion of the image features of the second image and the features of the mask image;

[0008] processing, by the image segmentation model, image features of the first image and prototype features of the target object in the second image to obtain a third mask image, wherein the prototype features are partial features of the first attention features, the partial features correspond to the mask image of the second image, and the third mask image is used to indicate the object segmented from the first image;

[0009] The image segmentation model is iteratively trained based on the third mask image and the first mask image of each of the multiple sample groups.

[0010] In another aspect, an image segmentation method is provided, the method comprising:

[0011] Acquire image features of a first target image, image features of a second target image, and a mask image of the second target image, wherein the first image and the second image respectively include a target object, and a foreground of the mask image corresponds to the target object;

[0012] Processing the image features of the second target image and the mask image using an image segmentation model to obtain a first target attention feature, wherein the image segmentation model is obtained using the training method in any of the above embodiments, and the first target attention feature is a fusion of the image features of the second image and the features of the mask image;

[0013] The image features of the first target image and the prototype features of the target object in the first target image are processed through the image segmentation model to obtain a mask image of the first target image, wherein the prototype features are partial features in the first target attention features, and the partial features correspond to the mask image of the second target image, and the mask image of the first target image is used to indicate the target object segmented from the first target image.

[0014] In another aspect, a training device for an image segmentation model is provided, the device comprising:

[0015] an acquisition module, configured to acquire a plurality of sample groups, each sample group including image features and a mask image of a first image and image features and a mask image of a second image, the first image and the second image respectively including a target object, and foregrounds of the mask image of the first image and the mask image of the second image respectively corresponding to the target object;

[0016] a first processing module configured to process, for each sample group, the image features of the second image and the mask image in the sample group using an image segmentation model to obtain a first attention feature, wherein the image segmentation model is used to segment the target object from the first image, and the first attention feature is a fusion of the image features of the second image and the mask image;

[0017] a second processing module, configured to process the image features of the first image and the prototype features of the target object in the second image using the image segmentation model to obtain a third mask image, wherein the prototype features are partial features of the first attention features, the partial features correspond to the mask image of the second image, and the third mask image is used to indicate the object segmented from the first image;

[0018] The training module is configured to iteratively train the image segmentation model based on the third mask image and the first mask image of each of the multiple sample groups.

[0019] In some embodiments, the second processing module is configured to:

[0020] Processing the image features of the first image and the prototype features through an attention module in the image segmentation model to obtain a second attention feature, where the second attention feature is a feature that combines the prototype features and the image features of the first image;

[0021] The second attention feature is processed by a decoding module in the image segmentation model to obtain the third mask image, and the decoding module is used to obtain the mask image based on the input feature.

[0022] In some embodiments, the second processing module is configured to:

[0023] Processing the image features of the first image and the prototype features by the attention module to obtain cross-attention features corresponding to the image features of the first image and cross-attention features corresponding to the prototype features, where any cross-attention feature is a feature that combines the image features of the first image and the prototype features;

[0024] Processing the two cross-attention features respectively to obtain two self-attention features corresponding to the two cross-attention features;

[0025] Processing the two self-attention features to obtain two cross-attention features corresponding to the two self-attention features respectively;

[0026] The process of iterating a preset number of times to obtain cross-attention features and self-attention features; wherein, the two self-attention features used in the i+1th iteration are the two self-attention features obtained in the i-th iteration, i is an integer greater than 0, and the second attention feature is the self-attention feature corresponding to the image feature of the first image obtained in the last iteration.

[0027] In some embodiments, each sample group includes two image features of a first image and two image features of a second image, and the apparatus further includes:

[0028] An input-output module, configured to input the first image into a convolution module, wherein the outputs of the last two convolution layers of the convolution module are two image features of the first image, and the convolution module is configured to extract features of the input image; and input the second image into the convolution module, wherein the outputs of the last two convolution layers are two image features of the second image;

[0029] The training module is used to train the image segmentation model for each group of samples based on the average of the loss values ​​between the third mask images corresponding to each of the last two convolutional layers and the first mask image, where each third mask image is obtained based on the image features of the first image and the image features of the second image of the same convolutional layer.

[0030] In some embodiments, the apparatus further comprises:

[0031] a first determining module, configured to determine an image mask feature corresponding to a mask image of a first image, wherein the image mask feature is used to represent the mask image of the first image, and the image mask feature includes block features of a plurality of blocks in the mask image of the first image;

[0032] an adding module, configured to add a first noise to at least one block feature and add a second noise to the remaining block features of the plurality of block features to obtain a noise image feature, wherein a variance of the second noise is greater than a variance of the first noise;

[0033] a third processing module, configured to process the image features of the first image, the prototype features, and the noise image features through the attention module in the image segmentation model to obtain a third attention feature;

[0034] a fourth processing module, configured to process the third attention feature through a decoding module in the image segmentation model to obtain a fourth mask image, wherein the fourth mask image is used to indicate a block to which the first noise is added and a block to which the second noise is added in the mask image of the first image, and the decoding module is configured to obtain the mask image based on the input feature;

[0035] a second determining module, configured to determine a loss value between the fourth mask image and a fifth mask image, wherein the fifth mask image is a mask image corresponding to features of the noise image, and a foreground of the fifth mask image corresponds to a block to which the first noise is added;

[0036] The training module is configured to train the image segmentation model for each sample group based on a loss value between the third mask image and the first mask image and a loss value between the fourth mask image and the fifth mask image.

[0037] In some embodiments, the third processing module is configured to:

[0038] Processing the image features of the first image, the prototype features, and the noise image features through the attention module to obtain cross-attention features corresponding to the image features of the first image, cross-attention features corresponding to the prototype features, and cross-attention features corresponding to the noise image features, wherein any cross-attention feature is a feature that combines two of the image features of the first image, the prototype features, and the noise image features;

[0039] Processing the three cross-attention features respectively to obtain three self-attention features corresponding to the three cross-attention features;

[0040] Processing the three self-attention features to obtain three cross-attention features corresponding to the three self-attention features respectively;

[0041] The process of iterating a preset number of times to obtain cross-attention features and self-attention features; wherein, the three self-attention features used in the j+1th iteration are the three self-attention features obtained in the jth iteration, j is an integer greater than 0, and the third attention feature includes the self-attention feature corresponding to the image feature of the first image and the self-attention feature corresponding to the noise image feature obtained in the last iteration.

[0042] In another aspect, an image segmentation apparatus is provided, the apparatus comprising:

[0043] an acquisition module, configured to acquire image features of a first target image, image features of a second target image, and a mask image of the second target image, wherein the first image and the second image respectively include a target object, and a foreground of the mask image corresponds to the target object;

[0044] a first processing module, configured to process the image features of the second target image and the mask image using an image segmentation model to obtain a first target attention feature, wherein the image segmentation model is obtained using the training method in any of the above embodiments, and the first target attention feature is a fusion of the image features of the second image and the features of the mask image;

[0045] The second processing module is used to obtain a mask image of the first target image by using the image segmentation model to obtain the image features of the first target image and the prototype features of the target object in the first target image, wherein the prototype features are partial features in the first target attention features, and the partial features correspond to the mask image of the second target image, and the mask image of the first target image is used to indicate the target object segmented from the first target image.

[0046] In some embodiments, the second processing module is configured to:

[0047] Processing the image features of the first target image and the prototype features through the attention module in the image segmentation model to obtain a second target attention feature, where the second attention feature is a feature that combines the prototype features and the image features of the first image;

[0048] The second target attention feature is processed by a decoding module in the image segmentation model to obtain the mask image, and the decoding module is used to obtain the mask image based on the input feature.

[0049] In some embodiments, the second processing module is configured to:

[0050] Processing the image features of the first target image and the prototype features by the attention module to obtain cross-attention features corresponding to the image features of the first target image and cross-attention features corresponding to the prototype features, wherein any cross-attention feature is a feature that combines the image features of the first target image and the prototype features;

[0051] Processing the two cross-attention features respectively to obtain two self-attention features corresponding to the two cross-attention features;

[0052] Processing the two self-attention features to obtain two cross-attention features corresponding to the two self-attention features respectively;

[0053] The process of iterating a preset number of times to obtain cross-attention features and self-attention features; wherein, the two self-attention features used in the i+1th iteration are the two self-attention features obtained in the i-th iteration, i is an integer greater than 0, and the second target attention feature is the self-attention feature corresponding to the image feature of the first target image obtained in the last iteration.

[0054] On the other hand, a computer device is provided, which includes a processor and a memory, wherein the memory is used to store at least one program, and the at least one program is loaded and executed by the processor to implement the image segmentation model training method or image segmentation method in the embodiment of the present application.

[0055] On the other hand, a computer-readable storage medium is provided, in which at least one program is stored. The at least one program is loaded and executed by a processor to implement the image segmentation model training method or image segmentation method in the embodiment of the present application.

[0056] On the other hand, a computer program product is provided, which includes at least one program segment, and the at least one program segment is stored in a computer-readable storage medium. The processor of a computer device reads the at least one program segment from the computer-readable storage medium, and the processor executes the at least one program segment, so that the computer device executes the image segmentation model training method or image segmentation method described in any of the above-mentioned implementation methods.

[0057] An embodiment of the present application provides a training method for an image segmentation model, which performs an attention mechanism calculation on the image features of a second image and the mask image of the second image to obtain an attention feature. In this way, in the process of performing the attention mechanism calculation on the two features, since the foreground of the mask image points to the target object in the second image, the mask image mainly focuses on the features of the target object with high similarity to its foreground in the image features during the calculation process, and the features corresponding to the mask image in the obtained attention features can represent the target object in the second image in more dimensions and more refined attribute features. Therefore, the features corresponding to the mask image in the attention features are used as the prototype features of the target object, so that the reference value of the prototype features is high, and then the image segmentation model refers to the prototype features to process the first image, and the accuracy of the segmented object indicated by the obtained third mask image is high. In this way, the image segmentation model is iteratively trained based on the third mask image, so that each training effect is good and the accuracy is high, thereby reducing the number of training times and improving training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0059] Figure 1 This is a schematic diagram of an implementation environment provided by an embodiment of the present application;

[0060] Figure 2 This is a flowchart of a method for training an image segmentation model provided in an embodiment of the present application;

[0061] Figure 3 This is a flowchart of another image segmentation model training method provided in an embodiment of the present application;

[0062] Figure 4 is a schematic diagram of a first image, a second image, and respective mask images provided in an embodiment of the present application;

[0063] Figure 5 This is a flowchart of another image segmentation model training method provided in an embodiment of the present application;

[0064] Figure 6 This is a flowchart of another image segmentation model training method provided in an embodiment of the present application;

[0065] Figure 7 This is a flowchart of an image segmentation method provided by an embodiment of the present application;

[0066] Figure 8 is a flowchart of another image segmentation method provided in an embodiment of the present application;

[0067] Figure 9 is a schematic diagram of a segmentation result of a first target image provided in an embodiment of the present application;

[0068] Figure 10 This is an effect comparison diagram provided by an embodiment of the present application;

[0069] Figure 11 This is a block diagram of a training device for an image segmentation model provided in an embodiment of the present application;

[0070] Figure 12 is a block diagram of an image segmentation device provided in an embodiment of the present application;

[0071] Figure 13 This is a block diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION

[0072] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0073] In this application, the terms "first", "second", etc. are used to distinguish identical or similar items with substantially the same effects and functions. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor is there any limitation on the quantity and execution order.

[0074] In the present application, the term "at least one" means one or more, and the term "plurality" means two or more.

[0075] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, storage, display, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the sample groups involved in this application were all obtained with full authorization.

[0076] The following is an introduction to the professional terms involved in this application:

[0077] Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive field of computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. AI technology is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI domains. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0078] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning. Pretrained models are the latest development in deep learning, integrating these techniques.

[0079] Computer vision (CV) is the science of making machines "see." Specifically, it refers to using cameras and computers to replace the human eye in identifying, tracking, and measuring objects, and then performing image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems that can extract information from images or multidimensional data. Large model technology has brought significant changes to the development of computer vision technology. Pre-trained models in the field of vision, such as the Swin Transformer, ViT, V-MOE, and MAE, can be fine-tuned to quickly and widely apply to specific downstream tasks. Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / action recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies. It also includes common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0080] Base class: also known as the known class. The categories in the training dataset used to train the image segmentation model are collectively referred to as base classes.

[0081] New class: also called unseen class, which is different from other categories of base class.

[0082] Query image (i.e., first image): The input of the image segmentation model is an RGB image containing the target object to be segmented in the query image.

[0083] Query image mask (i.e., mask image): The mask map of the target object to be segmented in the query image is a 0-1 map, where a value of 0 represents that the pixel position belongs to the background area, and a value of 1 represents that the pixel position belongs to the foreground (target object) area.

[0084] Support image (also known as the second image): The input of the image segmentation model is an RGB image containing the target object to be segmented in the query image.

[0085] Support image mask (i.e. mask image): The input of the image segmentation model is a 0-1 image, where a value of 0 represents that the pixel position belongs to the background area, and a value of 1 represents that the pixel position belongs to the foreground (target object) area.

[0086] Image segmentation: The process of subdividing an image into multiple sub-regions (a collection of pixels), that is, the technology and process of dividing the image into several specific regions with unique properties and proposing targets of interest. It is a key step from image processing to image analysis.

[0087] Few-shot segmentation: During the training phase, both support and query images are sampled from the base class. Given a query image, a small number of support images, and their mask images, the image segmentation model first calculates the prototype features of the target object in the support images and then segments the target object in the query image based on the prototype features. During the inference phase, both support and query images are sampled from the new class, and the segmentation process is the same as during the training phase.

[0088] Patch feature (token): Each image is divided into N non-overlapping and same-sized patches of size patch_size × patch_size. The patch_size × patch_size pixels in each patch are encoded to obtain a C token Dimension token, the image consists of a total of patches num ×patch num (N) tokens, where N is an integer greater than 0.

[0089] MAP (Mask Average Pooling) algorithm: The prototype features are obtained by multiplying the image features of the support image by the mask of the support image, and then calculating the mean of each channel in the support image.

[0090] Prototype features: These are the collective characteristics of an object or category. For example, for the object "pants," prototype features could include fabric, color, texture, size, category, and other characteristics. They can also be implicit features learned by a neural network. Each object (or category) can include features across multiple dimensions.

[0091] The following is an introduction to the implementation environment involved in this application:

[0092] The training method of the image segmentation model provided in the embodiment of the present application can be executed by a computer device, which can be provided as a server or a terminal. The following is a schematic diagram of the implementation environment of the training method of the image segmentation model provided in the embodiment of the present application.

[0093] See also Figure 1 , Figure 1A schematic diagram of an implementation environment for a training method for an image segmentation model provided in an embodiment of the present application, the implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 can be directly or indirectly connected via wired or wireless communication, which is not limited in this application. In some embodiments, the server 102 is used to train the image segmentation model, and the trained image segmentation model is used to segment the target object from the image. A target application is installed on the terminal 101, and the target application is used to segment the target object from the image. In some embodiments, the trained image segmentation model is embedded in the terminal 101, and the terminal 101 segments the target object from the image through the image segmentation model. In other embodiments, the terminal 101 segments the target object from the image through the image segmentation model on the server 102.

[0094] In some embodiments, the terminal 101 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, an intelligent voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, a VR (Virtual Reality) device, an AR (Augmented Reality) device, etc., but is not limited thereto. In some embodiments, the server 102 is an independent server or a server cluster or distributed system composed of multiple servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In some embodiments, the server 102 is primarily responsible for computing tasks, and the terminal 101 is responsible for secondary computing tasks; alternatively, the server 102 is responsible for secondary computing services, and the terminal 101 is responsible for primary computing tasks; alternatively, the server 102 and the terminal 101 use a distributed computing architecture for collaborative computing.

[0095] See also Figure 2 , Figure 2 A flowchart of a method for training an image segmentation model provided in an embodiment of the present application, the method comprising the following steps.

[0096] 201. A computer device obtains multiple sample groups, each sample group including image features and a mask image of a first image, and image features and a mask image of a second image, the first image and the second image respectively including a target object, and foregrounds of the mask image of the first image and the mask image of the second image respectively corresponding to the target object.

[0097] In an embodiment of the present application, the first image and the second image include the same target object. The first image is an image from which the target object is to be segmented, the second image is a reference image of the first image, and the mask image of the second image is used to indicate the same object as the first image in the second image, that is, to indicate the target object to be segmented from the first image. The appearance of the target object in the first image and the second image may be the same or different, and the appearance includes but is not limited to the color, movement, size, etc. of the target object. The first image and the second image are different images. Optionally, the first image and the second image are different in other elements except the target object, that is, if the first image and the second image have the target object as the foreground, the backgrounds of the first image and the second image are different. Alternatively, the appearance of the target object in the first image and the second image is different, and the other elements except the target object are the same, that is, if the first image and the second image have the target object as the foreground, the backgrounds of the first image and the second image are the same.

[0098] In the embodiment of the present application, the mask image is a 0-1 image, where each pixel is represented by 0 or 1, 1 represents the foreground area in the image, that is, the area where the target object is located, and 0 represents the background area, that is, the area in the image other than the foreground area.

[0099] In an embodiment of the present application, the image feature of each image is a matrix, the image feature includes multiple vectors, the image is divided into multiple non-overlapping blocks, each vector represents a block, that is, each vector is a block feature of a block, and the image feature includes multiple block features.

[0100] 202. For each sample group, the computer device processes the image features of the second image and the mask image in the sample group through an image segmentation model to obtain a first attention feature. The image segmentation model is used to segment the target object from the first image. The first attention feature is a fusion of the image features of the second image and the features of the mask image.

[0101] In the embodiment of the present application, processing the image features of the second image and the mask image in the sample group refers to performing calculations on the two using a self-attention mechanism.

[0102] In an embodiment of the present application, the first attention feature includes features corresponding to image features of the second image and features corresponding to the mask image. By performing a self-attention mechanism on the two, the image features are made to focus on the foreground features of the target object indicated by the mask image, and the mask image is made to focus on the features of the target object in the image features. The features corresponding to the image features in the first attention feature are also image features that focus on the foreground features of the target object indicated by the mask image, and the features corresponding to the mask image in the first attention feature are also features of the mask image that focus on the features of the target object in the image features.

[0103] Since the mask image indicates the target object in the second image, the self-attention mechanism is calculated on the mask image and the image features. Since the mask image mainly focuses on the features of the target object in the image features that have high similarity with its foreground during the calculation process, the features corresponding to the mask image can represent the target object in the second image in more dimensions and more refined attribute features. Therefore, the features corresponding to the mask image in the first attention feature can be used as the prototype features of the target object in the second image.

[0104] 203. The computer device processes the image features of the first image and the prototype features of the target object in the second image through an image segmentation model to obtain a third mask image. The prototype features are partial features in the first attention features, and the partial features correspond to the mask image of the second image. The third mask image is used to indicate the object segmented from the first image.

[0105] In an embodiment of the present application, the image segmentation model is used to segment an object that is the same as the target object in the second image from the first image, that is, the image segmentation model is used to determine the characteristics of the target object in the image characteristics of the first image based on an indication of the prototype characteristics of the target object, and then segment the target object from the first image.

[0106] 204. The computer device iteratively trains the image segmentation model based on the third mask image and the first mask image of each of the multiple sample groups.

[0107] In an embodiment of the present application, for each sample group, the model parameters of the image segmentation model are adjusted based on the loss value between the third mask image and the first mask image of the sample group.

[0108] In an embodiment of the present application, a computer device iteratively trains an image segmentation model based on the third mask image and the first mask image of each of the multiple sample groups until a preset requirement is met. Meeting the preset requirement may mean that the loss value between the third mask image and the first mask image converges, or that the loss value reaches a preset threshold, or that the number of iterations reaches a preset number, which are not specifically limited herein.

[0109] In an embodiment of the present application, an image segmentation model is trained based on the image features and mask images of two images. Since the mask image indicates the target object to be segmented from the image, the trained image segmentation model can segment the object identical to the target object from the image.

[0110] An embodiment of the present application provides a training method for an image segmentation model, which performs an attention mechanism calculation on the image features of a second image and the mask image of the second image to obtain an attention feature. In this way, in the process of calculating the attention mechanism for the two features, since the foreground of the mask image points to the target object in the second image, the mask image mainly focuses on the features of the target object with high similarity to its foreground in the image features during the calculation process, and the features corresponding to the mask image in the obtained attention features can represent the target object in the second image in more dimensions and more refined attribute features. Therefore, the features corresponding to the mask image in the attention features are used as the prototype features of the target object, so that the reference value of the prototype features is high, and then the image segmentation model refers to the prototype features to process the first image, and the accuracy of the segmented object indicated by the obtained third mask image is high. In this way, the image segmentation model is iteratively trained based on the third mask image, so that each training effect is good and the accuracy is high, thereby reducing the number of training times and improving training efficiency.

[0111] above Figure 2 The basic process of the training method for the image segmentation model is as follows: Figure 3 Further introduction to the training method of image segmentation model. Figure 3 , Figure 3 A flowchart of a method for training an image segmentation model provided in an embodiment of the present application, the method comprising the following steps.

[0112] 301. A computer device obtains multiple sample groups, each sample group including image features and a mask image of a first image, and image features and a mask image of a second image, the first image and the second image respectively including a target object, and foregrounds of the mask image of the first image and the mask image of the second image respectively corresponding to the target object.

[0113] In some embodiments, the computer device obtains multiple sample groups based on the training data set. The computer device samples the multiple data groups from the training data set, and each data group includes a first image, a second image, a first mask image, and a second mask image. The computer device performs feature extraction on the images in each data group to obtain a sample group corresponding to the data group. Each sample group includes the image features and mask image of the first image and the image features and mask image of the second image. For example, see Figure 4 , Figure 4 Schematic diagram of a first image, a second image, and respective mask images provided in an embodiment of the present application.

[0114] In some embodiments, the computer device augments images in the training dataset to increase the number of images in the training dataset. Optionally, the computer device augments the images by rotating, cropping, translating, scaling, etc. The computer device may augment the images in the training dataset before sampling or after sampling, without specific limitation.

[0115] In an embodiment of the present application, the multiple data groups may include the same first image and different second images, and image features may be extracted once from the first image to obtain multiple sample groups.

[0116] In some embodiments, the computer device extracts image features using a convolution module, which includes multiple convolution layers. Optionally, the convolution module is a ResNet-50 network. The convolution module can be a functional module within the image segmentation model or a separate functional module outside the image segmentation model. Optionally, if the convolution module is a functional module within the image segmentation model, the computer device inputs the image into the image segmentation model and obtains image features using the convolution module within the image segmentation model.

[0117] In some embodiments, the computer device uses a convolution module to extract two image features of different spatial sizes from each image. The convolution module includes multiple convolution layers. As an image is processed sequentially through these layers, the features output by the final convolution layer become more accurate. Therefore, the outputs of the last two convolution layers in the convolution module are used as the two image features of the image.

[0118] In some embodiments, the image features of each image include tile features of multiple tiles in the image, each tile feature represents one of the multiple tiles, and the multiple tiles do not overlap. Accordingly, the process of extracting features from any image by a computer device includes the following steps: the computer device inputs the image into a convolution module to obtain initial image features, and then performs non-overlapping cropping on the initial image features to obtain image features including multiple tile features.

[0119] In some embodiments, each sample group includes two image features of a first image and two image features of a second image. The process of obtaining the image features by the computer device includes the following steps: the computer device inputs the first image into a convolution module in an image segmentation model, the outputs of the last two convolution layers of the convolution module are the two image features of the first image, and the second image is input into the convolution module, the outputs of the last two convolution layers are the two image features of the second image.

[0120] Optionally, different convolutional layers have different sizes, so the number of channels of the image features output by the last two convolutional layers is different. For example, the sizes of the two image features are: 1024×h×w; 2048×h×w, where 1024 and 2048 represent the number of channels, h represents the height of the image feature, and w represents the width of the image feature. Therefore, one of the image features is transformed through a convolutional network to keep its number of channels consistent with the other image feature.

[0121] 302. For each sample group, the computer device processes the image features of the second image and the mask image in the sample group through an image segmentation model to obtain a first attention feature. The image segmentation model is used to segment the target object from the first image. The first attention feature is a fusion of the image features of the second image and the features of the mask image.

[0122] Optionally, the computer device processes the image features and mask image of the second image in the sample group using an image segmentation model to obtain the first attention feature, comprising the following steps: the computer device performs convolution processing on the mask image of the second image using the image segmentation model to obtain image mask features of the mask image of the second image, wherein the image mask features represent the mask image of the second image; splicing the image features of the second image and the image mask features to obtain a spliced ​​feature, and processing the spliced ​​feature to obtain the first attention feature. Optionally, processing the spliced ​​feature refers to extracting self-attention features from the spliced ​​feature using the attention module.

[0123] It should be noted that the image features of the second image include block features of multiple blocks in the second image, and the image mask features include block features of multiple blocks in the mask image. The number of block features included in the image mask features is the same as the number of block features included in the image features, and their positions in the second image and the mask image correspond one-to-one. Furthermore, the dimensions of the multiple block features included in the image mask features are the same as the dimensions of the multiple block features included in the image features.

[0124] Each block feature is a vector, so splicing the image feature and the image mask feature of the second image means splicing multiple block features in the image feature and multiple block features in the image mask feature into a matrix, and the matrix is ​​the splicing feature.

[0125] It should be noted that the first attention feature includes multiple vectors, which correspond one-to-one to multiple vectors in the splicing features. Therefore, some vectors in these multiple vectors corresponding to the image feature positions are used as features corresponding to the image features, and some vectors in these multiple vectors corresponding to the image mask feature positions are used as features corresponding to the image mask features. The features corresponding to the image mask features are also the features corresponding to the mask image.

[0126] Optionally, the computer device can be configured to perform a step size of patch size , the convolution layer with a convolution kernel size of patch_size performs convolution processing on the mask image of the second image to obtain the P tokens The image mask features of tokens. Then include S tokens The image features of the second image with tokens are spliced ​​with the image mask features to obtain the spliced ​​features. Among them, tokens represents multiple block features, that is, multiple vectors, P tokens and S tokens is an integer greater than 1, and the two can be the same or different. The receptive field of each block feature in the spliced ​​feature is the image feature of the entire second image and the mask feature of the entire image. Finally, the P in the first attention feature is tokens The partial vectors corresponding to the tokens are taken out to obtain the prototype features.

[0127] For example, the computer device obtains the first attention feature by the following formula (1). The attention module includes multiple attention layers, and any attention layer calculates the self-attention feature by the following formula (1), and the input of the current attention layer is the output of the previous attention layer.

[0128]

[0129] Q=MLP(tokens), V=MLP(tokens), K=MLP(tokens) (1)

[0130] Among them, atten(Q,K,V) represents the first attention feature, Q represents the query vector, which is the product of the input of the current attention layer and the query weight parameter in the current attention layer. V represents the value vector, which is the product of the input of the current attention layer and the value weight parameter in the current attention layer. K represents the key vector, which is the product of the input of the current attention layer and the key weight parameter in the current attention layer. T represents transpose. Softmax represents the activation function. D represents the matrix dimension of K, and tokens represents the concatenated features including multiple vectors. MLP represents the Multi-Layer Perceptron (MLP) in the attention layer, which is a feedforward artificial neural network model that can map multiple input data sets to a single output data set.

[0131] 303. The computer device processes the image features and prototype features of the first image through the attention module in the image segmentation model to obtain a second attention feature. The second attention feature is a feature that combines the prototype feature and the image features of the first image. The prototype feature is a partial feature in the first attention feature, and the partial feature corresponds to the mask image of the second image.

[0132] In some embodiments, the computer device performs multiple iterative extractions of attention features on the image features and prototype features of the first image to improve the accuracy and significance of the obtained attention features. Optionally, the computer device processes the image features and prototype features of the first image using an attention module in an image segmentation model to obtain a second attention feature, including the following steps.

[0133] The computer device processes the image features and prototype features of the first image through the attention module to obtain cross-attention features corresponding to the image features of the first image and cross-attention features corresponding to the prototype features, where any cross-attention feature is a feature that integrates the image features and prototype features of the first image; processes the two cross-attention features separately to obtain two self-attention features corresponding to the two cross-attention features; processes the two self-attention features to obtain two cross-attention features corresponding to the two self-attention features; iterates the process of obtaining cross-attention features and self-attention features a preset number of times; wherein the two self-attention features used in the i+1th iteration are the two self-attention features obtained in the i-th iteration, i is an integer greater than 0, and the second attention feature is the self-attention feature corresponding to the image feature of the first image obtained in the last iteration.

[0134] The attention module is used to implement the attention mechanism, which includes a self-attention mechanism and a cross-attention mechanism. Processing two cross-attention features separately means performing a self-attention mechanism calculation on each cross-attention feature to obtain a self-attention feature corresponding to the cross-attention feature, with one cross-attention feature corresponding to one self-attention feature. Processing two self-attention features means performing a cross-calculation on the two self-attention features so that each obtained cross-attention feature is a fusion of the two self-attention features, with one self-attention feature corresponding to one cross-attention feature.

[0135] In an embodiment of the present application, cross-attention features and self-attention features are extracted multiple times from the image features and prototype features of the first image. By extracting features layer by layer, the features of the target object in the obtained second attention features are made more prominent, and the second attention features are made more accurate.

[0136] It should be noted that the cross-attention features and self-attention features corresponding to the image features have the same dimension as the image features, that is, they include the same number of vectors, and vectors with the same position in different features correspond to the same patch in the first image. The cross-attention features and self-attention features corresponding to the prototype features have the same dimension as the prototype features, that is, they include the same number of vectors, and vectors with the same position in different features correspond to the same patch in the mask image of the second image.

[0137] Optionally, the computer device obtains the cross-attention feature by the following formula (2). Wherein, the attention module includes multiple attention layers, the multiple attention layers include an attention layer for calculating the cross-attention feature, and the computer device calculates the cross-attention feature by the following formula (2).

[0138]

[0139] Q=MLP(tokens), V=MLP(F tokens ), K=MLP(F tokens ) (2)

[0140] Among them, atten(Q,K,V) represents the cross attention feature, tokens represents the prototype feature including multiple vectors, and F tokens Represents an image feature of a first image including a plurality of vectors. In the process of calculating the cross-attention feature corresponding to the image feature, the query vector is the product of the image feature and the query weight parameter in the current attention layer. The value vector is the product of the prototype feature and the median weight parameter in the current attention layer. The key vector is the product of the prototype feature and the key weight parameter in the current attention layer. In the process of calculating the cross-attention feature corresponding to the prototype feature, the query vector is the product of the prototype feature and the query weight parameter in the current attention layer. The value vector is the product of the image feature and the median weight parameter in the current attention layer. The key vector is the product of the image feature and the key weight parameter in the current attention layer.

[0141] In some embodiments, the computer device directly uses the cross-attention features obtained by the above formula as the cross-attention features corresponding to the image features and the prototype features respectively. In other embodiments, the computer device uses the cross-attention features obtained by the above formula as the initial cross-attention features. For image features, the sum of the initial cross-attention features corresponding to the image features and the image features is used as the cross-attention features corresponding to the image features. For prototype features, the sum of the initial cross-attention features corresponding to the prototype features and the prototype features is used as the cross-attention features corresponding to the prototype features. Similarly, in the subsequent iterative process, the sum of the initial cross-attention features corresponding to the self-attention features and the self-attention features is used as the cross-attention features corresponding to the self-attention features. The process of the computer device determining the self-attention features is implemented by the above formula (1), which will not be repeated here.

[0142] In some embodiments, after the computer device obtains the self-attention feature corresponding to the image feature in the self-attention feature output during the last iteration, it extracts the self-attention feature again to further enhance the significance of the target object feature in the image feature.

[0143] 304. The computer device processes the second attention feature through a decoding module in the image segmentation model to obtain a third mask image, where the decoding module is used to obtain the mask image based on the input feature, and the third mask image is used to indicate the object segmented from the first image.

[0144] In an embodiment of the present application, the second attention feature is a feature of the target object in the first image determined based on the prototype feature of the target object, and then the decoding module determines the area matching the second attention feature from the first image based on the second attention feature, and segments the area to obtain the object segmented from the first image.

[0145] In this embodiment of the present application, steps 303-304 are used to process the image features of the first image and the prototype features of the target object in the second image using the image segmentation model to obtain a third mask image. In this embodiment, the image features of the first image and the prototype features of the target object are input into the attention module, and the saliency of the features of the target object in the first image is enhanced by calculating the attention features. Image segmentation is then performed based on the second attention features with strong saliency of the target object features, resulting in more accurate segmented objects.

[0146] It should be noted that steps 303-304 are only an optional implementation method for implementing the process. The computer device can also implement the process through other optional implementation methods, which will not be repeated here.

[0147] 305. The computer device iteratively trains the image segmentation model based on the third mask image and the first mask image of each of the multiple sample groups.

[0148] In some embodiments, the process of iteratively training the image segmentation model by the above-mentioned computer device based on the third mask image and the first mask image of each of multiple groups of sample groups includes the following steps: the computer device determines the loss value between the third mask image and the first mask image for each group of sample groups, and adjusts the model parameters of the image segmentation model based on the loss value.

[0149] Optionally, the computer device adjusts model parameters of the image segmentation model based on a loss value between the third mask image and the first mask image using an SGD (Stochastic Gradient Descent) optimizer. The loss value in the embodiment of the present application can be a loss value of a mean square error loss function, a mean absolute loss function, a cross entropy cost function, or a regularization loss function, and is not specifically limited here.

[0150] In some embodiments, the loss value between the third mask image and the first mask image refers to a loss value between image mask features of the third mask image and image mask features of the first mask image.

[0151] In some embodiments, each sample group includes two image features of the first image and two image features of the second image. The two image features output by the same convolutional layer are used to determine a third mask image. Therefore, the process of the computer device training the image segmentation model based on the third mask images and the first mask images of the multiple sample groups includes the following steps: for each sample group, the computer device adjusts the model parameters of the image segmentation model based on the average of the loss values ​​between the third mask images corresponding to the last two convolutional layers and the first mask images, and each third mask image is obtained based on the image features of the first image and the image features of the second image of the same convolutional layer.

[0152] In this embodiment, the increase in segmentation results is achieved by increasing the number of image features, and then the image segmentation model is trained by integrating multiple segmentation results. That is, the model parameters of the image segmentation model are adjusted based on multiple segmentation results, so that the adjustment of the model parameters is more accurate, thereby improving the training effect and training efficiency of the model.

[0153] In other embodiments, the image features of the two convolutional layers can also be cross-combined, that is, a third mask image is obtained based on the image features of the first image and the image features of the second image of the convolutional layers of different layers, so that the image segmentation model is trained based on the average of the loss values ​​between the four third mask images and the first mask image.

[0154] An embodiment of the present application provides a training method for an image segmentation model, which performs an attention mechanism calculation on the image features of a second image and the mask image of the second image to obtain an attention feature. In this way, in the process of calculating the attention mechanism for the two features, since the foreground of the mask image points to the target object in the second image, the mask image mainly focuses on the features of the target object with high similarity to its foreground in the image features during the calculation process, and the features corresponding to the mask image in the obtained attention features can represent the target object in the second image in more dimensions and more refined attribute features. Therefore, the features corresponding to the mask image in the attention features are used as the prototype features of the target object, so that the reference value of the prototype features is high. The attention mechanism is calculated on the prototype features and the image features of the first image to obtain the attention features. In the process of calculating the attention mechanism on the two features, since the prototype features point to the target object, the image segmentation model refers to the prototype features to extract the features of the target object in the image features of the first image layer by layer, so that the features of the target object in the obtained second attention features are more significant. Then, the image segmentation model performs image segmentation based on the second attention features, and the accuracy of the segmented object indicated by the obtained third mask image is high. In this way, the image segmentation model is iteratively trained based on the third mask image, so that each training effect is good and the accuracy is high, thereby reducing the number of training times and improving training efficiency.

[0155] Through the above Figure 2-3 The training process of the image segmentation model is introduced in the embodiment. The above embodiment is described as an example in which the image segmentation model is not trained for noise detection. Figure 5 The embodiment of the present invention takes the noise detection training of the image segmentation model as an example to illustrate. Figure 5 , Figure 5 This is a flowchart of a method for training an image segmentation model provided in an embodiment of the present application, which includes the following steps.

[0156] 501. A computer device obtains multiple sample groups, each sample group including image features and a mask image of a first image, and image features and a mask image of a second image, the first image and the second image respectively including a target object, and foregrounds of the mask image of the first image and the mask image of the second image respectively corresponding to the target object.

[0157] 502. For each sample group, the computer device processes the image features of the second image and the mask image in the sample group through an image segmentation model to obtain a first attention feature. The image segmentation model is used to segment the target object from the first image. The first attention feature is a fusion of the image features of the second image and the features of the mask image.

[0158] 503. The computer device processes the image features and prototype features of the first image through the attention module in the image segmentation model to obtain a second attention feature. The second attention feature is a feature that integrates the prototype feature and the image features of the first image. The prototype feature is a partial feature in the first attention feature, and the partial feature corresponds to the mask image of the second image.

[0159] 504. The computer device processes the second attention feature through a decoding module in the image segmentation model to obtain a third mask image, where the decoding module is used to obtain the mask image based on the input feature, and the third mask image is used to indicate the object segmented from the first image.

[0160] In the embodiment of the present application, steps 501-504 are similar to the above steps 301-304 and are not described again here.

[0161] 505. The computer device determines an image mask feature corresponding to the mask image of the first image, where the image mask feature is used to represent the mask image of the first image, and the image mask feature includes block features of multiple blocks in the mask image of the first image; a first noise is added to at least one block feature, and a second noise is added to the remaining block features of the multiple block features to obtain a noise image feature, where the variance of the second noise is greater than the variance of the first noise.

[0162] In some embodiments, the process of the above-mentioned computer device determining the image mask features corresponding to the mask image of the first image includes the following steps: the computer device performs convolution processing on the mask image of the first image through the image segmentation model to obtain the image mask features corresponding to the mask image of the first image. The image mask features of the mask image of the second image include block features of multiple blocks in the mask image. The number of multiple block features included in the image mask features is the same as the number of multiple block features included in the image features, and they correspond one to one, that is, the mask image and the first image are divided into blocks in the same way. Furthermore, the dimensions of the multiple block features included in the image mask features are the same as the dimensions of the multiple block features included in the image features of the first image. Each block feature is a vector. The remaining block features in the multiple block features are also the block features to which the first noise is not added.

[0163] In the embodiments of the present application, the first noise and the second noise can be set and modified as needed. For example, if both the first noise and the second noise are Gaussian noise, and the variance of the second noise is greater than the variance of the first noise, the variances of the two noises can be set and modified as needed. For example, if the variance of the first noise is 10 and the variance of the second noise is 25, the larger the variance, the more severe the noise, and the smaller the variance, the milder the noise.

[0164] In the embodiment of the present application, the number of tiles in the plurality of tiles to which the first noise is added can be set and changed as needed, and a portion of the tiles can be randomly selected to add the first noise. For example, half of the tiles in the mask image are added to the noise area. Optionally, the computer device can obtain the noise image features using the following formula (3).

[0165] m1=(m+noise)*w+(1-)*heavy_noise (3)

[0166] Where m1 represents the noise image feature, a matrix. m represents the image mask feature, a matrix. noise represents the first noise. w is a matrix that corresponds one-to-one with each pixel in the mask image. Pixels with the first noise added to the matrix have a value of 1, while pixels with the second noise added have a value of 0. heavy_noise represents the second noise.

[0167] 506. The computer device processes the image features, prototype features, and noise image features of the first image through the attention module in the image segmentation model to obtain a third attention feature.

[0168] In some embodiments, the above-mentioned computer device processes the image features, prototype features and noise image features of the first image through the attention module in the image segmentation model to obtain the third attention feature, which includes the following steps: the computer device processes the image features, prototype features and noise image features of the first image through the attention module to obtain the cross-attention feature corresponding to the image feature of the first image, the cross-attention feature corresponding to the prototype feature and the cross-attention feature corresponding to the noise image feature, and any cross-attention feature is a feature that integrates at least two of the image features, prototype features and noise image features of the first image; the three cross-attention features are processed respectively to obtain three self-attention features corresponding to the three cross-attention features; the three self-attention features are processed to obtain three cross-attention features corresponding to the three self-attention features; the process of iterating a preset number of times to obtain the cross-attention features and the self-attention features; wherein, the three self-attention features used in the j+1th iteration are the three self-attention features obtained in the jth iteration, j is an integer greater than 0, and the third attention feature includes the self-attention feature corresponding to the image feature of the first image and the self-attention feature corresponding to the noise image feature obtained in the last iteration.

[0169] Here, processing the three cross-attention features separately means performing a self-attention mechanism calculation on each cross-attention feature to obtain a self-attention feature corresponding to the cross-attention feature, that is, one cross-attention feature corresponds to one self-attention feature. Processing the three self-attention features means performing a cross-calculation on the three self-attention features so that each obtained cross-attention feature incorporates at least two of the three self-attention features, and one self-attention feature corresponds to one cross-attention feature.

[0170] In an embodiment of the present application, cross-attention features and self-attention features are extracted multiple times on the image features, prototype features and noise image features of the first image. By extracting features layer by layer, the features of the noise in the third attention feature are made more prominent, and the third attention feature is made more accurate.

[0171] It should be noted that, in the process of determining the cross-attention features corresponding to the image features of the first image, the initial cross-attention features between the image features and the prototype features, as well as the initial cross-attention features between the image features and the image noise features, are calculated respectively. Then, both initial cross-attention features are added to the image features of the first image to obtain the cross-attention features corresponding to the image features of the first image. The process of determining the cross-attention features corresponding to the prototype features and the noise image features is similar to this, and this process corresponds to the first iteration process. The subsequent iteration processes are similar to this. In the process of calculating the cross-attention based on any self-attention feature, the initial cross-attention feature corresponding to the self-attention feature is added to the self-attention feature to obtain the cross-attention feature corresponding to the self-attention feature.

[0172] It should be noted that in the process of calculating the cross-attention based on the image features, prototype features and noise image features of the first image, it is necessary to ensure that the noise image features can pay attention to the image features and prototype features of the first image, while the image features and prototype features of the first image cannot pay attention to the noise image features, that is, the cross-attention features corresponding to the image features are only fused with the prototype features, the cross-attention features corresponding to the prototype features are only fused with the image features, and the cross-attention images corresponding to the noise image features are fused with the image features and the prototype features.

[0173] Therefore, in the process of determining the cross-attention feature corresponding to the image feature of the first image, when calculating the initial cross-attention feature between the image feature and the image noise feature, the image noise feature is assigned a value of 0, or the sum of the initial cross-attention feature between the image feature and the prototype feature and the image feature is directly used as the cross-attention feature corresponding to the image feature. The process of determining the cross-attention feature corresponding to the prototype feature is similar. This process corresponds to the first iteration process, and subsequent iteration processes are similar.

[0174] Among them, in the process of calculating the initial cross-attention features of the noise image features relative to the image features, the query vector is the product of the noise image features and the query weight parameters in the current attention layer. The value vector is the product of the image features and the median weight parameters in the current attention layer. The key vector is the product of the image features and the key weight parameters in the current attention layer. In the process of calculating the initial cross-attention features of the noise image features relative to the prototype features, the query vector is the product of the noise image features and the query weight parameters in the current attention layer. The value vector is the product of the prototype features and the median weight parameters in the current attention layer. The key vector is the product of the prototype features and the key weight parameters in the current attention layer. These two initial cross-attention features are then added to the noise image features to obtain the cross-attention features corresponding to the noise image features. The subsequent iterative process is the same.

[0175] In some embodiments, after the computer device obtains the self-attention features corresponding to the image features and noise image features in the self-attention features output during the last iteration, it extracts the self-attention features of this part of the self-attention features again to further enhance the significance of the noise features in the image features and noise image features.

[0176] It should be noted that, since the process of calculating the second attention feature in step 504 also includes the process of calculating the cross-attention features and self-attention features corresponding to the image features and prototype features of the first image, and the cross-attention features and self-attention features of both do not focus on the noise image features. Therefore, step 506 can directly reuse the cross-attention features and self-attention features obtained in step 504. The number of iterations in step 504 and step 506 can be the same or different. In this embodiment, the same number of iterations is used as an example for explanation.

[0177] It should be noted that the serial numbers of step 504 and step 506 are only for ease of explanation and are not used to limit the execution order of the two. If step 504 is executed before step 506, step 506 can reuse the cross-attention features and self-attention features in step 504. If step 506 is executed before step 504, step 504 can reuse the cross-attention features and self-attention features in step 506. Alternatively, if the two are executed at the same time, it is only necessary to calculate the cross-attention features and self-attention features of the same parts of the two once, and the results are used separately in the two steps.

[0178] 507. The computer device processes the third attention feature through a decoding module in the image segmentation model to obtain a fourth mask image, where the fourth mask image is used to indicate the blocks with the first noise added and the blocks with the second noise added in the mask image of the first image. The decoding module is used to obtain the mask image based on the input feature.

[0179] In an embodiment of the present application, the image segmentation model is used to segment the blocks with the first noise added and the blocks with the second noise added from the mask image of the first image, that is, the image segmentation module is used to determine the blocks with the first noise added and the blocks with the second noise added in the mask image based on the third attention feature.

[0180] 508. The computer device determines a loss value between the fourth mask image and the fifth mask image, where the fifth mask image is a mask image corresponding to the noise image feature, and the foreground of the fifth mask image corresponds to the block to which the first noise is added.

[0181] In some embodiments, the loss value between the fourth mask image and the fifth mask image refers to a loss value between the image mask features of the fourth mask image and the image mask features of the fifth mask image.

[0182] The image mask feature of the fifth mask image corresponding to the noise image feature can be expressed as m′=w*m, where m′ represents the image mask feature of the fifth mask image.

[0183] 509. The computer device trains an image segmentation model for each sample group based on a loss value between the third mask image and the first mask image and a loss value between the fourth mask image and the fifth mask image.

[0184] It should be noted that the computer device iteratively trains the image segmentation model based on multiple groups of sample groups, and the computer device iteratively trains the image segmentation model based on the third mask image, first mask image, fourth mask image and fifth mask image corresponding to the multiple sample groups respectively.

[0185] See also Figure 6 , Figure 6 This is a training flow chart for an image segmentation model provided in an embodiment of the present application. The image segmentation model's attention module first performs self-attention calculation on the image features of the second image and the mask image to obtain a first attention feature, and the features of the first attention feature corresponding to the mask image are used as prototype features. Cross-attention and self-attention calculations are then performed on the prototype features, the image features of the first image, and the noise image features to obtain a third attention feature. Finally, the third attention feature is input into the decoding module of the image segmentation model to obtain a third mask image.

[0186] The method provided in the embodiment of the present application mainly includes three parts, one part is the extraction of image features. One part is the extraction of prototype features, which is implemented by an encoder. One part is the image segmentation guided by denoising, which is implemented by an attention module and a decoding module. Typically, given a first image, a second image and its corresponding mask image, a shared network is used to extract features corresponding to the features respectively. The shared network is pre-trained on ImageNet (a data set), and its parameters are fixed during training and reasoning.

[0187] In an embodiment of the present application, the trained image segmentation model can determine the areas with severe noise and slight noise in the image. Since there is generally noise interference at the edge of the target object in the image, the image segmentation model is not accurate in segmenting the edge of the target object. In this embodiment, since the image segmentation model can automatically detect severe noise and slight noise in the image, and then based on the rule that the area with slight noise is generally the foreground area where the target object is located and the area with severe noise is generally the background area, the auxiliary image segmentation model segments the target object from the image, divides the area with slight noise into the foreground area, and divides the area with severe noise into the background area, so that the edge of the segmented object is clearer, and the segmented target object is more accurate.

[0188] Figure 2 、 3 , 6 is a process of training an image segmentation model, based on Figure 7 The embodiment of the present invention introduces the use process of the image segmentation model. Figure 7 , Figure 7 This is a flowchart of an image segmentation method provided in an embodiment of the present application. The image segmentation model used in this method is the image segmentation model trained by any of the above embodiments. The method includes the following steps.

[0189] 701. A computer device obtains image features of a first target image, image features of a second target image, and a mask image of the second target image. The first target image and the second target image respectively include target objects, and a foreground of the mask image corresponds to the target object.

[0190] In an embodiment of the present application, the first target image and the second target image include the same target object. The first target image is the image from which the target object is to be segmented, the second image is a reference image for the first image, and the mask image is used to indicate the same object in the second target image as in the first target image, that is, to indicate the target object to be segmented from the first target image.

[0191] 702. The computer device processes the image features of the second target image and the mask image through an image segmentation model to obtain a first target attention feature, where the first target attention feature is a fusion of the image features of the second image and the features of the mask image.

[0192] In the embodiment of the present application, processing the image features and the mask image refers to performing calculations on the two using a self-attention mechanism.

[0193] 703. The computer device processes the image features of the first target image and the prototype features of the target object in the first target image through an image segmentation model to obtain a mask image of the first target image, where the prototype features are partial features of the first target attention features, and the partial features correspond to the mask image of the second target image. The mask image of the first target image is used to indicate the target object segmented from the first target image.

[0194] In an embodiment of the present application, the image segmentation model extracts the features of the target object from the image features of the first target image based on the indication of the prototype features, and then segments the target object from the first target image.

[0195] An embodiment of the present application provides an image segmentation method, which performs an attention mechanism calculation on the image features of the second target image and the mask image of the second target image to obtain an attention feature. In this way, in the process of calculating the attention mechanism for the two features, since the foreground of the mask image points to the target object in the second target image, the mask image mainly focuses on the features of the target object with high similarity to its foreground in the image features during the calculation process, and the features corresponding to the mask image in the obtained attention features can represent the target object in the second target image in more dimensions and more refined attribute features. Therefore, the features corresponding to the mask image in the attention features are used as the prototype features of the target object, so that the reference value of the prototype features is high, and then the image segmentation model refers to the prototype features to process the first target image, and the accuracy of the segmented object indicated by the obtained mask image is high.

[0196] above Figure 7 This is the basic process of the image segmentation method. Figure 8 The embodiment of the image segmentation method is further introduced. Figure 8 , Figure 8 This is a flowchart of an image segmentation method provided in an embodiment of the present application, which includes the following steps.

[0197] 801. A computer device obtains image features of a first target image, image features of a second target image, and a mask image of the second target image. The first target image and the second target image respectively include target objects, and a foreground of the mask image corresponds to the target object.

[0198] In the embodiment of the present application, step 801 is similar to step 701 and will not be described again here.

[0199] 802. The computer device processes the image features of the second target image and the mask image through an image segmentation model to obtain a first target attention feature, where the first target attention feature is a fusion of the image features of the second image and the features of the mask image.

[0200] In an embodiment of the present application, the computer device processes the image features and mask image of the second target image through an image segmentation model to obtain the first target attention feature. The process is the same as the process in step 302 in which the computer device processes the image features and mask image of the second image in the sample group through an image segmentation model to obtain the first attention feature, and will not be repeated here.

[0201] 803. The computer device processes the image features and prototype features of the first target image through the attention module in the image segmentation model to obtain a second target attention feature. The prototype feature is a partial feature of the first target attention feature, and the partial feature corresponds to the mask image of the second target image. The second target attention feature is a feature that integrates the prototype feature and the image feature of the first target image.

[0202] In some embodiments, the above-mentioned computer device processes the image features and prototype features of the first target image through the attention module in the image segmentation model to obtain the second target attention features, which includes the following steps: the computer device processes the image features and prototype features of the first target image through the attention module to obtain cross-attention features corresponding to the image features of the first target image and cross-attention features corresponding to the prototype features, where any cross-attention feature is a feature that integrates the image features and prototype features of the first target image; processes the two cross-attention features separately to obtain two self-attention features corresponding to the two cross-attention features; processes the two self-attention features to obtain two cross-attention features corresponding to the two self-attention features; iterates a preset number of times to obtain cross-attention features and self-attention features; wherein the two self-attention features used in the i+1th iteration are the two self-attention features obtained in the i-th iteration, i is an integer greater than 0, and the second target attention feature is the self-attention feature corresponding to the image features of the first target image obtained in the last iteration.

[0203] In an embodiment of the present application, cross-attention features and self-attention features are extracted multiple times from the image features and prototype features of the first target image. By extracting features layer by layer, the features of the target object in the second attention feature are made more prominent, and the second target attention feature is made more accurate.

[0204] 804. The computer device processes the second target attention feature through a decoding module in the image segmentation model to obtain a mask image of the first target image. The decoding module is used to obtain the mask image based on the input feature. The mask image of the first target image is used to indicate the target object segmented from the first target image.

[0205] In an embodiment of the present application, steps 803-804 are implemented to process the image features of the first target image and the prototype features of the target object in the second target image using an image segmentation model to obtain a mask image. In this embodiment, the image features of the first target image and the prototype features of the target object are input into the attention module, and the saliency of the features of the target object in the first target image is enhanced by calculating the attention features. Image segmentation is then performed based on the second target attention features with strong saliency of the target object features, resulting in more accurate segmented objects.

[0206] It should be noted that steps 803-804 are only an optional implementation method for implementing the process. The computer device can also implement the process through other optional implementation methods, which will not be described in detail here.

[0207] The embodiment of the present application provides an image segmentation method, which performs an attention mechanism calculation on the image features of the second target image and the mask image of the second target image to obtain an attention feature. In this way, in the process of calculating the attention mechanism for the two features, since the foreground of the mask image points to the target object in the second target image, the mask image mainly focuses on the features of the target object with high similarity to its foreground in the image features during the calculation process, and the features corresponding to the mask image in the obtained attention feature can represent the target object in the second target image in more dimensions and more refined attribute features. Therefore, the features corresponding to the mask image in the attention feature are used as the prototype features of the target object, so that the reference value of the prototype feature is high. The attention mechanism is calculated on the prototype feature and the image feature of the first image to obtain the attention feature. In this way, in the process of calculating the attention mechanism for the two features, since the prototype feature points to the target object, the image segmentation model refers to the prototype feature to extract the features of the target object in the image features of the first target image layer by layer, so that the features of the target object in the obtained second attention feature are more significant, and thus the image segmentation model performs image segmentation based on the second attention feature, and the accuracy of the segmented object indicated by the obtained mask image is high.

[0208] In an embodiment of the present application, the image features of the second image and its mask image are converted into block features (tokens) of the same dimension, and then the prototype features corresponding to the mask image are obtained based on the transformer framework autoregression. In this way, the embodiment of the present application avoids the use of the MAP algorithm, so it will not cause the loss of image spatial information. In addition, the embodiment of the present application inputs the prototype features and the image features of the first image into the transformer, and enhances the saliency of the features of the target object in the image by calculating attention, so that the target object in the first image can be accurately segmented. In addition, in order to enhance the robustness of the image segmentation model to diverse inputs and learn clearer decision boundaries, the embodiment of the present application puts the mask image (GT mask) of the first image into an independent branch similar to the prototype features, and randomly adds noise, so that the method is applied to the field of small sample segmentation, which can improve the decision boundary of the image segmentation model, that is, make the edges of the segmented target objects clearer.

[0209] For example, see Figure 9 , Figure 9 FIG. 1 is a schematic diagram of a segmentation result of a first target image provided by an embodiment of the present application, wherein a tree is a target object segmented from the first target image.

[0210] Small-sample image segmentation has great application value in a variety of businesses. In the early stages of a new project, when data is scarce, a small-sample segmentation model trained on a base class (previous base class data) can be quickly migrated to the current new project, allowing for rapid iteration to produce a model that demonstrates the effect. In the mid-term, some target categories may present long-tail problems due to their low frequency of occurrence. Small-sample image segmentation models can be specifically designated for these categories to achieve better results. In the later stages of a project, if new user requirements emerge, the small-sample image segmentation model can quickly iterate on the new category and provide feedback, significantly shortening the project cycle.

[0211] In the embodiment of the present application, the trained image segmentation model is used to perform image segmentation on the new class to verify the network effect of the trained image segmentation model based on Transformer denoising guidance. Figure 10 , Figure 10 This is a comparison chart of the effects provided by the embodiment of this application. The left column shows various image segmentation methods. Score 0 to score 3 are the results obtained by four cross-validations, namely the miou (Mean Intersection over Union) scores. Among them, the dataset PASCAL-5 is used. i For verification. PASCAL-5i includes images from PASCAL VOC 2012 (an image library) and additional annotated images from SBD (an image library). This dataset contains 20 categories. When performing verification, these 20 categories are evenly divided into four parts for cross-validation. Based on this dataset, the method provided in this application and several other methods were verified respectively. DifFSS (diffusion model for semantic segmentation) among other methods uses a diffusion model with the help of external data. The rest of the methods are based on PASCAL-5. i Train and test above. Figure 10 It can be seen that the miou score of the method provided in the embodiment of the present application is higher than that of other methods, indicating that the effect of this method is better than other methods.

[0212] Figure 11 This is a block diagram of a training device for an image segmentation model provided in accordance with an embodiment of the present application. Figure 11 , the device comprises:

[0213] An acquisition module 1101 is configured to acquire multiple sample groups, each sample group including image features and a mask image of a first image, and image features and a mask image of a second image, wherein the first image and the second image respectively include a target object, and the foreground of the mask image of the first image and the foreground of the mask image of the second image respectively correspond to the target object;

[0214] A first processing module 1102 is configured to process, for each sample group, the image features of the second image and the mask image in the sample group using an image segmentation model to obtain a first attention feature. The image segmentation model is configured to segment the target object from the first image. The first attention feature is a fusion of the image features of the second image and the mask image.

[0215] a second processing module 1103 configured to process the image features of the first image and the prototype features of the target object in the second image using an image segmentation model to obtain a third mask image, wherein the prototype features are partial features of the first attention features, the partial features correspond to the mask image of the second image, and the third mask image is used to indicate the object segmented from the first image;

[0216] The training module 1104 is configured to iteratively train the image segmentation model based on the third mask image and the first mask image of each of the multiple sample groups.

[0217] In some embodiments, the second processing module 1103 is configured to:

[0218] The image features and prototype features of the first image are processed by the attention module in the image segmentation model to obtain a second attention feature, which is a feature that combines the prototype feature and the image feature of the first image;

[0219] The second attention feature is processed by a decoding module in the image segmentation model to obtain a third mask image. The decoding module is used to obtain the mask image based on the input feature.

[0220] In some embodiments, the second processing module 1103 is configured to:

[0221] Processing the image features and the prototype features of the first image through the attention module to obtain cross-attention features corresponding to the image features of the first image and cross-attention features corresponding to the prototype features, where any cross-attention feature is a feature that combines the image features of the first image and the prototype features;

[0222] The two cross-attention features are processed separately to obtain two self-attention features corresponding to the two cross-attention features;

[0223] The two self-attention features are processed to obtain two cross-attention features corresponding to the two self-attention features;

[0224] The process of iterating a preset number of times to obtain cross-attention features and self-attention features; wherein, the two self-attention features used in the i+1th iteration are the two self-attention features obtained in the i-th iteration, i is an integer greater than 0, and the second attention feature is the self-attention feature corresponding to the image feature of the first image obtained in the last iteration.

[0225] In some embodiments, each sample group includes two image features of a first image and two image features of a second image, and the apparatus further includes:

[0226] The input-output module is used to input the first image into the convolution module. The outputs of the last two convolution layers of the convolution module are two image features of the first image. The convolution module is used to extract the features of the input image. The second image is input into the convolution module. The outputs of the last two convolution layers are two image features of the second image.

[0227] The training module 1104 is used to adjust the model parameters of the image segmentation model for each group of samples based on the average of the loss values ​​between the third mask images corresponding to each of the last two convolutional layers and the first mask image, where each third mask image is obtained based on the image features of the first image and the image features of the second image of the same convolutional layer.

[0228] In some embodiments, the apparatus further comprises:

[0229] a first determining module, configured to determine an image mask feature corresponding to the mask image of the first image, wherein the image mask feature is used to represent the mask image of the first image, and the image mask feature includes block features of a plurality of blocks in the mask image of the first image;

[0230] an adding module, configured to add a first noise to at least one block feature and add a second noise to the remaining block features of the plurality of block features to obtain a noise image feature, wherein a variance of the second noise is greater than a variance of the first noise;

[0231] A third processing module is configured to process the image features, prototype features, and noise image features of the first image through the attention module in the image segmentation model to obtain a third attention feature;

[0232] a fourth processing module, configured to process the third attention feature through a decoding module in the image segmentation model to obtain a fourth mask image, wherein the fourth mask image is used to indicate a block to which the first noise is added and a block to which the second noise is added in the mask image of the first image, and the decoding module is configured to obtain the mask image based on the input feature;

[0233] a second determining module, configured to determine a loss value between a fourth mask image and a fifth mask image, wherein the fifth mask image is a mask image corresponding to features of the noise image, and a foreground of the fifth mask image corresponds to a block to which the first noise is added;

[0234] The training module 1104 is configured to:

[0235] For each sample group, an image segmentation model is trained based on a loss value between the third mask image and the first mask image and a loss value between the fourth mask image and the fifth mask image.

[0236] In some embodiments, the third processing module is configured to:

[0237] Processing the image features, prototype features, and noise image features of the first image through an attention module to obtain cross-attention features corresponding to the image features of the first image, cross-attention features corresponding to the prototype features, and cross-attention features corresponding to the noise image features, wherein any cross-attention feature is a feature that integrates at least two of the image features of the first image, the prototype features, and the noise image features;

[0238] The three cross-attention features are processed separately to obtain three self-attention features corresponding to the three cross-attention features;

[0239] The three self-attention features are processed to obtain three cross-attention features corresponding to the three self-attention features;

[0240] The process of iterating a preset number of times to obtain cross-attention features and self-attention features; wherein, the three self-attention features used in the j+1th iteration are the three self-attention features obtained in the jth iteration, j is an integer greater than 0, and the third attention feature includes the self-attention feature corresponding to the image feature of the first image and the self-attention feature corresponding to the noise image feature obtained in the last iteration.

[0241] An embodiment of the present application provides a training device for an image segmentation model, which performs an attention mechanism calculation on the image features of a second image and the mask image of the second image to obtain an attention feature. In this way, in the process of calculating the attention mechanism for the two features, since the foreground of the mask image points to the target object in the second image, the mask image mainly focuses on the features of the target object with high similarity to its foreground in the image features during the calculation process, and the features corresponding to the mask image in the obtained attention feature can represent the target object in the second image in more dimensions and more refined attribute features. Therefore, the features corresponding to the mask image in the attention feature are used as the prototype features of the target object, so that the reference value of the prototype feature is high, and then the image segmentation model refers to the prototype feature to process the first image, and the accuracy of the segmented object indicated by the obtained third mask image is high. In this way, the image segmentation model is iteratively trained based on the third mask image, so that each training effect is good and the accuracy is high, thereby reducing the number of training times and improving training efficiency.

[0242] Figure 12 This is a block diagram of an image segmentation device provided according to an embodiment of the present application. Figure 12 , the device comprises:

[0243] An acquisition module 1201 is configured to acquire image features of a first target image, image features of a second target image, and a mask image of the second target image, wherein the first image and the second image respectively include target objects, and a foreground of the mask image corresponds to the target object;

[0244] A first processing module 1202 is configured to process the image features of the second target image and the mask image using an image segmentation model to obtain a first target attention feature, wherein the image segmentation model is obtained using the training method in any of the above embodiments, and the first target attention feature is a fusion of the image features of the second image and the features of the mask image;

[0245] The second processing module 1203 is used to obtain a mask image of the first target image by using an image segmentation model to obtain the image features of the first target image and the prototype features of the target object in the first target image, where the prototype features are partial features in the first target attention features, and the partial features correspond to the mask image of the second target image. The mask image of the first target image is used to indicate the target object segmented from the first target image.

[0246] In some embodiments, the second processing module 1203 is configured to:

[0247] The image features and prototype features of the first target image are processed by the attention module in the image segmentation model to obtain the second target attention features;

[0248] The second target attention feature is processed by the decoding module in the image segmentation model to obtain a mask image. The decoding module is used to obtain the mask image based on the input feature.

[0249] In some embodiments, the second processing module 1203 is configured to:

[0250] Processing the image features and prototype features of the first target image through the attention module to obtain cross-attention features corresponding to the image features of the first target image and cross-attention features corresponding to the prototype features, where any cross-attention feature is a feature that combines the image features of the first target image and the prototype features;

[0251] The two cross-attention features are processed separately to obtain two self-attention features corresponding to the two cross-attention features;

[0252] The two self-attention features are processed to obtain two cross-attention features corresponding to the two self-attention features;

[0253] The process of iterating a preset number of times to obtain cross-attention features and self-attention features; wherein, the two self-attention features used in the i+1th iteration are the two self-attention features obtained in the i-th iteration, i is an integer greater than 0, and the second target attention feature is the self-attention feature corresponding to the image feature of the first target image obtained in the last iteration.

[0254] An embodiment of the present application provides an image segmentation device, which performs an attention mechanism calculation on the image features of the second target image and the mask image of the second target image to obtain an attention feature. In this way, in the process of calculating the attention mechanism for the two features, since the foreground of the mask image points to the target object in the second target image, the mask image mainly focuses on the features of the target object with high similarity to its foreground in the image features during the calculation process, and the features corresponding to the mask image in the obtained attention features can represent the target object in the second target image in more dimensions and more refined attribute features. Therefore, the features corresponding to the mask image in the attention features are used as the prototype features of the target object, so that the reference value of the prototype features is high, and then the image segmentation model refers to the prototype features to process the first target image, and the accuracy of the segmented object indicated by the obtained mask image is high.

[0255] In the embodiments of the present application, the computer device may be a terminal or a server. When the computer device is a terminal, the terminal serves as the execution subject to implement the technical solution provided in the embodiments of the present application; when the computer device is a server, the server serves as the execution subject to implement the technical solution provided in the embodiments of the present application; or, the technical solution provided in the present application may be implemented through interaction between the terminal and the server, which is not limited in the embodiments of the present application.

[0256] Figure 13 This is a structural diagram of a server provided in accordance with an embodiment of the present application. The server 1300 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 1301 and one or more memories 1302, wherein the memory 1302 is used to store executable program code, and the processor 1301 is configured to execute the above executable program code to implement the image segmentation model training method or image segmentation method provided in the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The server may also include other components for implementing device functions, which will not be described in detail here.

[0257] An embodiment of the present application also provides a computer-readable storage medium, in which at least one program is stored, and the at least one program is loaded and executed by a processor to implement the image segmentation model training method or image segmentation method of any of the above-mentioned implementation methods.

[0258] An embodiment of the present application also provides a computer program product, which includes at least one program, and the at least one program is stored in a computer-readable storage medium. The processor of the computer device reads the at least one program from the computer-readable storage medium, and the processor executes the at least one program, so that the computer device executes the image segmentation model training method or image segmentation method of any of the above-mentioned implementation methods.

[0259] In some embodiments, the computer program product involved in the embodiments of the present application can be deployed and executed on a computer device, or on multiple computer devices located at one location, or on multiple computer devices distributed at multiple locations and interconnected through a communication network. Multiple computer devices distributed at multiple locations and interconnected through a communication network can constitute a blockchain system.

[0260] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here. The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A training method for an image segmentation model, characterized in that: The method comprises: Acquire multiple sample groups, each sample group including image features and a mask image of a first image, and image features and a mask image of a second image, wherein the first image and the second image respectively include a target object, and foregrounds of the mask image of the first image and the mask image of the second image respectively correspond to the target object; For each sample group, processing the image features of the second image and the mask image in the sample group using an image segmentation model to obtain a first attention feature, wherein the image segmentation model is used to segment the target object from the first image, and the first attention feature is a fusion of the image features of the second image and the features of the mask image; processing, by the image segmentation model, image features of the first image and prototype features of the target object in the second image to obtain a third mask image, wherein the prototype features are partial features of the first attention features, the partial features correspond to the mask image of the second image, and the third mask image is used to indicate the object segmented from the first image; The image segmentation model is iteratively trained based on the third mask image and the first mask image of each of the multiple sample groups.

2. The method according to claim 1, characterized in that The step of processing the image features of the first image and the prototype features of the target object in the second image by the image segmentation model to obtain a third mask image includes: Processing the image features of the first image and the prototype features through an attention module in the image segmentation model to obtain a second attention feature, where the second attention feature is a feature that combines the prototype features and the image features of the first image; The second attention feature is processed by a decoding module in the image segmentation model to obtain the third mask image, and the decoding module is used to obtain the mask image based on the input feature.

3. The method according to claim 2, characterized in that The processing of the image features and the prototype features of the first image by the attention module in the image segmentation model to obtain a second attention feature includes: Processing the image features of the first image and the prototype features by the attention module to obtain cross-attention features corresponding to the image features of the first image and cross-attention features corresponding to the prototype features, where any cross-attention feature is a feature that combines the image features of the first image and the prototype features; Processing the two cross-attention features respectively to obtain two self-attention features corresponding to the two cross-attention features; Processing the two self-attention features to obtain two cross-attention features corresponding to the two self-attention features respectively; The process of iterating a preset number of times to obtain cross-attention features and self-attention features; wherein, the two self-attention features used in the i+1th iteration are the two self-attention features obtained in the i-th iteration, i is an integer greater than 0, and the second attention feature is the self-attention feature corresponding to the image feature of the first image obtained in the last iteration.

4. The method according to claim 1, wherein Each sample group includes two image features of the first image and two image features of the second image, and the method further includes: Input the first image into a convolution module, the outputs of the last two convolution layers of the convolution module are two image features of the first image, and the convolution module is used to extract features of the input image; input the second image into the convolution module, the outputs of the last two convolution layers are two image features of the second image; The iteratively training the image segmentation model based on the third mask image and the first mask image of each of the multiple sample groups includes: For each sample group, the model parameters of the image segmentation model are adjusted based on the average of the loss values ​​between the third mask images corresponding to each of the last two convolutional layers and the first mask image, where each third mask image is obtained based on the image features of the first image and the image features of the second image of the same convolutional layer.

5. The method according to claim 1, wherein The method further comprises: determining an image mask feature corresponding to the mask image of the first image, where the image mask feature is used to represent the mask image of the first image, and the image mask feature includes block features of a plurality of blocks in the mask image of the first image; adding a first noise to at least one block feature and adding a second noise to the remaining block features of the plurality of block features to obtain a noise image feature, wherein a variance of the second noise is greater than a variance of the first noise; Processing the image features of the first image, the prototype features, and the noise image features through an attention module in the image segmentation model to obtain a third attention feature; Processing the third attention feature through a decoding module in the image segmentation model to obtain a fourth mask image, where the fourth mask image is used to indicate the image block to which the first noise is added and the image block to which the second noise is added in the mask image of the first image, wherein the decoding module is used to obtain the mask image based on the input feature; determining a loss value between the fourth mask image and a fifth mask image, where the fifth mask image is a mask image corresponding to features of the noise image, and a foreground of the fifth mask image corresponds to the block to which the first noise is added; The iteratively training the image segmentation model based on the third mask image and the first mask image of each of the multiple sample groups includes: For each sample group, model parameters of the image segmentation model are adjusted based on a loss value between the third mask image and the first mask image and a loss value between the fourth mask image and the fifth mask image.

6. The method according to claim 5, characterized in that The processing of the image features of the first image, the prototype features, and the noise image features by the attention module in the image segmentation model to obtain a third attention feature includes: Processing the image features of the first image, the prototype features, and the noise image features through the attention module to obtain cross-attention features corresponding to the image features of the first image, cross-attention features corresponding to the prototype features, and cross-attention features corresponding to the noise image features, wherein any cross-attention feature is a feature that integrates at least two of the image features of the first image, the prototype features, and the noise image features; Processing the three cross-attention features respectively to obtain three self-attention features corresponding to the three cross-attention features; Processing the three self-attention features to obtain three cross-attention features corresponding to the three self-attention features respectively; The process of iterating a preset number of times to obtain cross-attention features and self-attention features; wherein, the three self-attention features used in the j+1th iteration are the three self-attention features obtained in the jth iteration, j is an integer greater than 0, and the third attention feature includes the self-attention feature corresponding to the image feature of the first image and the self-attention feature corresponding to the noise image feature obtained in the last iteration.

7. An image segmentation method, characterized in that: The method comprises: Acquire image features of a first target image, image features of a second target image, and a mask image of the second target image, wherein the first target image and the second target image respectively include a target object, and a foreground of the mask image corresponds to the target object; Processing the image features of the second target image and the mask image through an image segmentation model to obtain a first target attention feature, wherein the image segmentation model is obtained by the training method of any one of claims 1 to 6, and the first target attention feature is a fusion of the image features of the second image and the features of the mask image; The image features of the first target image and the prototype features of the target object in the first target image are processed through the image segmentation model to obtain a mask image of the first target image, wherein the prototype features are partial features in the first target attention features, and the partial features correspond to the mask image of the second target image, and the mask image of the first target image is used to indicate the target object segmented from the first target image.

8. The image segmentation method according to claim 7, characterized in that: The step of processing the image features of the first target image and the prototype features of the target object in the first target image by the image segmentation model to obtain a mask image of the first target image includes: Processing the image features of the first target image and the prototype features through the attention module in the image segmentation model to obtain a second target attention feature, where the second target attention feature is a feature that is a fusion of the prototype features and the image features of the first target image; The second target attention feature is processed by a decoding module in the image segmentation model to obtain the mask image, and the decoding module is used to obtain the mask image based on the input feature.

9. The image segmentation method according to claim 8, characterized in that: The processing of the image features and the prototype features of the first target image by the attention module in the image segmentation model to obtain a second target attention feature includes: Processing the image features of the first target image and the prototype features by the attention module to obtain cross-attention features corresponding to the image features of the first target image and cross-attention features corresponding to the prototype features, wherein any cross-attention feature is a feature that combines the image features of the first target image and the prototype features; Processing the two cross-attention features respectively to obtain two self-attention features corresponding to the two cross-attention features; Processing the two self-attention features to obtain two cross-attention features corresponding to the two self-attention features respectively; The process of iterating a preset number of times to obtain cross-attention features and self-attention features; wherein, the two self-attention features used in the i+1th iteration are the two self-attention features obtained in the i-th iteration, i is an integer greater than 0, and the second target attention feature is the self-attention feature corresponding to the image feature of the first target image obtained in the last iteration.

10. A training device for an image segmentation model, characterized in that: The device comprises: an acquisition module, configured to acquire a plurality of sample groups, each sample group including image features and a mask image of a first image and image features and a mask image of a second image, the first image and the second image respectively including a target object, and foregrounds of the mask image of the first image and the mask image of the second image respectively corresponding to the target object; a first processing module configured to process, for each sample group, the image features of the second image and the mask image in the sample group using an image segmentation model to obtain a first attention feature, wherein the image segmentation model is used to segment the target object from the first image, and the first attention feature is a fusion of the image features of the second image and the mask image; a second processing module, configured to process the image features of the first image and the prototype features of the target object in the second image using the image segmentation model to obtain a third mask image, wherein the prototype features are partial features of the first attention features, the partial features correspond to the mask image of the second image, and the third mask image is used to indicate the object segmented from the first image; The training module is configured to iteratively train the image segmentation model based on the third mask image and the first mask image of each of the multiple sample groups.

11. An image segmentation device, characterized in that: The device comprises: an acquisition module, configured to acquire image features of a first target image, image features of a second target image, and a mask image of the second target image, wherein the first image and the second image respectively include a target object, and a foreground of the mask image corresponds to the target object; a first processing module, configured to process the image features of the second target image and the mask image using an image segmentation model to obtain a first target attention feature, wherein the image segmentation model is obtained using the training method of any one of claims 1 to 6, and the first target attention feature is a fusion of the image features of the second image and the features of the mask image; The second processing module is used to process the image features of the first target image and the prototype features of the target object in the first target image through the image segmentation model to obtain a mask image of the first target image, where the prototype features are partial features in the first target attention features, and the partial features correspond to the mask image of the second target image, and the mask image of the first target image is used to indicate the target object segmented from the first target image.

12. A computer device, characterized in that: The computer device includes a processor and a memory, the memory is used to store at least one program, and the at least one program is loaded by the processor and executes the training method of the image segmentation model described in any one of claims 1 to 6 or the image segmentation method described in any one of claims 7-9.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store at least one program, and the at least one program is used to execute the training method of the image segmentation model described in any one of claims 1 to 6 or the image segmentation method described in any one of claims 7-9.

14. A computer program product, characterized in that The computer program product includes at least one program segment, which is stored in a computer-readable storage medium. The processor of the computer device reads the at least one program segment from the computer-readable storage medium, and the processor executes the at least one program segment, so that the computer device executes the image segmentation model training method described in any one of claims 1 to 6 or the image segmentation method described in any one of claims 7 to 9.