Image segmentation method and device, electronic equipment and storage medium
By modeling interactive segmentation into a binary conduction process and using diffusion partial differential equations, the problems of low accuracy and high computational cost in the prior art are solved, and more accurate and efficient image segmentation is achieved.
Patent Information
- Application Number
- CN202410095885.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-23
- Publication Date
- 2025-07-25
AI Technical Summary
The existing deep learning-based image segmentation method has the problem of low accuracy of image segmentation results. Especially in interactive segmentation tasks, the range of click information propagation is limited and the lack of interpretability is caused, resulting in high computing costs.
The interactive segmentation is modeled as a binary conduction process, and the diffusion partial differential equation is mathematically modeled, and the diffusion partial differential equation is constructed using the interactive infographic marked with the click position, and the solution is optimized to obtain the target category feature map, so as to realize image segmentation.
It improves the accuracy and efficiency of image segmentation, reduces additional computing costs, and effectively embeds click information in the network, which has good interpretability.
Smart Images

Figure CN120374966A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology. Specifically, this application relates to a method, apparatus, electronic device, and storage medium for image segmentation. Background Art
[0002] The interactive segmentation task aims to extract the region of the target object marked by the user from an image. The common type of user interaction is click interaction, where the user clicks positive points in the target region and negative points in the background region, with the expectation of achieving effective segmentation of the target object with the least number of clicks.
[0003] Currently, the vast majority of deep learning-based methods adopt existing semantic segmentation model frameworks. These frameworks generally include two parts: an encoder and a decoder. The encoder usually uses an existing semantic segmentation main network as a feature extractor. However, there are often problems with low accuracy of image segmentation results. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide a method, apparatus, electronic device, and storage medium for image segmentation that can improve the accuracy of image segmentation results. To achieve the above purpose, the technical solutions provided by the embodiments of this application are as follows:
[0005] In a first aspect, a method for image segmentation is provided, including:
[0006] Obtain a target image to be segmented and an interactive information map with click position identifiers generated based on click operations on the target image. The click operations include at least one of a positive click operation on a target pixel point or a negative click operation on a non-target pixel point;
[0007] After splicing the target image and the target interactive information map, perform feature extraction to obtain an initial category feature map and an initial image feature map. The feature value of each category feature point in the initial category feature map is used to represent the probability that the pixel point corresponding to the category feature point belongs to the target pixel point;
[0008] Use the category feature values in the category feature map as the concentration of the diffusing substance, and the gradient of the image feature values in the image feature map as the influencing factor of the diffusion coefficient. Construct a diffusion partial differential equation, where the category feature values in the initial category feature map are used as the initial concentration of the diffusing substance, and the gradient of the image feature values in the initial image feature map is used as the initial influencing factor of the diffusion coefficient. The diffusion coefficient is negatively correlated with the gradient;
[0009] Based on the Dirichlet boundary condition, the diffusion partial differential equation is optimized and solved to obtain a target category feature map, and image segmentation processing is performed based on the target category feature map to obtain the image segmentation result of the target image, where the condition includes that the eigenvalue of the category feature point corresponding to a positive click is constantly 1, and that of a negative click is constantly 0.
[0010] In a possible implementation manner, the optimizing and solving the diffusion partial differential equation based on the Dirichlet boundary condition to obtain a target category feature map includes:
[0011] Taking the initial category feature map and the initial image feature map as the first category feature map and the first image feature map of the first operation for the first time, repeating the first operation until a preset condition is satisfied, and taking the second category feature map obtained from the last first operation as the target category feature map;
[0012] Wherein, the first operation includes the following steps:
[0013] For each category eigenvalue in the first category feature map, determine the category eigenvalue difference between this category eigenvalue and the neighborhood eigenvalues of this category eigenvalue;
[0014] Determine the category eigenvalue difference between each category eigenvalue in the first category feature map and the neighborhood eigenvalues of this category eigenvalue, and determine the image eigenvalue difference between each image eigenvalue in the first image feature map and the neighborhood eigenvalues of this image eigenvalue;
[0015] According to the category eigenvalue differences and image eigenvalue differences corresponding to the respective category eigenvalues, determine the eigenvalue adjustment coefficient map corresponding to the first category feature map;
[0016] Adjust the first category feature map based on the eigenvalue adjustment coefficient map to obtain a second category feature map, and obtain a second image feature map by performing feature extraction on the second category feature map and the first image feature map, and take the second category feature map and the second image feature map as the first category feature map and the first image feature map of the next first operation.
[0017] In another possible implementation manner, the first operation further includes:
[0018] Respectively determine each first eigenvalue corresponding to the positive click operation and each second eigenvalue corresponding to the negative click operation in the first category feature map;
[0019] Perform feature extraction on the respective first eigenvalues to obtain the positive click feature corresponding to the positive click operation, and perform feature extraction on the respective second eigenvalues to obtain the negative click feature corresponding to the negative click operation;
[0020] Fuse the positive click feature and the negative click feature to obtain a click feature;
[0021] Based on the correlation between the click feature and the first category feature map, determine an adjustment weight map for the first category feature map;
[0022] Weight the first category feature map using the adjustment weight map to obtain a new category feature map;
[0023] Among them, determining the difference in category feature values between each category feature value in the first category feature map and the domain feature value of this category feature value includes:
[0024] For each category feature map in the new category feature map, determine the difference in category feature values between this category feature value and the neighborhood feature value of this category feature value.
[0025] In another possible implementation, the positive click feature and the negative click feature are obtained through the following method:
[0026] Take the first feature values and the second feature values as the to-be-processed feature values respectively, and perform the following operations on the processed feature values to obtain the click feature corresponding to the to-be-processed feature:
[0027] Extract features from the to-be-processed feature values;
[0028] Perform a linear mapping on the feature map obtained by feature extraction to obtain a corresponding mapping result, and collect the self-attention mechanism to determine the self-correlation between the feature values in the mapping result;
[0029] Weight the mapping result based on the self-correlation to obtain the click feature corresponding to the to-be-processed feature value. In another possible implementation, the adjustment of the first category feature map based on the eigenvalue adjustment coefficient of each category feature value to obtain the second category feature map includes:
[0030] Adjust the first category feature map based on the eigenvalue adjustment coefficient of each category feature value to obtain an adjusted first category feature map;
[0031] Based on the fusion process of the adjusted first category feature map and the first category feature map, obtain the second category feature map.
[0032] In another possible implementation, the obtaining of the second image feature map by performing feature extraction on the first image feature map and the second category feature map includes:
[0033] Merge the first image feature map and the second category feature map along the channel dimension to obtain a merged feature map;
[0034] Feature extraction is performed on the merged feature map to obtain a new image feature;
[0035] The new image feature is fused with the first image feature map to obtain the second image feature map.
[0036] In another possible implementation, the step of splicing the target image and the target interaction information map and then performing feature extraction to obtain the initial class feature map and the initial image feature map of the target image includes:
[0037] The target image is downsampled to obtain a first downsampling result of the target image;
[0038] The target interaction information map is downsampled to obtain a second downsampling result of the target interaction information map;
[0039] The first downsampling result and the second sampling result are spliced to obtain a spliced feature map;
[0040] Feature extraction is performed on the spliced feature map to obtain the initial class feature map and the initial image feature map of the target image.
[0041] In another possible implementation, the step of performing feature extraction on the spliced feature map to obtain the initial class feature map and the initial image feature map of the target image includes:
[0042] Feature extraction is performed on the spliced feature map to obtain the initial class feature map of the target image;
[0043] The initial class feature map is convolved, and the convolved class feature map and the initial class feature map are processed through a non-linear mapping to obtain the initial image feature map.
[0044] In another possible implementation, the step of splicing the target image and the target interaction information map and then performing feature extraction to obtain the initial class feature map and the initial image feature map of the target image includes:
[0045] Obtain a historical segmentation map obtained by historical segmentation of the target image;
[0046] The target interaction information map and the historical segmentation map are spliced to obtain a spliced image;
[0047] The target image and the spliced image are spliced and then feature extraction is performed to obtain the initial class feature map and the initial image feature map of the target image.
[0048] In a second aspect, an apparatus for image segmentation is provided. The apparatus includes:
[0049] An acquisition module, configured to acquire a target image to be segmented and an interactive information map with click position identifiers generated based on a click operation on the target image, where the click operation includes at least one of a positive click operation on a target pixel point or a negative click operation on a non-target pixel point;
[0050] A feature extraction module, configured to splice the target image and the target interactive information map and then perform feature extraction to obtain an initial category feature map and an initial image feature map, where the feature value of each category feature point in the initial category feature map is used to represent the probability that the pixel point corresponding to the category feature point belongs to the target pixel point;
[0051] A construction module, configured to use the category feature value in the category feature map as the concentration of the diffusing substance and the gradient of the image feature value in the image feature map as the influencing factor of the diffusion coefficient to construct a diffusion partial differential equation, where the category feature value in the initial category feature map is used as the initial concentration of the diffusing substance, and the gradient of the image feature value in the initial image feature map is used as the initial influencing factor of the diffusion coefficient, and the diffusion coefficient is negatively correlated with the gradient;
[0052] An optimization and solution module, configured to optimize and solve the diffusion partial differential equation based on the Dirichlet boundary condition to obtain a target category feature map, and perform image segmentation processing based on the target category feature map to obtain an image segmentation result of the target image, where the condition includes that the feature value of the category feature point corresponding to the positive click is constantly 1 and the negative click is constantly 0.
[0053] In a possible implementation manner, when the optimization and solution module optimizes and solves the diffusion partial differential equation based on the Dirichlet boundary condition to obtain a target category feature map, it is specifically configured to:
[0054] Use the initial category feature map and the initial image feature map as the first category feature map and the first image feature map of the first first operation, repeat the first operation until a preset condition is met, and use the second category feature map obtained by the last first operation as the target category feature map; where the first operation includes the following steps:
[0055] Determine the category feature value difference between each category feature value in the first category feature map and the domain feature value of the category feature value, and determine the image feature value difference between each image feature value in the first image feature map and the domain feature value of the image feature value;
[0056] Determine the eigenvalue adjustment coefficient map corresponding to the first category feature map according to the category feature value differences and image feature value differences corresponding to each of the category feature values;
[0057] Adjust the first category feature map based on the eigenvalue adjustment coefficient map to obtain a second category feature map, and perform feature extraction on the second category feature map and the first image feature map to obtain a second image feature map, and use the second category feature map and the second image feature map as the first category feature map and the first image feature map for the next first operation.
[0058] In a possible implementation manner, when the optimization solving module optimizes and solves the diffusion partial differential equation based on the Dirichlet boundary condition to obtain a target category feature map and performs image segmentation processing based on the target category feature map to obtain the image segmentation result of the target image, it is specifically used for:
[0059] Use the initial category feature map and the initial image feature map as the first category feature map and the first image feature map for the first first operation, repeat the first operation until a preset condition is met, and perform image segmentation processing based on the second category feature map obtained from the last first operation to obtain the image segmentation result of the target image;
[0060] Wherein, the first operation includes the following steps:
[0061] For each category eigenvalue in the first category feature map, determine the category eigenvalue difference between the category eigenvalue and the neighborhood eigenvalues of the category eigenvalue;
[0062] For each image eigenvalue in the first image feature map, determine the image eigenvalue difference between the image eigenvalue and the neighborhood eigenvalues of the image eigenvalue;
[0063] For each category eigenvalue, determine the eigenvalue adjustment coefficient of the category eigenvalue based on the category eigenvalue difference corresponding to the category eigenvalue and the image eigenvalue difference corresponding to the image eigenvalue corresponding to the category eigenvalue;
[0064] Adjust the first category feature map based on the eigenvalue adjustment coefficients of each category eigenvalue to obtain a second category feature map, and perform feature extraction on the first image feature map and the second category feature map to obtain a second image feature map, and use the second category feature map and the second image feature map as the first category feature map and the first image feature map for the next first operation.
[0065] In another possible implementation manner, when performing the first operation, the optimization solving module is further used for:
[0066] Determine each first eigenvalue corresponding to the positive click operation and each second eigenvalue corresponding to the negative click operation in the first category feature map respectively;
[0067] Perform feature extraction on each of the first eigenvalues to obtain a positive click feature corresponding to the positive click operation, and perform feature extraction on each of the second eigenvalues to obtain a negative click feature corresponding to the negative click operation;
[0068] Fuse the positive click feature and the negative click feature to obtain a click feature;
[0069] Determine an adjustment weight map of the first category feature map based on the correlation between the click feature and the first category feature map;
[0070] Weight the first category feature map using the adjustment weight map to obtain a new category feature map;
[0071] Wherein, when the optimization solving module determines the difference between each category eigenvalue in the first category feature map and the neighborhood eigenvalue of this category eigenvalue, it specifically is used for:
[0072] For each category feature map in the new category feature map, determine the difference between this category eigenvalue and the neighborhood eigenvalue of this category eigenvalue.
[0073] In another possible implementation manner, the positive click feature and the negative click feature are obtained through the following method:
[0074] Take each of the first eigenvalues and each of the second eigenvalues as the to-be-processed eigenvalues respectively, and perform the following operations on the processed eigenvalues to obtain the click feature corresponding to the to-be-processed feature:
[0075] Perform feature extraction on the to-be-processed eigenvalue;
[0076] Perform a linear mapping on the feature map obtained by feature extraction to obtain a corresponding mapping result, and collect a self-attention mechanism to determine the self-correlation between the eigenvalues in the mapping result;
[0077] Weight the mapping result based on the self-correlation to obtain the click feature corresponding to the to-be-processed eigenvalue.
[0078] In another possible implementation manner, when the optimization solving module performs feature extraction on each of the first eigenvalues to obtain a positive click feature corresponding to the positive click operation, and performs feature extraction on each of the second eigenvalues to obtain a negative click feature corresponding to the negative click operation, it specifically is used for:
[0079] Feature extraction is performed on each of the first eigenvalues to obtain a feature map corresponding to a positive click operation, and feature extraction is performed on each of the second eigenvalues to obtain a feature map corresponding to a negative click operation;
[0080] Perform a linear mapping on the feature map corresponding to the positive click operation to obtain a mapping result corresponding to the positive click operation, and collect a self-attention mechanism to determine the self-correlation between eigenvalues in the mapping result corresponding to the positive click operation;
[0081] Weight the mapping result corresponding to the positive click operation based on the self-correlation to obtain the positive click feature corresponding to the positive click operation; and,
[0082] Perform a linear mapping on the feature map corresponding to the negative click operation to obtain a mapping result corresponding to the negative click operation, and collect a self-attention mechanism to determine the self-correlation between eigenvalues in the mapping result corresponding to the negative click operation;
[0083] Weight the mapping result corresponding to the negative click operation based on the self-correlation to obtain the negative click feature corresponding to the negative click operation.
[0084] In another possible implementation, when the optimization solving module adjusts the first category feature map based on the eigenvalue adjustment coefficient map to obtain the second category feature map, it is specifically used for:
[0085] Adjust the first category feature map based on the eigenvalue adjustment coefficient map to obtain an adjusted first category feature map;
[0086] Perform a fusion process on the adjusted first category feature map and the first category feature map to obtain the second category feature map.
[0087] In another possible implementation, when the optimization solving module obtains the second image feature map by performing feature extraction on the first image feature map and the second category feature map, it is specifically used for:
[0088] Merge the first image feature map and the second category feature map along the channel dimension to obtain a merged feature map;
[0089] Perform feature extraction on the merged feature map to obtain a new image feature;
[0090] Perform a fusion process on the new image feature and the first image feature map to obtain the second image feature map.
[0091] In another possible implementation, when the feature extraction module performs feature extraction after splicing the target image and the target interaction information map to obtain the initial category feature map and the initial image feature map of the target image, it specifically is used for:
[0092] Perform downsampling processing on the target image to obtain a first downsampling result of the target image;
[0093] Perform downsampling processing on the target interaction information map to obtain a second downsampling result of the target interaction information map;
[0094] Perform splicing processing on the first downsampling result and the second sampling result to obtain a feature map after splicing processing;
[0095] Perform feature extraction on the feature map after splicing processing to obtain the initial category feature map and the initial image feature map of the target image.
[0096] In another possible implementation, when the feature extraction module performs feature extraction on the feature map after splicing processing to obtain the initial category feature map and the initial image feature map of the target image, it specifically is used for:
[0097] Perform feature extraction on the feature map after splicing processing to obtain the initial category feature map of the target image;
[0098] Perform convolution processing on the initial category feature map, and perform non-linear mapping processing on the convolved category feature map and the initial category feature map to obtain the initial image feature map.
[0099] In another possible implementation, when the feature extraction module performs feature extraction after splicing the target image and the target interaction information map to obtain the initial category feature map and the initial image feature map of the target image, it specifically is used for:
[0100] Obtain a historical segmentation map obtained by historical segmentation of the target image;
[0101] Perform splicing processing on the target interaction information map and the historical segmentation map to obtain a spliced image;
[0102] Perform feature extraction after splicing the target image and the spliced image to obtain the initial category feature map and the initial image feature map of the target image.
[0103] In a third aspect, an embodiment of the present application further provides an electronic device, which includes a memory and a processor. A computer program is stored in the memory, and the processor executes the computer program to implement the method for image segmentation provided by any possible implementation manner of the first aspect.
[0104] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements the image segmentation method provided by any possible implementation manner of the first aspect.
[0105] In a fifth aspect, an embodiment of the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the image segmentation method provided by any possible implementation manner of the fifth aspect.
[0106] The beneficial effects brought by the technical solution provided by the embodiment of the present application are as follows:
[0107] An embodiment of the present application provides an image segmentation method, device, electronic device, and storage medium. In the embodiment of the present application, by obtaining a target image to be segmented and an interactive information map with click position identifiers generated based on a click operation on the target image, the click operation includes at least one of a positive click operation on a target pixel point or a negative click operation on a non-target pixel point, and further obtaining an initial category feature map and an initial image feature map, it is possible to use the category feature value in the category feature map as the concentration of the diffusing substance, and the gradient of the image feature value in the image feature map as the influencing factor of the diffusion coefficient, to construct a diffusion partial differential equation. Among them, the category feature value in the initial category feature map is used as the initial concentration of the diffusing substance, and the gradient of the image feature value in the initial image feature map is used as the initial influencing factor of the diffusion coefficient. The diffusion coefficient is negatively correlated with the gradient. Thus, by optimizing and solving the diffusion partial differential equation based on the Dirichlet boundary condition, a target category feature map can be obtained. And the condition includes that the feature value of the category feature point corresponding to the positive click is always 1, and the negative click is always 0. Then, based on the target category feature map, image segmentation is performed to obtain the corresponding segmentation result. That is to say, in the embodiment of the present application, the interactive segmentation model is modeled as a binary conduction process and modeled by a diffusion partial differential equation to embed the interaction information with the target object in the network, so as to achieve a more accurate segmentation result. BRIEF DESCRIPTION OF THE DRAWINGS
[0108] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for description in the embodiments of the present application.
[0109] Figure 1a It is a schematic diagram of an image segmentation in the related art;
[0110] Figure 1b It is another schematic diagram of an image segmentation in the related art;
[0111] Figure 1c Schematic diagram of a system for image segmentation in an embodiment of this application;
[0112] Figure 2 Schematic diagram of a process for image segmentation in an embodiment of this application;
[0113] Figure 3 Schematic diagram of a target interaction information graph in an embodiment of this application;
[0114] Figure 4 Schematic diagram of another process for image segmentation in an embodiment of this application;
[0115] Figure 5 Schematic diagram of a process for solving boundary conditions in an embodiment of this application;
[0116] Figure 6 Schematic diagram of the specific implementation of a process for image segmentation in an embodiment of this application;
[0117] Figure 7 Schematic diagram of a process for applying an image segmentation method in an embodiment of this application;
[0118] Figure 8 Schematic diagram of the image segmentation achieved by applying the image segmentation method in an embodiment of this application;
[0119] Figure 9a Schematic example flowchart of a method for image segmentation in an embodiment of this application;
[0120] Figure 9b Schematic example flowchart of a process for solving a partial differential equation to obtain a target category feature map in an embodiment of this application;
[0121] Figure 9c Schematic diagram of the overall framework of a method for image segmentation in an embodiment of this application;
[0122] Figure 10a Schematic diagram of a (t + 1)-th differential equation solving module in an embodiment of this application;
[0123] Figure 10b Schematic diagram of the segmentation effect of the method shown in an embodiment of this application and a comparative method;
[0124] Figure 11 Schematic diagram of the structure of a device for image segmentation in an embodiment of this application;
[0125] Figure 12 Schematic flowchart of an electronic device in an embodiment of this application. Detailed implementation manners
[0126] Embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions of the embodiments of the present application.
[0127] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components, and / or their combinations supported by the art of the present technology. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by the term. For example, "A and / or B" can be implemented as "A", or as "B", or as "A and B". When describing multiple (two or more) items, if the relationship between the multiple items is not clearly defined, the multiple items can refer to one, multiple, or all of the multiple items. For example, for the description of "parameter A includes A1, A2, A3", it can be implemented that parameter A includes A1 or A2 or A3, or it can also be implemented that parameter A includes at least two of the three items of parameter A1, A2, and A3.
[0128] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0129] Among them, computer vision technology (CV): Computer vision is a science that studies how to enable machines to "see". Further, it refers to machine vision that uses cameras and computers to replace human eyes for target recognition and measurement, and further performs graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition. Machine learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0130] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role. The solution provided by the embodiments of this application involves technologies such as computer vision and machine learning in artificial intelligence, and specifically relates to an image segmentation method.
[0131] In related technologies, existing semantic segmentation model frameworks are adopted. These frameworks generally include two parts: an encoder and a decoder. The encoder usually uses an existing semantic segmentation main network as a feature extractor. The network for semantic segmentation is suitable for multi-class segmentation, has no special design for interactive segmentation, and lacks the embedding of information about the clicked location in the network.
[0132] Specifically, the interactive segmentation tasks involved in the related technologies may include: Reviving iterative training with mask guidance for interactive segmentation (RITM) and FocusCut: Diving Into a Focus View in Interactive Segmentation, both of which are deep learning-based technologies. As Figure 1a shown, RITM first converts the target object click into a click map, which can also be called the target interaction information map. Its two channels represent the foreground and background classes respectively. The values inside a disk with a radius of 5 pixels centered on each foreground or background click are 1, and the rest are 0. RITM inputs the click map and the segmentation result Prev.mask (also called: historical segmentation map) output in the previous round into the main network for interactive segmentation after convolution with the image respectively. The main network generally uses ResNet or HRNet as the network framework.
[0133] As Figure 1b shown, FocusCut uses the global view and the local view to input the whole image and the area around the click (after magnification) into the shared network for segmentation respectively, and then patches (crop and paste) the fine result output by the local view to the segmentation result of the global view to output the final target segmentation result. Its main network structure is similar to that of RITM and uses ResNet as the network framework.
[0134] However, both RITM and FocusCut adopt a convolutional-based main network structure, the range of click information propagation is relatively limited, and there is a lack of interpretability. In addition, FocusCut needs to perform additional segmentation on the image near each click, which greatly consumes computing costs, especially when the number of clicks is large.
[0135] To solve the above technical problems, the image segmentation method provided in the embodiments of the present application can be applied to such as Figure 1cIn the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can be set separately, integrated on the server 104, placed on the cloud or other servers. The server 104 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smartphones, tablets, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc.
[0136] In the embodiment of the present application, the embodiment of the present application provides a method for image segmentation, innovatively modeling interactive segmentation as a binary information conduction process, and mathematically modeling it with diffusion partial differential equations, and inspiring the design of a new network structure that can effectively embed click information in the network, achieving excellent segmentation effects; moreover, in the embodiment of the present application, there is no need to additionally input the model near the clicked area to increase the computing cost, which is relatively lightweight; furthermore, the network inspired by the mathematical model in the embodiment of the present application has good interpretability, and the click information can be introduced and well propagated in the network through the designed attention module.
[0137] Furthermore, the method for image segmentation provided in the embodiment of the present application can be integrated into a matte extraction software, enabling the target object to finely adjust the matte extraction result; this method can also be used for the dataset annotation task of the image segmentation task, and only by clicking can the annotation of the target object be efficiently completed, which can collect high-quality training data annotations with higher efficiency for tasks such as medical images that are difficult to obtain high-quality annotations.
[0138] The embodiment of the present application introduces in detail a method for image segmentation in combination with the accompanying drawings. Executed by an electronic device, in the embodiment of the present application, the electronic device can be a terminal device or a server, such as Figure 2 shown, the method may include:
[0139] Step S201, obtain a target image to be segmented and an interactive information graph with click position identifiers generated based on click operations on the target image.
[0140] For the embodiments of the present application, the target image to be segmented can be obtained first, and then the interactive information map with click position identifiers generated based on the click operation on the target image can be obtained. Alternatively, the interactive information map with click position identifiers generated based on the click operation on the target image can be obtained first, and then the target image to be segmented can be obtained. Of course, the target image to be segmented and the interactive information map with click position identifiers generated based on the click operation on the target image can also be obtained simultaneously. In a possible implementation manner, the target object performs a click interaction operation on the target image to be segmented, so as to perform image segmentation on the target image to be segmented. That is, at this time, the target image to be segmented and the interactive information map can be split from the target image to be segmented where the click interaction operation is performed.
[0141] Among them, the click operation includes at least one of a positive click operation on a target pixel point or a negative click operation on a non-target pixel point. In the embodiments of the present application, in order to segment the foreground image and the background image in the target image to be segmented, the positive click operation on the target pixel point can be a click operation in the foreground image, and the negative click operation on the non-target pixel point can be a click operation on the background image.
[0142] Specifically, the interactive information map records the positions of the pixels affected by the interaction performed on the target image. The interaction can be in various forms such as click, scribble, box, etc. In the embodiments of the present application, the interaction in the form of click is taken as an example for introduction, but it is not a limitation on the embodiments of the present application.
[0143] The interaction performed on the target image can be an actual interaction operation or a simulated interaction on the target image. The actual interaction operation refers to the operation during the actual interaction process of the user on the target image, and the simulated interaction is used to simulate the user's interaction operation. The pixels affected by the interaction performed on the target image can be one or more, and multiple means at least two. The positions of the pixels affected by the interaction performed on the target image can be in at least one of the foreground area or the background area. The foreground area is the position where the target object to be segmented in the target image is located. When the position of the pixel affected by the interaction performed on the target image is in the foreground area, the interaction is a positive interaction; when the position of the pixel affected by the interaction performed on the target image is in the background area, the interaction is a negative interaction. In a specific application, the target interactive information map can also record the interaction type of the pixel affected by the interaction, and the interaction type can indicate whether the interaction is a positive interaction or a negative interaction.
[0144] In some specific embodiments, the interaction with the target image is the user's click operation, and the target interaction information map can be generated through the following steps: Convert the click on the target object into a Click map. Its two channels respectively represent the foreground and background classes. Taking each foreground or background click as the center, the value inside a circular area with a radius of a preset number of pixels is 1, and the rest is 0. The preset number can be 5, for example. For illustration, as Figure 3 shown, assume the click position on the target image is as Figure 3 shown in Figure (a) of Figure 3 , where 1 and 2 are positions in the foreground area, and 3 and 4 are positions in the background area. Then, the target interaction information map shown in Figure (b) of
[0145] can be generated. The left figure is the foreground channel map, and the right figure is the background channel map.
[0146] Among them, the feature value of each category feature point in the initial category feature map is used to represent the probability that the pixel point corresponding to this category feature point belongs to the target pixel point. Of course, at this time, the feature value of the initial category feature point can also represent the probability that the pixel point corresponding to this category feature point belongs to a non-target pixel point.
[0147] Step S202: After splicing the target image and the target interaction information map, perform feature extraction to obtain the initial category feature map and the initial image feature map.
[0148] Among them, the category feature value in the initial category feature map is used as the initial concentration of the diffusing substance, and the gradient of the image feature value in the initial image feature map is used as the initial influencing factor of the diffusion coefficient. The diffusion coefficient is negatively correlated with the gradient.
[0149] Specifically, the diffusion partial differential equation is a class of partial differential equations that describe the propagation process of substances, heat, information, etc. in a medium. A typical diffusion equation is the heat conduction equation, which describes the diffusion process of heat and substances. The intuitive meaning of the diffusion partial differential equation is: In a given medium, substances, heat, or information, etc., spread from high-concentration regions to low-concentration regions. The rate of the diffusion process is proportional to the gradient of the concentration, and finally reaches a steady-state distribution. Therefore, the interactive segmentation problem can be regarded as a process of binary information conduction and thus modeled as a diffusion partial differential equation:
[0150]
[0151]
[0152] Among them, U(t; p) represents the probability that a point at coordinate p belongs to the positive class at time t; V(t; p) is an image feature map, which is the feature extracted from the image in this application; is the diffusion coefficient, and the function g(·) is a positive-valued decreasing function, usually set in the form of and div are the gradient and divergence operators in space, and can be approximated using the finite difference method. The partial differential equation has Dirichlet boundary conditions that the U value at positive class points is always 1, and the U value at negative class points is always 0. In the embodiments of the present application, positive class points represent that the pixel points corresponding to the category feature points belong to the target pixel points, and negative class points represent that the pixel points corresponding to the category feature points belong to non-target pixel points. represents the gradient of the image feature value in the initial image feature map; represents the gradient of the category feature value in the initial category feature map, represents the diffusion coefficient of the gradient of the category feature value in the initial category feature map; t represents time.
[0153] Further, it can be seen from equation (1) that if is larger, it means that it is at the boundary of the object in the image, and the information conduction rate is smaller, so the object boundary tends to divide the image into two sides with high / low positive class probabilities. We hope to solve equation (1) to obtain U in the steady state, that is, to obtain the probability that each position belongs to the positive / negative class.
[0154] Step S204: Based on the Dirichlet boundary conditions, optimize and solve the diffusion partial differential equation to obtain the target category feature map, and perform image segmentation processing based on the target category feature map to obtain the image segmentation result of the target image.
[0155] Among them, the Dirichlet boundary conditions include that the feature value of the category feature point corresponding to the positive click is always 1, and the negative click is always 0. In the embodiments of the present application, the calculation of the Dirichlet boundary conditions is detailed in the following embodiments.
[0156] Specifically, in the embodiments of the present application, the finite difference method is used to solve equation (1) to obtain the target category feature map. That is, after initializing the initial category feature map (U) and the initial image feature map (V), optimize and solve to obtain the target category feature map. In the embodiments of the present application, image segmentation is performed through the diffusion differential equation, which can avoid the need to input the model near the click again to increase the computing cost, and is relatively lightweight.
[0157] Specifically, based on the Dirichlet boundary condition, the diffusion partial differential equation is optimized and solved to obtain the target category feature map, which may specifically include: using the initial category feature map and the initial image feature map as the first category feature map and the first image feature map of the first operation for the first time, repeating the first operation until a preset condition is met, and using the second category feature map obtained from the last first operation as the target category feature map. Further, based on the second category feature map obtained from the last first operation, image segmentation processing is performed to obtain the image segmentation result of the target image.
[0158] For the embodiments of the present application, after obtaining the first category feature map and the first image feature map of the first operation for the first time, the differential equation can be solved T times (T first operations) to implement the process of image segmentation of the target image to be segmented, so as to obtain the image segmentation result of the target image.
[0159] It should be noted that T can be preset. For example, it can be 8 or determined based on a preset condition. In the embodiments of the present application, when it is detected that the U obtained multiple times t+1 tends to fit, that is, the finally obtained U T is used as the target segmentation result of the target image.
[0160] Among them, in order to increase the receptive field of the finite difference method, finite difference updates are performed on three resolutions of downsampling the initial image feature map and the initial category feature map by 2 times and 4 times, and the update method is performed according to the following first operation (step Sa, step Sb, step Sc, and step Sd), where, as Figure 4 shown, the first operation includes the following steps (step Sa, step Sb, step Sc, and step Sd), where,
[0161] Step Sa: Determine the category feature value difference between each category feature value in the first category feature map and the neighborhood feature value of the category feature value.
[0162] For the embodiments of the present application, for each category feature value in the first category feature map, determine the category feature value difference between the category feature value and the neighborhood feature value of the category feature value.
[0163] Specifically, determine the neighborhood of each category feature value in the first category feature map, which includes multiple feature points, and then determine the difference between the category feature value and the feature values of each feature point in its neighborhood.
[0164] Among them, N(m,n) is a neighborhood set of the pixel coordinates (m, n), which can usually be set as the 8 most adjacent pixels around; (m, n) is used to represent each category feature value.
[0165] Among them, Characterize the difference between the feature value of this category and the feature values of any feature points within its neighborhood. Characterize the feature value of this category. Characterize the feature value of any feature point within the neighborhood.
[0166] Step Sb: Determine the difference in image feature values between each image feature value in the first image feature map and the neighborhood feature values of this image feature value.
[0167] For the embodiments of this application, for each image feature value in the first image feature map, determine the difference in image feature values between this image feature value and the neighborhood feature values of this image feature value.
[0168] For the embodiments of this application, each image feature value in the first image feature map corresponds one-to-one with each category feature value in the first category feature map. In the embodiments of this application, after determining each image feature value in the first image feature map, determine the neighborhood feature values of each image feature value, and then determine, for each image feature value, the difference in image feature values between this image feature value and each of these neighborhood feature values.
[0169] Where N(m,n) is a neighborhood set of the pixel coordinates (m,n), which can usually be set as the 8 closest adjacent pixels around; (m,n) is used to characterize each category feature value.
[0170] Where Characterize the difference between this image feature value and the feature values of any feature points within its neighborhood. Characterize this image feature value. Characterize the feature value of any feature point within the neighborhood.
[0171] It should be noted that in the embodiments of this application, step Sa can be executed first and then step Sb, or step Sb can be executed first and then step Sa, or steps Sa and Sb can be executed simultaneously, and there is no limitation in the embodiments of this application.
[0172] Step Sc: Determine the feature value adjustment coefficient map corresponding to the first category feature map according to the category feature value differences and image feature value differences corresponding to each category feature value.
[0173] That is to say, for each category feature value, based on the category feature value difference corresponding to this category feature value and the image feature value difference corresponding to the image feature value corresponding to this category feature value, determine the feature value adjustment coefficient of this category feature value.
[0174] For the embodiments of the present application, for each category eigenvalue, determining the eigenvalue adjustment coefficient of the category eigenvalue is based on the category eigenvalue difference corresponding to the category eigenvalue, the difference between the image eigenvalues corresponding to the image eigenvalues of the category eigenvalue, and the category eigenvalue. In the embodiments of the present application, after determining the eigenvalue adjustment coefficient corresponding to each category eigenvalue, that is, determining the category adjustment coefficient corresponding to the first category feature map, the first category feature map is adjusted.
[0175] For example, the eigenvalue adjustment coefficient of the category eigenvalue is determined by formula (2), where
[0176]
[0177] where represents the difference between the category eigenvalue and the eigenvalue of any feature point in its neighborhood, represents the category eigenvalue, represents the eigenvalue of any feature point in the neighborhood; represents the difference between the image eigenvalue and the eigenvalue of any feature point in its neighborhood, represents the image eigenvalue, represents the eigenvalue of any feature point in the neighborhood, δ represents the eigenvalue adjustment coefficient of the category value, and g() represents the diffusion function.
[0178] Step Sd: Adjust the first category feature map based on the eigenvalue adjustment coefficient map to obtain the second category feature map, and perform feature extraction on the second category feature map and the first image feature map to obtain the second image feature map. The second category feature map and the second image feature map are used as the first category feature map and the first image feature map for the next first operation.
[0179] That is, in the embodiments of the present application, the first category feature map is adjusted based on the eigenvalue adjustment coefficient of each category eigenvalue to obtain the second category feature map, and feature extraction is performed on the first image feature map and the second category feature map to obtain the second image feature map. The second category feature map and the second image feature map are used as the first category feature map and the first image feature map for the next first operation.
[0180] Specifically, in the embodiments of the present application, after obtaining the first category feature and the first image feature map for the next first operation, steps Sa - Sd are continued until a preset condition is reached. That is, in the embodiments of the present application, in the solution process, steps Sa - Sd are looped until a preset condition is reached to finally obtain the target segmentation result of the target image.
[0181] Further, the first operation further includes: respectively determining each first eigenvalue corresponding to the positive click operation and each second eigenvalue corresponding to the negative click operation in the first category feature map; performing feature extraction on each first eigenvalue to obtain a positive click feature corresponding to the positive click operation, and performing feature extraction on each second eigenvalue to obtain a negative click feature corresponding to the negative click operation; fusing the positive click feature and the negative click feature to obtain a click feature; determining an adjustment weight map of the first category feature map based on the correlation between the click feature and the first category feature map; weighting the first category feature map with the adjustment weight map to obtain a new category feature map. In the embodiments of the present application, feature extraction may be first performed on each first eigenvalue to obtain a positive click feature corresponding to the positive click operation, and then feature extraction may be performed on each second eigenvalue to obtain a negative click feature corresponding to the negative click operation. Alternatively, feature extraction may be first performed on each second eigenvalue to obtain a negative click feature corresponding to the negative click operation, and then feature extraction may be performed on each first eigenvalue to obtain a positive click feature corresponding to the positive click operation. Alternatively, feature extraction may be simultaneously performed on each first eigenvalue to obtain a positive click feature corresponding to the positive click operation, and feature extraction may be performed on each second eigenvalue to obtain a negative click feature corresponding to the negative click operation.
[0182] It should be noted that the way to obtain the new category feature map is implemented by the following formula (3), where
[0183]
[0184] where, U t represents the first category feature map, represents the new category feature map. In the embodiments of the present application, the first category feature map (U) is made to satisfy the Dirichlet boundary condition through formula (3).
[0185] Specifically, in the embodiments of the present application, the positive click feature and the negative click feature are obtained in the following manner: taking each first eigenvalue and each second eigenvalue as the to-be-processed eigenvalues respectively, performing the following operations on the to-be-processed eigenvalues to obtain a click feature corresponding to the to-be-processed feature; performing feature extraction on the to-be-processed eigenvalues; performing a linear mapping on the feature map obtained by the feature extraction to obtain a corresponding mapping result, and collecting a self-attention mechanism to determine the self-correlation between the eigenvalues in the mapping result; weighting the mapping result based on the self-correlation to obtain a click feature corresponding to the to-be-processed eigenvalue.
[0186] That is to say, specifically, feature extraction is performed on each first eigenvalue to obtain the positive click features corresponding to the positive click operation, and feature extraction is performed on each second eigenvalue to obtain the negative click features corresponding to the negative click operation. Specifically, it may include: performing feature extraction on each first eigenvalue to obtain a feature map corresponding to the positive click operation, and performing feature extraction on each second eigenvalue to obtain a feature map corresponding to the negative click operation; performing a linear mapping on the feature map corresponding to the positive click operation to obtain a mapping result corresponding to the positive click operation, and collecting the self-attention mechanism to determine the self-correlation between the eigenvalues in the mapping result corresponding to the positive click operation; weighting the mapping result corresponding to the positive click operation based on the self-correlation to obtain the positive click features corresponding to the positive click operation; and performing a linear mapping on the feature map corresponding to the negative click operation to obtain a mapping result corresponding to the negative click operation, and collecting the self-attention mechanism to determine the self-correlation between the eigenvalues in the mapping result corresponding to the negative click operation; weighting the mapping result corresponding to the negative click operation based on the self-correlation to obtain the negative click features corresponding to the negative click operation.
[0187] Specifically, as Figure 5 shown, extract the features c t corresponding to positive and negative clicks from U t , and perform a linear mapping through a layer of MLP to obtain Then, for the positive and negative class points, respectively pass through their respective attention modules to obtain and where are all learnable parameters in the network, and are combined by and to obtain e t , perform a ClickAttn module on U t and e t to obtain This boundary condition module can make the pixels more similar to the positive (negative) class points in after attention have more positive class features by separately learning the k and v corresponding to the positive and negative class points, and promote that the positive (negative) class points can be expressed as linear combination of, and can be easily classified as the positive class. In the embodiments of the present application, the network designed inspired by the mathematical model has good interpretability, and click information can be introduced into the network through the designed attention module and propagated well.
[0188] Among them, represents the positive click features corresponding to the positive click operation and represents the negative click features corresponding to the negative click operation; e t represents the click features corresponding to the click operation.
[0189] Specifically, in the embodiments of the present application, the new category feature map is characterized by In order to meet the boundary conditions and improve the accuracy of image segmentation, based on the above embodiments, for each category feature value in the first category feature map, the category feature value difference between the category feature value and the neighborhood feature value of the category feature value is determined. Specifically, it may include: for each category feature map in the new category feature map, the category feature value difference between the category feature value and the neighborhood feature value of the category feature value is determined. That is, through The category feature value difference between the category feature value and the category feature value is determined.
[0190] After determining Based on The difference between each eigenvalue in the first image feature map and each eigenvalue in the neighborhood is determined, and then the new category feature map The eigenvalue adjustment coefficient corresponding to each category feature value in can be determined, that is, the eigenvalue adjustment coefficient corresponding to each category feature value is determined through the following formula (4).
[0191]
[0192] Wherein, Represents the eigenvalue of each feature point in the new category feature map, Represents the eigenvalue of each feature point in its neighborhood, and δ represents the eigenvalue adjustment coefficient for each eigenvalue in the new category feature map, or represents the adjustment coefficient for the new category feature map. Further, after obtaining the adjustment coefficients for each category feature value, the first category feature map is adjusted based on the eigenvalue adjustment coefficients of each category feature value to obtain the second category feature map. Specifically, it may include: the first category feature map is adjusted based on the eigenvalue adjustment coefficients of each category feature value to obtain the adjusted first category feature map; the adjusted first category feature map and the first category feature map are fused to obtain the second category feature map.
[0193] Further, after the above boundary processing, that is, after formulas (2)-(4), after obtaining the eigenvalue adjustment coefficient δ of each category feature value, the second category feature map is obtained based on formula (5), wherein,
[0194]
[0195] Wherein, U t+1 Represents the second category feature map, Represents the new category feature map.
[0196] Further, after obtaining the second category feature map, feature extraction is performed on the first image feature map and the second category feature map to obtain the second image feature map. Specifically, the second image feature map is obtained through formula (6), where formula (6) is as follows:
[0197] V t+1 = Update(V t ; U t+1 ), formula (6)
[0198] where V t+1 represents the second image feature map, V t represents the first image feature map, and U t+1 represents the second category feature map.
[0199] Specifically, by performing feature extraction on the first image feature map and the second category feature map, the second image feature map can be obtained, which specifically may include: merging the first image feature map and the second category feature map along the channel dimension to obtain a merged feature map; performing feature extraction on the merged feature map to obtain new image features; and performing a fusion process on the new image features and the first image feature map to obtain the second image feature map.
[0200] Specifically, V t and U t+1 are merged along the channel dimension, then passed through Conv3×3, Group Normalization (GroupNorm), Linear rectification function (ReLU), Conv3×3, GroupNorm, and the resulting result is added to V t , and a final V t+1 is obtained through a ReLU, that is, the second image feature map is obtained, where the channel dimension is set to 96 and the number of groups of GroupNorm is set to 1.
[0201] Further, by looping the solution method of the diffusion differential equation shown in the above embodiments until a preset condition is reached to obtain the final target category feature map. Further, after obtaining the final target category feature map, image segmentation processing is performed through a segmentation head (such as Conv1×1) to obtain the target segmentation result. In the process of solving the diffusion differential equation, the initial category feature map and the initial image feature map used are obtained by performing feature extraction based on the target image to be segmented and the target interaction information map. In order to further improve the accuracy of image segmentation, especially to improve the accuracy of the extracted initial category feature map and initial image feature map, the influence of the historical segmentation map of the target image on them can be considered when obtaining the initial category feature map and the initial image feature map.
[0202] Specifically, after splicing the target image and the target interaction information map in step S202, feature extraction is performed to obtain the initial category feature map and the initial image feature map of the target image, which may specifically include: obtaining the historical segmentation map obtained by performing historical segmentation on the target image; performing splicing processing on the target interaction information map and the historical segmentation map to obtain the spliced image; performing splicing on the target image and the spliced image and then performing feature extraction to obtain the initial category feature map and the initial image feature map of the target image.
[0203] Among them, the historical segmentation image is obtained by performing historical segmentation processing on the target image at least once. The interaction information map after channel superposition is the image obtained by superposing the historical segmentation image and the interaction information map in the channel direction.
[0204] Specifically, the electronic device can perform channel superposition on the historical segmentation image and the interaction information map to obtain the interaction information map after channel superposition, and then can perform downsampling on the interaction information map after channel superposition and the target image in the same way respectively to obtain two initial feature maps.
[0205] In the above embodiment, by performing channel superposition on the historical segmentation image and the interaction information map to obtain the interaction information map after channel superposition, the image segmentation model can perform image segmentation in combination with the historical segmentation image, further improving the accuracy of the image segmentation result.
[0206] Specifically, after splicing the target image and the target interaction information map and then performing feature extraction to obtain the initial category feature map and the initial image feature map of the target image, it may specifically include: performing downsampling processing on the target image to obtain the first downsampling result of the target image; performing downsampling processing on the target interaction information map to obtain the second downsampling result of the target interaction information map; performing splicing processing on the first downsampling result and the second sampling result to obtain the feature map after splicing processing; performing feature extraction on the feature map after splicing processing to obtain the initial category feature map and the initial image feature map of the target image. In the embodiments of the present application, the target image can be first downsampled to obtain the first downsampling result of the target image, and then the target interaction information map can be downsampled to obtain the second downsampling result of the target interaction information map. It is also possible to first downsample the target interaction information map to obtain the second downsampling result of the target interaction information map, and then downsample the target image to obtain the first downsampling result of the target image. It is also possible to simultaneously downsample the target image to obtain the first downsampling result of the target image, and, downsample the target interaction feature map to obtain the second downsampling result of the target interaction information map.
[0207] Furthermore, in the above embodiments, it is pointed out that the influence of the historical segmentation result of the target image to be segmented on the segmentation result of the target image can also be considered. Therefore, after considering splicing the target interaction information map and the historical segmentation map to obtain the spliced image, further, after performing splicing processing on the target interaction information map and the historical segmentation map to obtain the spliced image, the target image and the spliced image are spliced and then feature extraction is performed to obtain the initial category feature map and the initial image feature map of the target image, which specifically may include: performing downsampling processing on the target image to obtain the first downsampling result of the target image; performing downsampling processing on the spliced image to obtain the second downsampling result of the spliced image; performing splicing processing again on the first downsampling result and the second sampling result to obtain the feature map after the second splicing processing; performing feature extraction on the spliced feature map to obtain the initial category feature map and the initial image feature map of the target image.
[0208] For example, the downsampling processing is a convolution operation with a convolution kernel size of 4 and a stride of 4, and the number of channels is 96. That is, the target image is passed through a convolution operation with a convolution kernel size of 4 and a stride of 4 (the number of channels is 96) to obtain the first downsampling result of the target image, and the spliced image is passed through a convolution operation with a convolution kernel size of 4 and a stride of 4 (the number of channels is 96) to obtain the second downsampling result of the spliced image.
[0209] Specifically, performing feature extraction on the spliced feature map to obtain the initial category feature map and the initial image feature map of the target image may specifically include: performing feature extraction on the spliced feature map to obtain the initial category feature map of the target image; performing convolution processing on the initial category feature map, and performing non-linear mapping processing on the convolved category feature map and the initial category feature map to obtain the initial image feature map.
[0210] That is, performing feature extraction on the spliced feature map through the initial convolution block to obtain the initial category feature map of the target image, and then performing convolution processing on the initial category feature map, and performing non-linear mapping processing on the convolved category feature map and the initial category feature map to obtain the initial image feature map.
[0211] Specifically, after passing through an initial convolution block, it sequentially passes through Conv3×3, GroupNorm, ReLU, Conv3×3, GroupNorm, ReLU, Conv3×3 to obtain the initial category feature map U 0 , where the number of channels is set to 96 for all, and the number of groups of GroupNorm is set to 1 for all, V 0 is obtained through V 0 =ReLU(U 0 +Conv3x3(U 0)) Calculate to obtain the initial category feature map and the initial image feature map.
[0212] As can be seen from the above embodiments: In the above embodiments, when performing image segmentation on the image to be segmented, it is implemented based on an interactive segmentation network of diffusion partial differential equations. Before performing image segmentation using the interactive network based on diffusion partial differential equations, the network needs to be trained. The specific training process can also refer to Figure 6 As shown, initialize the model parameters, then load the sample images (training images to be segmented), perform preprocessing and data augmentation and other preprocessing on the sample images. For the sample images obtained from the preprocessing, generate several simulated clicks to obtain the click map (interactive information map), and then predict the segmentation result through the image segmentation model. After obtaining the predicted segmentation result, update the model parameters in the image segmentation model according to the training objective function and the Adam algorithm, and determine whether the preset number of iterations is reached. If the preset number of iterations is not reached, continue the iterative training. If the preset number of iterations is reached, the training ends, and the trained image segmentation model is saved. Specifically, in the embodiments of the present application, 8,498 images and 20,172 foreground annotations in the publicly available Semantic Boundaries Dataset (SBD) are used as training data. Among them, some images have multiple objects, that is, there are multiple foreground annotations; normalize the pixel value range of the training data to the threshold range [0, 1], and randomly crop each training image into image patches of size 256×256, and then perform random horizontal mirror flipping and random vertical mirror flipping with a probability of 0.3 respectively.
[0213] Specifically, at the beginning of training, the network parameters are randomly initialized with Gaussian distribution, where The initialization of is relatively special. First, randomly initialize Then take As the Initialization.
[0214] Among them, the training objective function is shown in formula (7):
[0215]
[0216] Where Is the final output of the network, y gt Is the true label of the training sample, Is the Normalized focal loss.
[0217] Furthermore, in the embodiments of the present application, the optimization network parameters are updated by using the Adaptive Moment Estimation (Adam) algorithm. In each iteration process, the prediction result error is calculated and backpropagated to the convolutional neural network model, and the gradient is calculated and the parameters of the convolutional neural network model are updated. The training shown in the embodiments of the present application is carried out end-to-end, and the initial learning rate is set to 5×10 -5 , for a total of 100 epochs. The learning rate decays by 10 times at the 80th and 90th epochs. In each iteration process, the batch size is 16, and the patch size is 256×256. When generating training samples, simulated clicks are required, and the simulation algorithm is implemented as follows: the next click is generated at the center of the maximum error region.
[0218] After the training is completed, the trained image segmentation model can also be tested. The specific processes of training and testing can be referred to Figure 6 . After the process starts, it is determined whether it is training or testing. If it is training, the model parameters are initialized, and then the sample images (training images to be segmented) are loaded. Preprocessing and data augmentation and other preprocessing are performed on the sample images. For the sample images obtained from the preprocessing, several simulated clicks are generated to obtain the click map (interaction information map), and then the segmentation result is predicted through the image segmentation model. After obtaining the predicted segmentation result, the model parameters in the image segmentation model are updated according to the training objective function and the Adam algorithm. It is determined whether the preset number of iterations is reached. If the preset number of iterations is not reached, the iterative training continues. If the preset number of iterations is reached, the training ends, and the trained image segmentation model is saved.
[0219] Continuing to refer to Figure 6 , if it is determined that it is testing, the trained model is loaded, and then the image to be tested is loaded. The target object clicks on the foreground class or the background class on the image. The image segmentation model outputs the current segmentation result according to the image to be tested, the target object click, and the previous segmentation result. It is determined whether the foreground target object is segmented. If so, the final segmentation result is obtained. If not, the user continues to click on the foreground class or the background class on the image, and the image segmentation model continues to perform image segmentation until the foreground target object is segmented. In the embodiments of the present application, the method of performing image segmentation can be as shown in the above embodiments.
[0220] Furthermore, in the embodiments of the present application, a specific example is used to introduce the image segmentation method shown in the embodiments of the present application, which is specifically as follows:
[0221] As Figure 7As shown, the target object can perform a click operation on the image to be segmented through the terminal device A, and then the terminal device A transmits the image to be segmented and the target object click to the server, so that the server segments the foreground part and the background part through the algorithm shown in the embodiments of the present application (the interactive segmentation network based on the diffusion partial differential equation) to obtain the segmentation result and returns the segmentation result to the terminal device A;
[0222] Specifically, as Figure 8 shown, Figure 8 in (a) of is the image after the target object performs a click operation on the image to be segmented, including the click points for clicking on the foreground part pixels and the operation points for clicking on the background part, and the segmentation result is as Figure 8 shown in (b) of.
[0223] Specifically, a method of image segmentation is described through an example. For specific details, see the following embodiments. As Figure 9a shown, where
[0224] Step S901: Obtain the target image to be segmented and the interactive information map with click position marks generated based on the click operation on the target image.
[0225] Among them, the click operation includes at least one of a positive click operation on the target pixel point or a negative click operation on the non-target pixel point. In the embodiments of the present application, in order to segment the foreground image and the background image in the target image to be segmented, the positive click operation on the target pixel point can be a click operation in the foreground image, and the negative click operation on the non-target pixel point can be a click operation on the background image.
[0226] Specifically, the interactive information map records the positions of the pixels affected by the interaction performed on the target image. The interaction can be in various forms such as click, scribble, box, etc. In the embodiments of the present application, the interaction is taken as an example of the click form for introduction, but it is not a limitation on the embodiments of the present application.
[0227] Step S902: Perform downsampling processing on the target image to obtain the first downsampling result of the target image.
[0228] For example, the downsampling processing is a convolution operation with a convolution kernel size of 4, a stride of 4, and 96 channels. That is, the target image is passed through a convolution operation with a convolution kernel size of 4, a stride of 4 (96 channels) to obtain the first downsampling result of the target image.
[0229] Step S903: Obtain the historical segmentation map obtained by historical segmentation of the target image.
[0230] Among them, the historical segmentation image is obtained by performing at least one historical segmentation process on the target image. Among them, the historical segmentation image can be, for example, Figure 1a the Prev.mask shown.
[0231] Step S904: Perform a splicing process on the interaction information map and the historical segmentation map to obtain a spliced image.
[0232] Specifically, the electronic device can perform channel superposition on the historical segmentation image and the interaction information map to obtain an interaction information map after channel superposition. Furthermore, the interaction information map after channel superposition and the target image can be downsampled in the same way to obtain two initial feature maps.
[0233] In the above embodiment, by performing channel superposition on the historical segmentation image and the interaction information map to obtain an interaction information map after channel superposition, the image segmentation model can perform image segmentation in combination with the historical segmentation image, further improving the accuracy of the image segmentation result.
[0234] Step S905: Perform a downsampling process on the spliced image to obtain a downsampling result corresponding to the spliced image; in the embodiment of the present application, the downsampling result corresponding to the spliced image is also the second downsampling result.
[0235] Specifically, the spliced image is passed through a convolution operation with a convolution kernel size of 4 and a stride of 4 (the number of channels is 96) to obtain the second downsampling result of the spliced image.
[0236] Step S906: Perform a splicing process on the first downsampling result and the downsampling result corresponding to the spliced image (the second downsampling result) to obtain a spliced feature map.
[0237] Specifically, in the embodiment of the present application, the method of performing the splicing process is not limited.
[0238] Step S907: Perform feature extraction on the spliced feature map to obtain an initial class feature map of the target image.
[0239] That is to say, the spliced feature map is subjected to feature extraction through an initial convolution block to obtain an initial class feature map of the target image.
[0240] Specifically, after passing through an initial convolution block, it passes through Conv3×3, GroupNorm, ReLU, Conv3×3, GroupNorm, ReLU, Conv3×3 in sequence to obtain the initial class feature map U 0 , where the number of channels is set to 96, and the number of groups of GroupNorm is set to 1.
[0241] Step S908: Perform convolution processing on the initial class feature map, and perform non-linear mapping processing on the convolved class feature map and the initial class feature map to obtain the initial image feature map.
[0242] Specifically, after obtaining the initial class feature map, then perform convolution processing on the initial class feature map, and perform non-linear mapping processing on the convolved class feature map and the initial class feature map to obtain the initial image feature map.
[0243] Further, the initial image feature map can also be characterized by V 0 where V 0 is calculated by V 0 = ReLU(U 0 + Conv3×3(U 0 ))
[0244] Step S909: Use the class feature values in the class feature map as the concentration of the diffusing substance, and use the gradient of the image feature values in the image feature map as the influencing factor of the diffusion coefficient to construct a diffusion partial differential equation.
[0245] Among them, the class feature values in the initial class feature map are used as the initial concentration of the diffusing substance, the gradient of the image feature values in the initial image feature map is used as the initial influencing factor of the diffusion coefficient, and the diffusion coefficient is negatively correlated with the gradient.
[0246] Specifically, the diffusion partial differential equation is a type of partial differential equation that describes the propagation process of substances, heat, information, etc. in a medium. A typical diffusion equation is the heat conduction equation, which describes the diffusion process of heat and substances. The intuitive meaning of the diffusion partial differential equation is that in a given medium, substances, heat, or information, etc. spread from high-concentration regions to low-concentration regions, and the rate of the diffusion process is proportional to the gradient of the concentration, reaching the final steady-state distribution. Therefore, the interactive segmentation problem can be regarded as a binary information conduction process, and thus modeled as a diffusion partial differential equation:
[0247]
[0248]
[0249] where U(t; p) represents the probability that the point at coordinate p belongs to the positive class at time t; V(t; p) is an image feature map, which in this application is the feature extracted from the image; is the diffusion coefficient, and the function g(·) is a positive-valued decreasing function, usually set in the form of The operators ∇ and div represent the gradient and divergence in space, and can be approximated using the finite difference method. This partial differential equation has Dirichlet boundary conditions, where the value of U at positive class points is always 1, and the value of U at negative class points is always 0. In the embodiments of the present application, positive class points indicate that the pixel points corresponding to the class feature points belong to the target pixel points, and negative class points indicate that the pixel points corresponding to the class feature points belong to non-target pixel points. represents the gradient of the image feature value in the initial image feature map; represents the gradient of the class feature value in the initial class feature map, represents the diffusion coefficient of the gradient of the class feature value in the initial class feature map; t represents time.
[0250] Furthermore, it can be seen from Equation (1) that if is larger, it means that it is at the boundary of the object in the image, and the information conduction rate is smaller. Thus, the object boundary tends to divide the image into two sides with high / low positive class probabilities. We hope to solve Equation (1) to obtain U in the steady state, that is, to obtain the probability of each position belonging to the positive / negative class.
[0251] Step S910: Based on the Dirichlet boundary conditions, optimize and solve the diffusion partial differential equation to obtain the target class feature map, where the Dirichlet boundary conditions include that the eigenvalue of the class feature point corresponding to the positive click is always 1, and the eigenvalue of the negative click is always 0.
[0252] Specifically, in the embodiments of the present application, the finite difference method is used to solve Equation (1) to obtain the target class feature map. That is, after initializing the initial class feature map (U) and the initial image feature map (V), optimize and solve to obtain the target class feature map.
[0253] Specifically, the method of using the finite difference method to solve Equation (1) can be as shown in the above Equations (3), (4), (5) and (6), which will not be elaborated here.
[0254] Furthermore, in the embodiments of the present application, image segmentation can be performed through the diffusion differential equation, which does not require additional input of the model near the click to increase the computational cost, and is relatively lightweight. In the embodiments of the present application, the specific method of optimizing and solving the diffusion partial differential equation is described in detail in the following embodiments, which will not be elaborated here.
[0255] Step S911: Segment the final target class feature map U T through the segmentation head (Conv1×1) to obtain the target segmentation result.
[0256] It should be noted that T can be set in advance. For example, it can be 8, or it can be determined based on preset conditions. In the embodiments of the present application, when it is detected that the obtained U is obtained multiple times t+1 tends to fit, that is, the finally obtained U T is used as the target segmentation result of the target image.
[0257] Specifically, in step S910, based on the Dirichlet boundary condition, the diffusion partial differential equation is optimized and solved, and the specific implementation method in the target category feature map is as follows Figure 9b shown, where
[0258] Step S9101: Respectively determine each first eigenvalue corresponding to the positive click operation and each second eigenvalue corresponding to the negative click operation in the first category feature map;
[0259] Specifically, in the embodiments of the present application, each first eigenvalue corresponding to the positive click operation in the first category feature map can be determined first, and then each second eigenvalue corresponding to the negative click operation can be determined. It is also possible to first determine each second eigenvalue corresponding to the negative click operation, and then determine each first eigenvalue corresponding to the positive click operation. It is also possible to obtain each first eigenvalue and each second eigenvalue at the same time, or in other orders, to determine each first eigenvalue and each second eigenvalue.
[0260] Step S9102: Perform feature extraction on each first eigenvalue to obtain a feature map corresponding to the positive click operation, and perform feature extraction on each second eigenvalue to obtain a feature map corresponding to the negative click operation.
[0261] Specifically, in the embodiments of the present application, feature extraction can be performed on each first eigenvalue first to obtain the positive click feature corresponding to the positive click operation, and then feature extraction can be performed on each second eigenvalue to obtain the negative click feature corresponding to the negative click operation. It is also possible to first perform feature extraction on each second eigenvalue to obtain the negative click feature corresponding to the negative click operation, and then perform feature extraction on each first eigenvalue to obtain the positive click feature corresponding to the positive click operation. It is also possible to perform feature extraction on each first eigenvalue at the same time to obtain the positive click feature corresponding to the positive click operation, and perform feature extraction on each second eigenvalue to obtain the negative click feature corresponding to the negative click operation.
[0262] As Figure 5 shown, extract the feature c t corresponding to the positive click from U t , and extract the feature c t corresponding to the negative click from U t .
[0263] Step S9103: Perform a linear mapping on the feature map corresponding to the positive click operation to obtain the mapping result corresponding to the positive click operation, collect the self-attention mechanism, and determine the self-correlation between the eigenvalues in the mapping result corresponding to the positive click operation; weight the mapping result corresponding to the positive click operation based on the self-correlation to obtain the positive click feature corresponding to the positive click operation.
[0264] Specifically, as Figure 5 shown, extract the feature c t corresponding to the positive click from U t , perform a linear mapping on it through a layer of MLP to obtain Then, for the positive class points, pass through the corresponding attention module to obtain where are all learnable parameters in the network. Among them, represents the positive click feature corresponding to the positive click operation.
[0265] Step S9104: Perform a linear mapping on the feature map corresponding to the negative click operation to obtain the mapping result corresponding to the negative click operation, collect the self-attention mechanism, and determine the self-correlation between the eigenvalues in the mapping result corresponding to the negative click operation; weight the mapping result corresponding to the negative click operation based on the self-correlation to obtain the negative click feature corresponding to the negative click operation.
[0266] Furthermore, continue as Figure 5 shown, extract the feature c t corresponding to the negative click from U t , perform a linear mapping on it through a layer of MLP to obtain Then, for the negative class points, pass through the corresponding attention module to obtain where are all learnable parameters in the network. Among them, represents the negative click feature corresponding to the negative click operation.
[0267] Step S9105: Fuse the positive click feature and the negative click feature to obtain the click feature.
[0268] Continue, as Figure 5 shown, combine and to obtain e t , e t represents the click feature corresponding to the click operation.
[0269] Step S9106: Based on the correlation between the click feature and the first category feature map, determine the adjustment weight map of the first category feature map; weight the first category feature map using the adjustment weight map to obtain the new category feature map.
[0270] Specifically, for the first category feature map (Ut ) and click feature (e t ) is processed by self-attention through a ClickAttn module to obtain
[0271] Step S9107: For each category feature value in the new category feature map, determine the category feature value difference between the category feature value and the neighborhood feature values of the category feature value.
[0272] For the embodiments of the present application, by Determine the category feature value difference between the category feature value and the category feature value itself. Among them, Represents the feature value of each feature point in the new category feature map, Represents the feature value of each feature point in its neighborhood.
[0273] Step S9108: For each image feature value in the first image feature map, determine the image feature value difference between the image feature value and the neighborhood feature values of the image feature value.
[0274] Specifically, after determining Based on Determine the difference between each feature value in the first image feature map and each feature value in the neighborhood. Among them, Represents the image feature value, Represents the feature value of any feature point in the neighborhood.
[0275] Step S9109: For each category feature value, based on the category feature value difference corresponding to the category feature value and the image feature value difference corresponding to the image feature value corresponding to the category feature value, determine the feature value adjustment coefficient of the category feature value.
[0276] In the embodiments of the present application, after obtaining the category feature value difference corresponding to the category feature value and the image feature value difference corresponding to the image feature value corresponding to the category feature value through the above steps S9107 and S9108, the feature value adjustment coefficient corresponding to each category feature value can be determined by the following formula (4).
[0277]
[0278] Among them, Represents the feature value of each feature point in the new category feature map, Represents the feature value of each feature point in its neighborhood, and δ represents the feature value adjustment coefficient for each feature value in the new category feature map.
[0279] Step S9110: Adjust the first category feature map based on the eigenvalue adjustment coefficients of various category eigenvalues to obtain a second category feature map, and merge the first image feature map and the second category feature map along the channel dimension to obtain a merged feature map.
[0280] For the embodiments of the present application, the eigenvalue adjustment coefficients of various category eigenvalues, that is, δ, can be obtained through the above step S9109, and the second category feature map is obtained based on formula (5), where,
[0281]
[0282] where, U t+1 represents the second category feature map, represents the new category feature map.
[0283] Further, after obtaining the first image feature map and the second category feature map, that is, after obtaining V t and U t+1 and then, V t and U t+1 are merged along the channel dimension to obtain a merged feature map. Step S9111: Extract features from the merged feature map to obtain new image features.
[0284] For the embodiments of the present application, after obtaining the merged feature map, it is further processed through Conv3×3, Group Normalization (GroupNorm), Linear rectification function (ReLU), Conv3×3, GroupNorm to obtain a new feature map.
[0285] Further, the way of feature extraction from the merged feature map can be through the above - shown model structure for feature extraction, or through other model structures for feature extraction, or even through other methods to achieve feature extraction, which is not limited in the embodiments of the present application.
[0286] Step S9112: Fuse the new image features with the first image feature map to obtain a second image feature map, and use the second category feature map and the second image feature map as the first category feature map and the first image feature map for the next first operation.
[0287] In the embodiments of the present application, the new image features obtained through the above steps and the first image feature map (V t ) are added together, and a second image feature map (V t+1) That is, the second image feature map is obtained, where the channel dimension is set to 96 and the number of groups for GroupNorm is set to 1.
[0288] Further, after obtaining the second category feature map (U t+1 ) and the second image feature map (V t+1 ) through the above embodiments, the second category feature map (U t+1 ) and the second image feature map (V t+1 ) are used as the first category map and the first image feature map for the next first operation. Step S9113: Loop steps S9101 - S9112 until T loops are reached to obtain the final target category feature map U T .
[0289] Further, by looping the solution method of the diffusion differential equation shown in the above embodiments (steps S9101 - S9112) until T loops are reached to obtain the final target category feature map U T . In the embodiments of the present application, T can be preset. For example, T can be 8.
[0290] Specifically, introduced through a more detailed example, in the embodiments of the present application, the interactive segmentation problem (i.e., the algorithm shown in the embodiments of the present application) can be regarded as a binary information conduction process, and thus modeled as a diffusion partial differential equation:
[0291]
[0292] Among them, where U(t; p) represents the probability of belonging to the positive class at the coordinate point p at time t in this application; V(t; p) is a guiding feature map (i.e., an image feature map), containing features extracted from the image to be segmented; is the diffusion coefficient, the function g(·) is a positive-valued decreasing function, and its form is usually set to and div are the gradient and divergence operators in space, and the finite difference method can be used for approximation. The partial differential equation has Dirichlet boundary conditions that the U value is always 1 at the positive class points and always 0 at the negative class points.
[0293] It can be seen from equation (1) that if is larger, it means that at the boundary of the object in the image to be segmented, the information conduction rate is smaller, so the object boundary tends to divide the image to be segmented into two sides with high / low positive class probabilities. In the embodiments of the present application, by solving the above partial differential equation to obtain U in the steady state, the probabilities of belonging to the positive class and the negative class at each position can be obtained.
[0294] In the embodiments of the present application, asFigure 9c As shown, the image to be segmented is initially downsampled to obtain a first downsampling result. Moreover, after splicing the historical segmentation result and the Click Map, initial downsampling processing is performed to obtain a second downsampling result. The first downsampling result and the second downsampling result are spliced, and the spliced result is subjected to feature extraction through an initial convolutional block to obtain an initial class feature map U 0 and an initial image feature V 0 , and then U 0 and V 0 are subjected to T differential equation solutions through a differential equation solving module to obtain a final class feature map U T , and then the final class feature map U T is segmented through a segmentation head to obtain the target segmentation result of the image to be segmented;
[0295] Furthermore, in the process of subjecting U 0 and V 0 to T differential equation solutions through a differential equation solving module, to obtain the class feature map U corresponding to the T-th differential solution T , and then U T is segmented through a segmentation head (Conv1×1) to obtain the target segmentation result.
[0296] Specifically, the process of subjecting U 0 and V 0 to the first differential equation solution through a differential equation solving module may specifically include: from to obtain the first target class feature map corresponding to the first operation, and then from to obtain the adjusted eigenvalue of each pixel in the first target class feature map, that is, to obtain the adjusted first target class feature map. Then, based on to obtain the class feature map corresponding to the first differential solution, and then from V 1 = Update(V 0 ; U 1 ), that is, based on the class feature map corresponding to the first differential solution and the initial image special map, to obtain the image feature map corresponding to the first differential solution, and then perform the second differential solution, that is, from to obtain the first target class feature map corresponding to the second first operation, and then from to obtain the adjusted eigenvalue of each pixel in the first target class feature map in the second differential solution, that is, to obtain the second adjusted first target class feature map. Then, based on to obtain the category feature map corresponding to the third differential solution, and then by V 2 = Update(V 1 ; U 2 ), to obtain the image feature map corresponding to the second differential solution,..., in the (t + 1)-th differential solution process, it may specifically include: by to obtain Then by Then, based on to obtain U t+1 , further, V t+1 = Update(V t ; U t+1 ), to obtain V t+1 , as shown in Figure 10a , until it loops T times, to obtain U T .
[0297] Furthermore, in the embodiments of the present application, through the above embodiments, it can be ensured that the target image to be segmented can have a more accurate segmentation result at the click of the target object, and in the embodiments of the present application, by constructing a partial differential equation to perform image segmentation, a segmentation effect similar to the comparative method can be achieved with a parameter quantity one order of magnitude lower. As shown in Figure 10b , where the Dynamic Fusion Network for Multi-Domain End-to-end Task-Oriented Dialog (DF-Net) represents the image segmentation method shown in the embodiments of the present application, NoC@85 and NoC@90 are two evaluation indicators, indicating how many clicks are required on average on the current dataset to reach an intersection over union (IoU) of 85% and 90%, and the lower these two indicators are, the better the effect; #Params is the parameter quantity, and the floating point operations (FLOPs) represent the computational amount, both of which are the lower the better. It can be seen that the embodiments of the present application can achieve a segmentation effect similar to the comparative method with a parameter quantity one order of magnitude lower and a lower computational amount. Among them, the comparative Segment Anything Model (SAM) method is the latest popular method and uses a huge training dataset, while the method shown in the embodiments of the present application only uses a small-scale training dataset.
[0298] Based on the same principle as the image segmentation method provided in the embodiments of the present application, the embodiments of the present application also provide an image segmentation device, as shown in Figure 11As shown, the device 110 may include: an acquisition module 111, a feature extraction module 112, a construction module 113, and an optimization and solution module 114, where
[0299] The acquisition module 111 is configured to acquire a target image to be segmented and an interactive information map with click position identifiers generated based on a click operation on the target image. The click operation includes at least one of a positive click operation on a target pixel point or a negative click operation on a non-target pixel point;
[0300] The feature extraction module 112 is configured to splice the target image and the target interactive information map and then perform feature extraction to obtain an initial class feature map and an initial image feature map. The feature value of each class feature point in the initial class feature map is used to represent the probability that the pixel point corresponding to the class feature point belongs to the target pixel point;
[0301] The construction module 113 is configured to use the class feature value in the class feature map as the concentration of the diffusing substance and the gradient of the image feature value in the image feature map as the influencing factor of the diffusion coefficient to construct a diffusion partial differential equation, where the class feature value in the initial class feature map is used as the initial concentration of the diffusing substance, the gradient of the image feature value in the initial image feature map is used as the initial influencing factor of the diffusion coefficient, and the diffusion coefficient is negatively correlated with the gradient;
[0302] The optimization and solution module 114 is configured to optimize and solve the diffusion partial differential equation based on the Dirichlet boundary condition to obtain a target class feature map, and perform image segmentation processing based on the target class feature map to obtain an image segmentation result of the target image, where the condition includes that the feature value of the class feature point corresponding to the positive click is constantly 1 and the negative click is constantly 0.
[0303] In a possible implementation manner of the embodiment of the present application, when the optimization and solution module 114 optimizes and solves the diffusion partial differential equation based on the Dirichlet boundary condition to obtain a target class feature map, it is specifically configured to:
[0304] Use the initial class feature map and the initial image feature map as the first class feature map and the first image feature map of the first first operation, repeat the first operation until a preset condition is met, and use the second class feature map obtained from the last first operation as the target class feature map; where the first operation includes the following steps:
[0305] Determine the class feature value difference between each class feature value in the first class feature map and the domain feature value of the class feature value, and determine the image feature value difference between each image feature value in the first image feature map and the domain feature value of the image feature value;
[0306] Determine the eigenvalue adjustment coefficient map corresponding to the first category feature map according to the category eigenvalue difference and the image eigenvalue difference corresponding to each category eigenvalue;
[0307] Adjust the first category feature map based on the eigenvalue adjustment coefficient map to obtain the second category feature map, and perform feature extraction on the second category feature map and the first image feature map to obtain the second image feature map. Use the second category feature map and the second image feature map as the first category feature map and the first image feature map for the next first operation.
[0308] In a possible implementation manner of the embodiment of the present application, when the optimization solving module 114 optimizes and solves the diffusion partial differential equation based on the Dirichlet boundary condition to obtain the target category feature map and performs image segmentation processing based on the target category feature map to obtain the image segmentation result of the target image, it is specifically used for:
[0309] Use the initial category feature map and the initial image feature map as the first category feature map and the first image feature map for the first first operation, repeat the first operation until a preset condition is met, and perform image segmentation processing based on the second category feature map obtained from the last first operation to obtain the image segmentation result of the target image;
[0310] Wherein, the first operation includes the following steps:
[0311] For each category eigenvalue in the first category feature map, determine the category eigenvalue difference between this category eigenvalue and the neighborhood eigenvalues of this category eigenvalue;
[0312] For each image eigenvalue in the first image feature map, determine the image eigenvalue difference between this image eigenvalue and the neighborhood eigenvalues of this image eigenvalue;
[0313] For each category eigenvalue, determine the eigenvalue adjustment coefficient of this category eigenvalue based on the category eigenvalue difference corresponding to this category eigenvalue and the image eigenvalue difference corresponding to the image eigenvalue corresponding to this category eigenvalue;
[0314] Adjust the first category feature map based on the eigenvalue adjustment coefficients of each category eigenvalue to obtain the second category feature map, and perform feature extraction on the first image feature map and the second category feature map to obtain the second image feature map. Use the second category feature map and the second image feature map as the first category feature map and the first image feature map for the next first operation.
[0315] In another possible implementation manner of the embodiment of the present application, when performing the first operation, the optimization solving module 114 is further used for:
[0316] Determine each first eigenvalue corresponding to the positive click operation and each second eigenvalue corresponding to the negative click operation in the first category feature map respectively;
[0317] Extract features from each first eigenvalue to obtain the positive click feature corresponding to the positive click operation, and extract features from each second eigenvalue to obtain the negative click feature corresponding to the negative click operation;
[0318] Fuse the positive click feature and the negative click feature to obtain the click feature;
[0319] Determine the adjustment weight map of the first category feature map based on the correlation between the click feature and the first category feature map;
[0320] Weight the first category feature map with the adjustment weight map to obtain a new category feature map;
[0321] Among them, when the optimization solving module 114 determines the category feature value difference between each category feature value in the first category feature map and the domain feature value of this category feature value, it is specifically used for:
[0322] For each category feature map in the new category feature map, determine the category feature value difference between this category feature value and the neighborhood feature value of this category feature value.
[0323] In another possible implementation manner, the positive click feature and the negative click feature are obtained through the following method:
[0324] Take each first eigenvalue and each second eigenvalue as the to-be-processed eigenvalue respectively, and perform the following operations on the processed eigenvalue to obtain the click feature corresponding to the to-be-processed feature:
[0325] Extract features from the to-be-processed eigenvalue;
[0326] Perform a linear mapping on the feature map obtained by feature extraction to obtain the corresponding mapping result, and collect the self-attention mechanism to determine the self-correlation between the eigenvalues in the mapping result;
[0327] Weight the mapping result based on the self-correlation to obtain the click feature corresponding to the to-be-processed eigenvalue.
[0328] In another possible implementation manner of the embodiment of the present application, when the optimization solving module extracts features from each first eigenvalue to obtain the positive click feature corresponding to the positive click operation, and extracts features from each second eigenvalue to obtain the negative click feature corresponding to the negative click operation, it is specifically used for:
[0329] Extract features from each first eigenvalue to obtain the feature map corresponding to the positive click operation, and extract features from each second eigenvalue to obtain the feature map corresponding to the negative click operation;
[0330] Perform a linear mapping on the feature map corresponding to the positive click operation to obtain the mapping result corresponding to the positive click operation, and collect the self-attention mechanism to determine the self-correlation between the eigenvalues in the mapping result corresponding to the positive click operation;
[0331] Weight the mapping result corresponding to the positive click operation based on the self-correlation to obtain the positive click feature corresponding to the positive click operation; and,
[0332] Perform a linear mapping on the feature map corresponding to the negative click operation to obtain the mapping result corresponding to the negative click operation, and collect the self-attention mechanism to determine the self-correlation between the eigenvalues in the mapping result corresponding to the negative click operation;
[0333] Weight the mapping result corresponding to the negative click operation based on the self-correlation to obtain the negative click feature corresponding to the negative click operation.
[0334] Another possible implementation manner of the embodiment of the present application is that when the optimization solving module 114 adjusts the first category feature map based on the eigenvalue adjustment coefficient map to obtain the second category feature map, it specifically is used for:
[0335] Adjust the first category feature map based on the eigenvalue adjustment coefficient map to obtain the adjusted first category feature map;
[0336] Perform a fusion process on the adjusted first category feature map and the first category feature map to obtain the second category feature map.
[0337] Another possible implementation manner of the embodiment of the present application is that when the optimization solving module 114 extracts features from the first image feature map and the second category feature map to obtain the second image feature map, it specifically is used for:
[0338] Merge the first image feature map and the second category feature map along the channel dimension to obtain the merged feature map;
[0339] Extract features from the merged feature map to obtain new image features;
[0340] Perform a fusion process on the new image features and the first image feature map to obtain the second image feature map.
[0341] Another possible implementation manner of the embodiment of the present application is that when the feature extraction module 112 extracts features after splicing the target image and the target interaction information map to obtain the initial category feature map and the initial image feature map of the target image, it specifically is used for:
[0342] Perform downsampling processing on the target image to obtain the first downsampling result of the target image;
[0343] Downsample the target interaction information graph to obtain the second downsampling result of the target interaction information graph;
[0344] Concatenate the first downsampling result and the second sampling result to obtain the feature map after concatenation processing;
[0345] Extract features from the feature map after concatenation processing to obtain the initial category feature map and the initial image feature map of the target image.
[0346] In another possible implementation manner of the embodiment of the present application, when the feature extraction module 112 extracts features from the feature map after concatenation processing to obtain the initial category feature map and the initial image feature map of the target image, it specifically is used for:
[0347] Extract features from the feature map after concatenation processing to obtain the initial category feature map of the target image;
[0348] Perform convolution processing on the initial category feature map, and perform non-linear mapping processing on the convolved category feature map and the initial category feature map to obtain the initial image feature map.
[0349] In another possible implementation manner of the embodiment of the present application, when the feature extraction module 112 extracts features after concatenating the target image and the target interaction information graph to obtain the initial category feature map and the initial image feature map of the target image, it specifically is used for:
[0350] Obtain the historical segmentation map obtained by historical segmentation of the target image;
[0351] Perform concatenation processing on the target interaction information graph and the historical segmentation map to obtain the concatenated image;
[0352] Perform feature extraction after concatenating the target image and the concatenated image to obtain the initial category feature map and the initial image feature map of the target image.
[0353] The device in the embodiment of the present application can execute the method provided in the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device in each embodiment of the present application correspond to the steps in the method in each embodiment of the present application. For the detailed function description of each module of the device, reference can specifically be made to the description in the corresponding method shown above, and details are not described herein again.
[0354] Figure 12 Shows a schematic structural diagram of an electronic device applicable to the embodiment of the present application, as Figure 12 shown. This electronic device can be a server or a terminal device, and this electronic device can be used to implement the method provided in any embodiment of the present application.
[0355] Such as Figure 12As shown, the electronic device 2000 may mainly include at least one processor 2001 ( Figure 12 one is shown in), a memory 2002, a communication module 2003, an input / output interface 2004, etc. Optionally, the components may be connected and communicate with each other through a bus 2005. It should be noted that Figure 12 the structure of the electronic device 2000 shown in is only schematic and does not constitute a limitation on the electronic device applicable to the method provided in the embodiments of the present application.
[0356] Among them, the memory 2002 can be used to store an operating system and application programs, etc. The application programs can include computer programs that implement the methods shown in the embodiments of the present application when called by the processor 2001, and can also include programs for implementing other functions or services. The memory 2002 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and computer programs, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0357] The processor 2001 is connected to the memory 2002 via the bus 2005 and realizes corresponding functions by calling the application programs stored in the memory 2002. Among them, the processor 2001 can be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof, which can implement or execute various exemplary logic blocks, modules, and circuits described in combination with the disclosure of this application. The processor 2001 can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0358] The electronic device 2000 can be connected to a network through the communication module 2003 (which can include, but is not limited to, components such as a network interface) to communicate with other devices (such as user terminals or servers, etc.) through the network to achieve data interaction, such as sending data to other devices or receiving data from other devices. Among them, the communication module 2003 can include a wired network interface and / or a wireless network interface, etc., that is, the communication module can include at least one of a wired communication module or a wireless communication module.
[0359] The electronic device 2000 can be connected to the required input / output devices, such as a keyboard, a display device, etc., through the input / output interface 2004. The electronic device 2000 itself can have a display device and can also externally connect other display devices through the interface 2004. Optionally, a storage device, such as a hard disk, etc., can also be connected through this interface 2004 to store the data in the electronic device 2000 into the storage device, or read the data in the storage device, and can also store the data in the storage device into the memory 2002. It can be understood that the input / output interface 2004 can be a wired interface or a wireless interface. According to different actual application scenarios, the devices connected to the input / output interface 2004 can be components of the electronic device 2000 or external devices connected to the electronic device 2000 when needed.
[0360] The bus 2005 for connecting each component may include a path for transmitting information between the above components. The bus 2005 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. According to different functions, the bus 2005 may be divided into an address bus, a data bus, a control bus, etc.
[0361] Optionally, for the solution provided in the embodiments of the present application, the memory 2002 may be used to store a computer program for executing the solution of the present application, and the processor 2001 runs the computer program to implement the actions of the method or device provided in the embodiments of the present application.
[0362] Based on the same principle as the method provided in the embodiments of the present application, the embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the corresponding content of the foregoing method embodiments can be implemented.
[0363] The embodiments of the present application further provide a computer program product, which includes a computer program, and when the computer program is executed by a processor, the corresponding content of the foregoing method embodiments can be implemented.
[0364] It should be noted that the terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification, claims and the above drawings of the present application are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data may be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than that shown or described in words.
[0365] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0366] It should be understood that although the flowchart of the embodiments of the present application indicates various operation steps by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated in this article, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.
[0367] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, adopting other similar implementation means based on the technical idea of the present application also belongs to the protection scope of the embodiments of the present application.
Claims
1. A method for image segmentation, characterized in that, The method includes: Obtaining a target image to be segmented and an interactive information map with click position identifiers generated based on a click operation on the target image, where the click operation includes at least one of a positive click operation on a target pixel point or a negative click operation on a non-target pixel point; After splicing the target image and the target interactive information map, feature extraction is performed to obtain an initial class feature map and an initial image feature map, where the feature value of each class feature point in the initial class feature map is used to represent the probability that the pixel point corresponding to this class feature point belongs to the target pixel point; Taking the class feature values in the class feature map as the concentration of the diffusing substance and the gradient of the image feature values in the image feature map as the influencing factor of the diffusion coefficient, a diffusion partial differential equation is constructed, where the class feature values in the initial class feature map are used as the initial concentration of the diffusing substance, and the gradient of the image feature values in the initial image feature map is used as the initial influencing factor of the diffusion coefficient, and the diffusion coefficient is negatively correlated with the gradient; Based on the Dirichlet boundary conditions, the diffusion partial differential equation is optimized and solved to obtain a target class feature map, and image segmentation processing is performed based on the target class feature map to obtain the image segmentation result of the target image, where the conditions include that the feature value of the class feature point corresponding to the positive click is constantly 1 and the negative click is constantly 0.
2. The method according to claim 1, characterized in that, The optimizing and solving the diffusion partial differential equation based on the Dirichlet boundary conditions to obtain a target class feature map includes: Taking the initial class feature map and the initial image feature map as the first class feature map and the first image feature map of the first operation for the first time, repeating the first operation until a preset condition is met, and taking the second class feature map obtained by the last first operation as the target class feature map; where the first operation includes the following steps: Determining the class feature value difference between each class feature value in the first class feature map and the domain feature value of this class feature value, and determining the image feature value difference between each image feature value in the first image feature map and the domain feature value of this image feature value; According to the class feature value differences and image feature value differences corresponding to each class feature value, determining a feature value adjustment coefficient map corresponding to the first class feature map; Adjusting the first class feature map based on the feature value adjustment coefficient map to obtain a second class feature map, and by performing feature extraction on the second class feature map and the first image feature map, obtaining a second image feature map, and taking the second class feature map and the second image feature map as the first class feature map and the first image feature map of the next first operation.
3. The method according to claim 2, wherein The first operation further includes: Respectively determining each first feature value corresponding to the positive click operation and each second feature value corresponding to the negative click operation in the first class feature map; Performing feature extraction on each of the first feature values to obtain a positive click feature corresponding to the positive click operation, and performing feature extraction on each of the second feature values to obtain a negative click feature corresponding to the negative click operation; Fusing the positive click feature and the negative click feature to obtain a click feature; Determine an adjustment weight map of the first category feature map based on the correlation between the click feature and the first category feature map; Weight the first category feature map using the adjustment weight map to obtain a new category feature map; Among them, determining the category feature value difference between each category feature value in the first category feature map and the domain feature value of the category feature value includes: For each category feature map in the new category feature map, determine the category feature value difference between the category feature value and the neighborhood feature value of the category feature value.
4. The method according to claim 3, wherein The positive click feature and the negative click feature are obtained through the following method: Respectively use the first feature values and the second feature values as the to-be-processed feature values, and perform the following operations on the processed feature values to obtain the click feature corresponding to the to-be-processed feature: Extract features from the to-be-processed feature value; Perform a linear mapping on the feature map obtained by feature extraction to obtain a corresponding mapping result, and collect the self-attention mechanism to determine the self-correlation between the feature values in the mapping result; Weight the mapping result based on the self-correlation to obtain the click feature corresponding to the to-be-processed feature value.
5. The method according to claim 2, characterized in that, The adjusting the first category feature map based on the eigenvalue adjustment coefficient map to obtain a second category feature map includes: Adjust the first category feature map based on the eigenvalue adjustment coefficient map to obtain an adjusted first category feature map; Perform a fusion process on the adjusted first category feature map and the first category feature map to obtain the second category feature map.
6. The method according to claim 2, characterized in that, The obtaining the second image feature map by performing feature extraction on the first image feature map and the second category feature map includes: Merge the first image feature map and the second category feature map along the channel dimension to obtain a merged feature map; Perform feature extraction on the merged feature map to obtain new image features; Perform a fusion process on the new image features and the first image feature map to obtain the second image feature map.
7. The method according to claim 1, characterized in that, The obtaining the initial category feature map and the initial image feature map of the target image by performing feature extraction after splicing the target image and the target interaction information map includes: Perform downsampling on the target image to obtain the first downsampling result of the target image; Perform downsampling on the target interaction information map to obtain the second downsampling result of the target interaction information map; Perform a splicing process on the first downsampling result and the second sampling result to obtain a spliced feature map; Perform feature extraction on the spliced feature map to obtain the initial category feature map and the initial image feature map of the target image.
8. The method according to claim 7, characterized in that The performing feature extraction on the spliced feature map to obtain the initial category feature map and the initial image feature map of the target image includes: Perform feature extraction on the spliced feature map to obtain the initial category feature map of the target image; Perform convolution on the initial category feature map, and perform non-linear mapping processing on the convolved category feature map and the initial category feature map to obtain the initial image feature map.
9. The method according to claim 1, characterized in that, Performing feature extraction after splicing the target image and the target interaction information graph to obtain an initial class feature map and an initial image feature map of the target image, includes: Obtaining a historical segmentation graph obtained by performing historical segmentation on the target image; Performing splicing processing on the target interaction information graph and the historical segmentation graph to obtain a spliced image; Performing feature extraction after splicing the target image and the spliced image to obtain an initial class feature map and an initial image feature map of the target image.
10. An apparatus for image segmentation, characterized in that, The device includes: An acquisition module, configured to acquire a target image to be segmented and an interaction information graph with click position identifiers generated based on a click operation on the target image, where the click operation includes at least one of a positive click operation on a target pixel point or a negative click operation on a non-target pixel point; A feature extraction module, configured to perform feature extraction after splicing the target image and the target interaction information graph to obtain an initial class feature map and an initial image feature map, where the feature value of each class feature point in the initial class feature map is used to represent the probability that the pixel point corresponding to the class feature point belongs to a target pixel point; A construction module, configured to use the class feature value in the class feature map as the concentration of a diffusing substance and use the gradient of the image feature value in the image feature map as an influencing factor of the diffusion coefficient to construct a diffusion partial differential equation, where the class feature value in the initial class feature map is used as the initial concentration of the diffusing substance, and the gradient of the image feature value in the initial image feature map is used as the initial influencing factor of the diffusion coefficient, and the diffusion coefficient is negatively correlated with the gradient; An optimization and solution module, configured to perform optimization and solution on the diffusion partial differential equation based on Dirichlet boundary conditions to obtain a target class feature map, and perform image segmentation processing based on the target class feature map to obtain an image segmentation result of the target image, where the conditions include that the feature value of the class feature point corresponding to the positive click is constantly 1 and the negative click is constantly 0.
11. An electronic device, characterized in that, The electronic device includes a memory and a processor, where a computer program is stored in the memory, and the processor executes the image segmentation method according to any one of claims 1 to 9 when running the computer program.
12. A computer-readable storage medium, characterized in that, A computer program is stored in the storage medium, and when the computer program is executed by a processor, the image segmentation method according to any one of claims 1 to 9 is implemented.