Image Segmentation Method and Device
By decoupling feature extraction and user interaction processing in image segmentation, using a larger network to extract rich features and combining a lightweight network for interactive processing, the problems of computational overhead and time-consuming in interactive segmentation are solved, and fast and accurate segmentation results are achieved.
Patent Information
- Application Number
- CN202210239557.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-11
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-03-11
AI Technical Summary
Interactive segmentation technology often leads to very large computing overhead and operation time-consuming, especially due to the strong coupling of feature extraction and user interaction processing, which requires multiple iterations, resulting in a large network scale and high computing complexity.
By decoupling the image segmentation process into two independent stages: feature extraction and user interaction, a larger network is used to perform feature extraction to obtain rich image features, a relatively small network is used to perform user interaction processing, and the results of the two are fused to generate a mask, thereby achieving fast and accurate segmentation results.
While ensuring the segmentation speed, accurate segmentation results can be obtained based on the obtained mask, which reduces calculation overhead and operation time, and solves common performance bottleneck problems in interactive segmentation.
Smart Images

Figure CN114565767B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing, and in particular, to an image segmentation method and apparatus. Background Art
[0002] Image segmentation technology is a very important computer vision task, and it has many applications in image retrieval, picture editing, and film and television production. Interactive segmentation, as a specific segmentation method in the field of image segmentation, aims to distinguish the object of interest and the background with the least user input and inference time and achieve the best segmentation accuracy. Due to the diversity of user input information (such as clicks, scribbles, bounding boxes, etc.), interactive segmentation provides great flexibility for users and can effectively adjust the current segmentation result according to user guidance.
[0003] Interactive segmentation can be specifically divided into two subtasks: feature encoding of the input image and user interaction processing. Currently, interactive segmentation usually strongly couples the two subtasks and obtains the target mask by integrating the feature extraction of the input image and user interaction processing in an end-to-end manner. In this process, multiple iterations are required, that is, the entire network is repeatedly experienced multiple times to obtain the segmentation result. However, in order to obtain a better segmentation result, a relatively large network is generally used for interactive segmentation, that is, a relatively large network is used for feature extraction and user interaction processing, which will bring very large computational overhead and running time consumption. Summary of the Invention
[0004] The present disclosure provides an image segmentation method and apparatus to at least solve the problem that interactive segmentation in related technologies often leads to very large computational overhead and running time consumption.
[0005] According to a first aspect of an embodiment of the present disclosure, an image segmentation method is provided, including: inputting an image to be processed into a first image feature extraction network to obtain a first image feature; inputting preset reference information, first interaction information, and the image to be processed into a second image feature extraction network to obtain a second image feature, where the amount of information of the image features extracted from the image by the second image feature extraction network is less than that of the image features extracted from the image by the first image feature extraction network, and the first interaction information is position information for indicating the object to be segmented in the image to be processed; obtaining a first target mask for the object to be segmented based on the first image feature and the second image feature; and performing segmentation processing on the image to be processed based on the first target mask to obtain a first segmentation result for the object to be segmented.
[0006] Optionally, in the case where the first segmentation result does not meet the preset requirements, the image segmentation method further includes: inputting the first target mask, the second interaction information, and the image to be processed into a second image feature extraction network to obtain a third image feature, where the second interaction information is used to indicate the position information of the object to be segmented in the image to be processed; obtaining a second target mask for the object to be segmented based on the first image feature and the third image feature; and performing segmentation processing on the image to be processed based on the second target mask to obtain a second segmentation result for the object to be segmented.
[0007] Optionally, obtaining a first target mask for the object to be segmented based on the first image feature and the second image feature includes: fusing the first image feature and the second image feature to obtain a fused image feature; and inputting the fused image feature into an image mask extraction network to obtain the first target mask.
[0008] Optionally, fusing the first image feature and the second image feature to obtain a fused image feature includes: concatenating the first image feature and the second image feature to obtain a concatenated feature; and inputting the concatenated feature into a residual network to obtain the fused image feature.
[0009] Optionally, the first interaction information includes positive interaction information and / or negative interaction information, where the positive interaction information is used to indicate the region where the object to be segmented is located in the image to be processed, and the negative interaction information is used to indicate the region outside the region where the object to be segmented is located in the image to be processed.
[0010] Optionally, the second interaction information is determined based on the first segmentation result and the second interaction information includes positive interaction information and / or negative interaction information, where the positive interaction information is used to indicate the region where the object to be segmented is located in the image to be processed, and the negative interaction information is used to indicate the region outside the region where the object to be segmented is located in the image to be processed.
[0011] According to a second aspect of the embodiments of the present disclosure, there is provided an image segmentation apparatus, including: a first image feature acquisition unit configured to input an image to be processed into a first image feature extraction network to obtain a first image feature; a second image feature acquisition unit configured to input preset reference information, first interaction information, and the image to be processed into a second image feature extraction network to obtain a second image feature, where the amount of information of the image features extracted by the second image feature extraction network from the image is less than the amount of information of the image features extracted by the first image feature extraction network from the image, and the first interaction information is used to indicate the position information of the object to be segmented in the image to be processed; a mask acquisition unit configured to obtain a first target mask for the object to be segmented based on the first image feature and the second image feature; and a segmentation unit configured to perform segmentation processing on the image to be processed based on the first target mask to obtain a first segmentation result for the object to be segmented.
[0012] Optionally, when the first segmentation result does not meet the preset requirements, the second image feature acquisition unit is further configured to input the first target mask, the second interaction information, and the image to be processed into a second image feature extraction network to obtain a third image feature, where the second interaction information is used to indicate the position information of the object to be segmented in the image to be processed; the mask acquisition unit is further configured to obtain a second target mask for the object to be segmented based on the first image feature and the third image feature; the segmentation unit is further configured to perform segmentation processing on the image to be processed based on the second target mask to obtain a second segmentation result for the object to be segmented.
[0013] Optionally, the mask acquisition unit is further configured to fuse the first image feature and the second image feature to obtain a fused image feature; and input the fused image feature into an image mask extraction network to obtain a first target mask.
[0014] Optionally, the mask acquisition unit is further configured to splice the first image feature and the second image feature to obtain a spliced feature; and input the spliced feature into a residual network to obtain a fused image feature.
[0015] Optionally, the first interaction information includes positive interaction information and / or negative interaction information, where the positive interaction information is used to indicate the area where the object to be segmented is located in the image to be processed, and the negative interaction information is used to indicate the area outside the area where the object to be segmented is located in the image to be processed.
[0016] Optionally, the second interaction information is determined based on the first segmentation result and the second interaction information includes positive interaction information and / or negative interaction information, where the positive interaction information is used to indicate the area where the object to be segmented is located in the image to be processed, and the negative interaction information is used to indicate the area outside the area where the object to be segmented is located in the image to be processed.
[0017] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including: a processor; a memory for storing processor-executable instructions; where the processor is configured to execute the instructions to implement the image segmentation method according to the present disclosure.
[0018] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. When the instructions in the computer-readable storage medium are run by at least one processor, at least one processor is caused to execute the image segmentation method according to the present disclosure as described above.
[0019] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, including computer instructions that, when executed by a processor, implement the image segmentation method according to the present disclosure.
[0020] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0021] According to the image segmentation method and device of the present disclosure, the two tasks of feature extraction and user interaction processing are decoupled. That is, a relatively large network is used for feature extraction to obtain relatively rich image features, and a relatively small network is used for user interaction processing to ensure the segmentation speed. Then, the extracted image features are fused with the user interaction processing results, and a relatively accurate mask can be obtained according to the fusion result. Thus, accurate segmentation results can be obtained based on the obtained mask while ensuring the segmentation speed. Therefore, the present disclosure solves the problem that interactive segmentation often causes very large computational overhead and running time in the related art.
[0022] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation to the present disclosure.
[0024] Figure 1 is a schematic diagram of an implementation scenario of an image segmentation method according to an exemplary embodiment of the present disclosure;
[0025] Figure 2 is a flowchart of an image segmentation method shown according to an exemplary embodiment;
[0026] Figure 3 is a schematic structural diagram of a network adopted by an image segmentation method shown according to an exemplary embodiment;
[0027] Figure 4 is a block diagram of an image segmentation device shown according to an exemplary embodiment;
[0028] Figure 5 is a block diagram of an electronic device 500 according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0030] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present disclosure are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following examples do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0031] It should be noted here that "at least one of several items" in the present disclosure all represents three parallel situations, including "any one of the several items", "a combination of any multiple of the several items", and "the whole of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example is "performing at least one of step one and step two", which means the following three parallel situations: (1) performing step one; (2) performing step two; (3) performing step one and step two.
[0032] The present disclosure provides an image segmentation method, which can obtain an accurate segmentation result based on the obtained mask while ensuring the segmentation speed. Hereinafter, an example of the scene of human segmentation in image segmentation will be used for illustration.
[0033] Figure 1 is a schematic diagram of an implementation scenario of an image segmentation method according to an exemplary embodiment of the present disclosure, as Figure 1 described, the implementation scenario includes a server 100, a user terminal 110 and a user terminal 120. Among them, the number of user terminals is not limited to 2, including but not limited to devices such as mobile phones and personal computers. The user terminal can be equipped with a camera for obtaining images. The server can be a single server, or a server cluster composed of several servers, or a cloud computing platform or a virtualization center.
[0034] The user terminal 110 or the user terminal 120 acquires an image containing a person through a camera, and uploads the image as an image to be processed to the server 100. The server 100 inputs the image to be processed into a first image feature extraction network to obtain a first image feature, and inputs a preset reference information, a first interaction information, and the image to be processed into a second image feature extraction network to obtain a second image feature, where the amount of information of the image features extracted from the image by the second image feature extraction network is less than that of the image features extracted from the image by the first image feature extraction network, and the first interaction information is position information for indicating the object to be segmented in the image to be processed. Then, based on the first image feature and the second image feature, a first target mask for the object to be segmented is obtained, and then, based on the first target mask, the image to be processed is segmented to obtain a first segmentation result for the object to be segmented.
[0035] Next, an image segmentation method and apparatus according to an exemplary embodiment of the present disclosure will be described in detail with reference to the accompanying drawings.
[0036] Figure 2 is a flowchart of an image segmentation method shown according to an exemplary embodiment, as Figure 2 shown, the image segmentation method includes the following steps:
[0037] In step S201, the image to be processed is input into a first image feature extraction network to obtain a first image feature. The first image feature extraction network can be implemented by an encoder in an image processing network having an encoder-decoder structure.
[0038] For better understanding, the following will be described in detail in conjunction with Figure 3 For details, Figure 3 is a schematic structural diagram of a network adopted by an image segmentation method shown according to an exemplary embodiment, as Figure 3 shown, the network framework of the image segmentation method can be divided into two stages, stage I: Feature Extraction and stage II: Interactive Refinement.
[0039] Figure 3 The stage I shown is an optional feature extraction method for step S201. Denote i as the current interactive segmentation round, M i as the target mask for interactive segmentation in the i-th round, I RGB as the RGB information of the image to be processed, and F c as the above-mentioned first image feature. In the feature extraction stage, I RGBInput into a larger network (i.e., the first image feature extraction network mentioned above), and extract better semantic features through a partial structure of the larger network (such as an encoder). For example, use a partial structure of HRNet (such as an encoder) as the feature extraction network to extract high-resolution and detail-complete image features F from the image to be processed. c (equivalent to the first image feature mentioned above), F c Can be reused in subsequent user interaction processing.
[0040] In this embodiment, different from the related art where feature extraction and user interaction processing are regarded as a whole and integrated into a network with shared weights, this embodiment uses an independent larger network to extract semantic information of the image, so as to extract high-resolution and detail-complete image features to assist the interactive segmentation processing of the subsequent lightweight network and improve the accuracy of interactive segmentation.
[0041] In step S202, the preset reference information, the first interaction information, and the image to be processed are input into the second image feature extraction network to obtain the second image feature. Among them, the amount of information of the image features extracted by the second image feature extraction network from the image is less than that of the image features extracted by the first image feature extraction network from the image. The first interaction information is the position information used to indicate the object to be segmented in the image to be processed. The above preset reference information can be set as needed. For example, but not limited to, in order not to increase the computational complexity, the preset reference information can be set to be empty. The above second image feature extraction network can be implemented by the encoder in an image processing network with an encoder-decoder structure. It should be noted that the fact that the amount of information of the image features extracted by the second image feature extraction network from the image is less than that of the image features extracted by the first image feature extraction network from the image indicates that for the same image, the representation ability of the second image feature extraction network for the image is weak, and the representation ability of the first image feature extraction network for the image is strong. Generally, the strength and weakness of the representation ability can be distinguished by the number of parameters in the network. If the number of parameters in the network is greater than the first preset value, it represents to a certain extent that the representation ability of the network is strong. If the number of parameters in the network is less than the second preset value, it represents to a certain extent that the representation ability of the network is weak. Applied to the present disclosure, that is, the number of parameters in the second image feature extraction network is less than the second preset value, and the number of parameters in the first image feature extraction network is greater than the second preset value. The first preset value and the second preset value can be defined as needed.
[0042] According to an exemplary embodiment of the present disclosure, the first interaction information includes positive interaction information and / or negative interaction information. Among them, the positive interaction information is used to indicate the region where the object to be segmented is located in the image to be processed, and the negative interaction information is used to indicate the region outside the region where the object to be segmented is located in the image to be processed. According to this embodiment, the first interaction information may include both positive interaction information and negative interaction information to improve the final segmentation result.
[0043] For example, the interaction information can be obtained by recognizing user input information (such as clicking, scribbling, bounding boxes, etc.). Taking scribbling as an example, assuming that the user scribbles on the object to be segmented, the recognized interaction information at this time is positive interaction information, that is, it indicates the region where the object to be segmented is located in the image to be processed. Assuming that the user scribbles on the background of the object to be segmented, the recognized interaction information at this time is negative interaction information, that is, it indicates the region outside the region where the object to be segmented is located in the image to be processed.
[0044] For better understanding, still taking Figure 3 the structural schematic diagram for detailed description, as shown in Figure 3 Stage II: Interactive Refinement shown in the figure, where is the positive interaction information in the i-th round, the negative interaction information in the i-th round, M i-1 is the target mask used for interactive segmentation in the previous interactive segmentation round i - 1, is the above-mentioned second image feature or third image feature. It should be noted that the above M i-1 When i is 1, that is, when the current interactive segmentation round is the first round, there is no previous interactive segmentation round at this time, so M i-1 can be set to 0 or other preset values (that is, the above-mentioned preset reference information), as long as the computational complexity is not increased.
[0045] In the interactive refinement stage, to enhance the stability of the iterative adjustment, according to the interaction information given by the current user, the lightweight network iteratively adjusts the current target mask to obtain a good segmentation result. Specifically, the target mask M i-1 , (equivalent to the above positive interaction information) and (equivalent to the above negative interaction information) are input into the lightweight network (that is, the above-mentioned second image feature extraction network), and through part of the structure of the lightweight network (such as the encoder),
[0046] In this embodiment, an independent and relatively small network (i.e., a lightweight network) is adopted to extract the semantic information of the image with interaction information. For example, a partial structure (such as an encoder) of a spectral auto-associative feedback neural network (SARN) is used as a feature extraction network to extract low-resolution and incomplete-detail image features carrying interaction information from the interaction information and the image to be processed. These low-resolution and incomplete-detail image features can be combined with the previously extracted high-resolution and complete-detail image features for interactive segmentation processing to ensure the speed of the interactive segmentation processing. The specific combination method of the low-resolution and incomplete-detail image features and the high-resolution and complete-detail image features will be discussed in detail in the following steps and will not be elaborated here.
[0047] Return Figure 2 , in step S203, based on the first image feature and the second image feature, a first target mask for the object to be segmented is obtained.
[0048] According to an exemplary embodiment of the present disclosure, obtaining a first target mask for the object to be segmented based on the first image feature and the second image feature includes: fusing the first image feature and the second image feature to obtain a fused image feature; inputting the fused image feature into an image mask extraction network to obtain the first target mask. According to this embodiment, the first image feature extracted by a larger network enhances the second image feature to ensure the accuracy of the mask used for segmentation, so as to obtain an accurate segmentation result.
[0049] For example, since the first image feature is extracted by a larger network, the semantic information it contains is relatively complete. Therefore, fusing the first image feature with the second image feature can enrich the second image feature, and then inputting the fused image feature into the image mask extraction network to obtain the first target mask. For example, the image mask extraction network can be implemented by the decoder in an image processing network with an encoder-decoder structure. The second image feature extraction network and the image mask extraction network can belong to the same image processing network, that is, both belong to a relatively small network. In this way, the second image feature is obtained through the second image feature extraction network of the same image processing network, then the first image feature and the second image feature are fused to obtain a fused image feature, and then the fused image feature is continuously input into the image mask extraction network of the same image processing network, and a relatively accurate first target mask can be obtained.
[0050] According to an exemplary embodiment of the present disclosure, fusing a first image feature and a second image feature to obtain a fused image feature includes: splicing the first image feature and the second image feature to obtain a spliced feature; inputting the spliced feature into a residual network to obtain a fused image feature. According to this embodiment, the first image feature and the second image feature can be conveniently and quickly fused through the splicing operation and the residual network. It should be noted that the above residual network can adopt ResBlock, which can design networks with different structures and different complexities as long as it can help fuse the two features, and the present disclosure does not limit this.
[0051] For better understanding, still taking Figure 3 's structural schematic diagram as a detailed illustration, as Figure 3 shown in Stage II: Interactive Refinement, where is the above-mentioned fused image feature. In the iterative interactive adjustment stage, for each interactive segmentation, the feature F c in Stage I will be reused and feature fusion is performed in Stage II to achieve a highly accurate segmentation effect. For example, F c can be fused in an Adaptive Feature Fusion module to achieve the fusion of semantic features and interactive information. Specifically, the Adaptive Feature Fusion module splices F c and , that is, concatenation, and the spliced feature is sent into a Residual Network (ResBlock) module to output the fused feature After obtaining the fused feature , it is input into the network behind the encoder of the lightweight network (equivalent to the above-mentioned image mask extraction network) to obtain a target mask (equivalent to the above-mentioned first target mask) for target segmentation of the image to be processed.
[0052] In step S204, based on the first target mask, the image to be processed is segmented to obtain a first segmentation result for the object to be segmented. For example, the first target mask can be multiplied by the image to be processed to obtain the first segmentation result.
[0053] According to an exemplary embodiment of the present disclosure, when the first segmentation result does not meet the preset requirements, the image segmentation method further includes: inputting the first target mask, the second interaction information, and the image to be processed into a second image feature extraction network to obtain third image features, where the second interaction information is used to indicate the position information of the object to be segmented in the image to be processed; obtaining a second target mask for the object to be segmented based on the first image features and the third image features; and performing segmentation processing on the image to be processed based on the second target mask to obtain a second segmentation result for the object to be segmented. According to this embodiment, when the first segmentation result does not meet the user's requirements, a second round of segmentation can be performed. At this time, the first image features extracted in the previous round are still used without re-extraction, avoiding the time waste of secondary extraction using a large network. At the same time, combining the target mask of the previous round to obtain the current target mask also improves the accuracy of the second segmentation result.
[0054] For example, the above preset requirements can be set as needed or temporarily input according to the user's observation. For example, when the user observes that the first segmentation result does not meet the requirements, the user inputs indication information as the preset requirement to control the start of the second segmentation process. Here, it is assumed that the first segmentation result does not meet the user's needs, and a second segmentation process is required. At this time, for the same image to be processed, the first image features extracted by the larger network in the first segmentation process can be directly used without re-extraction. Moreover, the third image features with interaction information in the second segmentation process refer to the target mask of the previous segmentation process, which is equivalent to further adjusting the previous second image features. After obtaining the third image features, the first image features and the third image features are still fused to enrich the third image features, and the fused image features are continuously input into the network behind the encoder in the second network, and a relatively accurate second target mask can be obtained.
[0055] According to an exemplary embodiment of the present disclosure, the second interaction information is determined based on the first segmentation result, and the second interaction information includes positive interaction information and / or negative interaction information, where the positive interaction information is used to indicate the area where the object to be segmented is located in the image to be processed, and the negative interaction information is used to indicate the area outside the area where the object to be segmented is located in the image to be processed. According to this embodiment, the second interaction information can include both positive interaction information and negative interaction information to improve the final segmentation result, and the second interaction information is determined based on the first segmentation result, which can further improve the final segmentation result.
[0056] It should be noted that the second mutual information and the first interaction information can be the same or different. When they are different, the second interaction information can also be determined based on the first segmentation result. Specifically, when the first segmentation result is not ideal, the user can adjust the interaction information to be used in the next round as needed, such as increasing the smeared content to adjust the interaction information.
[0057] For better understanding, still taking the Figure 3 structural schematic diagram as an example for detailed description, as shown in Figure 3 Stage II: Interactive Refinement. In the interactive refinement stage, if the segmentation result of the first interactive segmentation round does not meet the preset requirements (such as not meeting the user's needs), the target mask M of the previous interactive segmentation round i-1 i-1 , (equivalent to the above forward interactive information) and (equivalent to the above reverse interactive information) are input into the lightweight network, and the of this interactive segmentation round is obtained through part of the structure of the lightweight network (such as the encoder). It should be noted that the and here may be the same as or different from those of the previous interactive segmentation round. For example, for the convenience of the user, the and of the previous interactive segmentation round are directly adopted. Another example is that when the first segmentation result is not ideal, the user can adjust the interactive information to be used in this round according to needs, such as increasing the smeared content to adjust the interactive information, that is, adopting and different from those of the previous interactive segmentation round. and
[0058] In this interactive segmentation round, the feature F of the previous interactive segmentation round is continued to be used c , and the adaptive feature fusion module concatenates F c and the of this interactive segmentation round, that is, concatenation. The concatenated feature is sent into a ResBlock module to output the fused feature of this interactive segmentation round. After obtaining the fused feature , it is input into the network behind the encoder of the lightweight network (equivalent to the above image mask extraction network) to obtain the target mask (equivalent to the above second target mask) for target segmentation of the image to be processed.
[0059] In summary, in view of the drawbacks caused by the strong coupling between the feature extraction and user interaction processing stages in the previous architecture design, the present disclosure decouples interactive segmentation into a feature extraction stage and a user interaction stage; for each image to be processed, the feature extraction process is only performed once, and a relatively large network can be used in this process, and this feature extraction process is Stage I. For each interactive segmentation process, a lightweight network can be selected, and according to the user interaction information (i.e., the interaction information), the target mask of the previous interactive segmentation, and an additional Adaptive Feature Fusion module is added to fuse the features of Stage I to perform interactive segmentation, so as to achieve a better balance between speed and segmentation accuracy and optimize the user experience.
[0060] Figure 4 is a block diagram of an image segmentation device shown according to an exemplary embodiment. Referring to Figure 4 , the device includes:
[0061] A first image feature acquisition unit 40, configured to input an image to be processed into a first image feature extraction network to obtain a first image feature; a second image feature acquisition unit 42, configured to input preset reference information, first interaction information, and an image to be processed into a second image feature extraction network to obtain a second image feature, wherein the amount of information of the image features extracted from the image by the second image feature extraction network is less than the amount of information of the image features extracted from the image by the first image feature extraction network, and the first interaction information is position information for indicating the position of an object to be segmented in the image to be processed; a mask acquisition unit 44, configured to obtain a first target mask for the object to be segmented based on the first image feature and the second image feature; a segmentation unit 46, configured to perform segmentation processing on the image to be processed based on the first target mask to obtain a first segmentation result for the object to be segmented.
[0062] According to an exemplary embodiment of the present disclosure, in the case where the first segmentation result does not meet the preset requirements, the second image feature acquisition unit 42 is further configured to input the first target mask, second interaction information, and an image to be processed into the second image feature extraction network to obtain a third image feature, wherein the second interaction information is position information for indicating the position of an object to be segmented in the image to be processed; the mask acquisition unit 44 is further configured to obtain a second target mask for the object to be segmented based on the first image feature and the third image feature; the segmentation unit 46 is further configured to perform segmentation processing on the image to be processed based on the second target mask to obtain a second segmentation result for the object to be segmented.
[0063] According to an exemplary embodiment of the present disclosure, the mask acquisition unit 44 is further configured to fuse the first image feature and the second image feature to obtain a fused image feature; and input the fused image feature into an image mask extraction network to obtain a first target mask.
[0064] According to an exemplary embodiment of the present disclosure, the mask acquisition unit 44 is further configured to splice the first image feature and the second image feature to obtain a spliced feature; and input the spliced feature into a residual network to obtain a fused image feature.
[0065] According to an exemplary embodiment of the present disclosure, the first interaction information includes positive interaction information and / or negative interaction information, wherein the positive interaction information is used to indicate the region where the object to be segmented in the image to be processed is located, and the negative interaction information is used to indicate the region outside the region where the object to be segmented in the image to be processed is located.
[0066] According to an exemplary embodiment of the present disclosure, the second interaction information is determined based on the first segmentation result and the second interaction information includes positive interaction information and / or negative interaction information, wherein the positive interaction information is used to indicate the region where the object to be segmented in the image to be processed is located, and the negative interaction information is used to indicate the region outside the region where the object to be segmented in the image to be processed is located.
[0067] According to an embodiment of the present disclosure, an electronic device may be provided. Figure 5 FIG. 500 is a block diagram of an electronic device 500 according to an embodiment of the present disclosure. The electronic device includes at least one memory 501 and at least one processor 502. A set of computer-executable instructions is stored in the at least one memory. When the set of computer-executable instructions is executed by the at least one processor, an image segmentation method according to an embodiment of the present disclosure is performed.
[0068] As an example, the electronic device 500 may be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instruction set. Here, the electronic device 1000 does not necessarily have to be a single electronic device, and may also be any assembly of devices or circuits capable of executing the above instructions (or instruction sets) alone or jointly. The electronic device 500 may also be a part of an integrated control system or a system manager, or may be configured as a portable electronic device that is interconnected with a local or remote device (e.g., via wireless transmission).
[0069] In the electronic device 500, the processor 502 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor 502 may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0070] The processor 502 can execute instructions or code stored in the memory, and the memory 501 can also store data. The instructions and data can also be sent and received via the network interface device through the network, where the network interface device can adopt any known transmission protocol.
[0071] The memory 501 can be integrated with the processor 502. For example, RAM or flash memory can be arranged within an integrated circuit microprocessor, etc. In addition, the memory 501 can include independent devices, such as external disk drives, storage arrays, or other storage devices that can be used by any database system. The memory 501 and the processor 502 can be operatively coupled or can communicate with each other, for example, through I / O ports, network connections, etc., such that the processor 502 can read the files stored in the memory 501.
[0072] In addition, the electronic device 500 can also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, mouse, touch input device, etc.). All components of the electronic device can be connected to each other via a bus and / or network.
[0073] According to an embodiment of the present disclosure, a computer-readable storage medium may also be provided, wherein when the instructions in the computer-readable storage medium are run by at least one processor, the at least one processor is caused to execute the image segmentation method of the embodiment of the present disclosure. Examples of the computer-readable storage medium here include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device, the any other device being configured to store a computer program and any associated data, data files and data structures in a non-transitory manner and provide the computer program and any associated data, data files and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium may run in an environment deployed in computer devices such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system such that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.
[0074] According to an embodiment of the present disclosure, a computer program product is provided, including computer instructions which, when executed by a processor, implement the image segmentation method of the embodiment of the present disclosure.
[0075] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only to be regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0076] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.
Claims
1. An image segmentation method, characterized in that, comprising: Inputting the image to be processed into a first image feature extraction network to obtain a first image feature; Inputting the preset reference information, the first interaction information, and the image to be processed into a second image feature extraction network to obtain a second image feature, wherein the amount of information of the image features extracted from the image by the second image feature extraction network is less than that of the image features extracted from the image by the first image feature extraction network, and the first interaction information is position information for indicating the object to be segmented in the image to be processed; Based on the first image feature and the second image feature, obtaining a first target mask for the object to be segmented; Performing segmentation processing on the image to be processed based on the first target mask to obtain a first segmentation result for the object to be segmented; Wherein, in the case that the first segmentation result does not meet the preset requirements, the image segmentation method further comprises: Inputting the first target mask, the second interaction information, and the image to be processed into the second image feature extraction network to obtain a third image feature, wherein the second interaction information is position information for indicating the object to be segmented in the image to be processed; Based on the first image feature and the third image feature, obtaining a second target mask for the object to be segmented; Performing segmentation processing on the image to be processed based on the second target mask to obtain a second segmentation result for the object to be segmented.
2. The image segmentation method according to claim 1, characterized in that, The obtaining a first target mask for the object to be segmented based on the first image feature and the second image feature includes: Fusing the first image feature and the second image feature to obtain a fused image feature; Inputting the fused image feature into an image mask extraction network to obtain the first target mask.
3. The image segmentation method according to claim 2, characterized in that, The fusing the first image feature and the second image feature to obtain a fused image feature includes: Stitching the first image feature and the second image feature to obtain a stitched feature; Inputting the stitched feature into a residual network to obtain the fused image feature.
4. The image segmentation method according to claim 1, characterized in that, The first interaction information includes forward interaction information and / or reverse interaction information, wherein the forward interaction information is used to indicate the region where the object to be segmented is located in the image to be processed, and the reverse interaction information is used to indicate the region outside the region where the object to be segmented is located in the image to be processed.
5. The image segmentation method according to claim 1, characterized in that, The second interaction information is determined based on the first segmentation result and the second interaction information includes forward interaction information and / or reverse interaction information, wherein the forward interaction information is used to indicate the region where the object to be segmented is located in the image to be processed, and the reverse interaction information is used to indicate the region outside the region where the object to be segmented is located in the image to be processed.
6. An image segmentation device, characterized in that, it includes: A first image feature acquisition unit configured to input an image to be processed into a first image feature extraction network to obtain first image features; A second image feature acquisition unit configured to input preset reference information, first interaction information, and the image to be processed into a second image feature extraction network to obtain second image features, where the amount of information of the image features extracted from the image by the second image feature extraction network is less than the amount of information of the image features extracted from the image by the first image feature extraction network, and the first interaction information is position information for indicating the object to be segmented in the image to be processed; A mask acquisition unit configured to obtain a first target mask for the object to be segmented based on the first image features and the second image features; A segmentation unit configured to perform segmentation processing on the image to be processed based on the first target mask to obtain a first segmentation result for the object to be segmented; wherein, in the case where the first segmentation result does not meet the preset requirements, the second image feature acquisition unit is further configured to input the first target mask, second interaction information, and the image to be processed into the second image feature extraction network to obtain third image features, where the second interaction information is used to indicate the position information of the object to be segmented in the image to be processed; the mask acquisition unit is further configured to obtain a second target mask for the object to be segmented based on the first image features and the third image features; the segmentation unit is further configured to perform segmentation processing on the image to be processed based on the second target mask to obtain a second segmentation result for the object to be segmented.
7. The image segmentation device according to claim 6, characterized in that, the mask acquisition unit is further configured to fuse the first image features and the second image features to obtain fused image features; input the fused image features into an image mask extraction network to obtain the first target mask.
8. The image segmentation device according to claim 7, characterized in that, the mask acquisition unit is further configured to splice the first image features and the second image features to obtain spliced features; input the spliced features into a residual network to obtain the fused image features.
9. The image segmentation device according to claim 6, characterized in that, the first interaction information includes forward interaction information and / or reverse interaction information, where the forward interaction information is used to indicate the area where the object to be segmented is located in the image to be processed, and the reverse interaction information is used to indicate the area outside the area where the object to be segmented is located in the image to be processed.
10. The image segmentation device according to claim 6, characterized in that, The second interaction information is determined based on the first segmentation result, and the second interaction information includes positive interaction information and / or negative interaction information, where the positive interaction information is used to indicate the region where the object to be segmented is located in the to-be-processed image, and the negative interaction information is used to indicate the region outside the region where the object to be segmented is located in the to-be-processed image.
11. An electronic device, characterized in that it includes: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the image segmentation method according to any one of claims 1 to 5.
12. A computer-readable storage medium, characterized in that when the instructions in the computer-readable storage medium are run by at least one processor, the at least one processor is caused to execute the image segmentation method according to any one of claims 1 to 5.
13. A computer program product comprising computer instructions, characterized in that when the computer instructions are executed by a processor, the image segmentation method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Integrated interactive image segmentation
US20210312635A1