Intelligent mask generation method, device and equipment for interactive image segmentation
By acquiring user interaction information for segmentation constraint signal transformation, feature fusion, and edge fusion, the accuracy and efficiency issues of image segmentation and mask generation technology in complex scenes are solved, achieving efficient and accurate mask generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA AI MEDIA&ENTERTAINMENT TECH CO LTD
- Filing Date
- 2025-12-29
- Publication Date
- 2026-05-05
AI Technical Summary
Existing image segmentation and mask generation technologies have low segmentation accuracy in complex scenes, cannot handle situations where the target and background are similar or occluded, and traditional interactive tools are cumbersome to operate, lack semantic understanding capabilities, are difficult to handle complex description requirements, and have rough mask edges, requiring manual image retouching.
By acquiring the initial image segmentation interaction information input by the user, segmentation constraint signal transformation, feature fusion, preliminary segmentation and edge fusion are performed. Combined with multimodal semantic parsing and dynamic segmentation adjustment, an accurate target mask is generated.
It improves the efficiency and accuracy of mask generation, enhances multimodal understanding capabilities, increases segmentation accuracy by 15-20%, improves interaction efficiency by 2-3 times, smooths mask edges, and reduces manual retouching workload by 90%.
Smart Images

Figure CN121982301A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image editing and processing technology, and in particular to an interactive image segmentation intelligent mask generation method, apparatus, and device. Background Technology
[0002] Currently, in image segmentation and mask generation technologies, segmentation accuracy determines image processing efficiency. Existing image segmentation and mask generation technologies have the following drawbacks: Automatic segmentation methods are poorly adaptable to complex scenes. When the target and background are similar, or when there is occlusion or blurred edges, the segmentation accuracy is low, and the generated mask has problems of missing or over-detecting, making it unsuitable for direct and precise editing.
[0003] Traditional interactive tools rely on a single interaction method (such as only supporting point selection or borders), requiring users to input multiple times to get close to the target, making the operation cumbersome; and they lack semantic understanding capabilities, unable to specify the segmentation target through text description (such as "segment all circular objects in the image"). They struggle to handle complex requirements with spatial location and attribute descriptions, such as "segment the third white cat on the left," and the alignment accuracy between the segmentation results and the user's intent is insufficient.
[0004] Existing methods generate masks with rough edges, resulting in poor segmentation of fine-grained or transparent targets such as hair, glass, and smoke. This necessitates manual retouching, leading to low efficiency. Summary of the Invention
[0005] The technical problem to be solved by this invention is to provide an interactive image segmentation intelligent mask generation method, apparatus, and device. This improves the efficiency, accuracy, and multimodal understanding capabilities of intelligent mask generation.
[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: An interactive image segmentation intelligent mask generation method includes: Obtain initial image segmentation interaction information input by the user; The initial image segmentation interaction information is subjected to segmentation constraint signal conversion processing to obtain the target image segmentation information; The target image segmentation information region constraint mask is fused to obtain a fused feature map; The fused feature map is initially segmented to obtain an initial mask; The initial mask is segmented based on the secondary interaction information to obtain the intermediate mask; The intermediate mask is then subjected to edge blending processing to obtain the target mask.
[0007] Optionally, the initial image segmentation interaction information is subjected to segmentation constraint signal transformation processing to obtain target image segmentation information, including: Based on the text instruction, semantic tags and attribute weights are extracted to obtain the target text instruction; The selected points are weighted and diffused to obtain the target selected points; The confidence level of the border data is checked to obtain the target border data.
[0008] Optionally, feature fusion is performed on the constraint mask of the target image segmentation information region to obtain a fused feature map, including: The feature weights are obtained based on the importance of each feature in the target image segmentation information; Based on the feature weights, a fused feature map is obtained.
[0009] Optionally, the fused feature map is initially segmented to obtain an initial mask, including: The bounding box of the fused feature map is located to obtain the guiding mask; The guiding mask is combined with the image visual features to obtain the initial mask.
[0010] Optionally, the initial mask is segmented based on the secondary interaction information to obtain an intermediate mask, including: Based on the secondary interaction information, the initial mask is evaluated for deviation, and the deviation evaluation result is obtained; Based on the deviation assessment results, a deviation heatmap is obtained; The initial mask is adjusted based on the deviation heatmap to obtain the intermediate mask.
[0011] Optionally, edge blending processing is performed on the intermediate mask to obtain the target mask, including: The edge gradient is obtained based on the horizontal and vertical gradients; The intermediate mask is edge-blending processed according to the edge gradient to obtain the target mask.
[0012] Optionally, performing edge blending processing on the intermediate mask to obtain the target mask further includes: Based on the semantic hierarchy of the target text instruction, the intermediate mask is split into layers to obtain the target mask.
[0013] Embodiments of the present invention also provide an interactive image segmentation intelligent mask generation apparatus, comprising: The acquisition module is used to acquire the initial image segmentation interaction information input by the user; The processing module is used to perform segmentation constraint signal conversion processing on the initial image segmentation interaction information to obtain target image segmentation information; perform feature fusion on the region constraint mask of the target image segmentation information to obtain a fused feature map; perform preliminary segmentation on the fused feature map to obtain an initial mask; segment the initial mask according to the secondary interaction information to obtain an intermediate mask; and perform edge fusion processing on the intermediate mask to obtain a target mask.
[0014] Embodiments of the present invention also provide a computing device, including: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the interactive image segmentation intelligent mask generation method of the present invention.
[0015] Embodiments of the present invention also provide a computer-readable storage medium storing a program that, when executed by a processor, implements the intelligent mask generation method for interactive image segmentation described in this invention.
[0016] The above-described technical solution of the present invention has at least the following technical effects: The intelligent mask generation method for interactive image segmentation described above in this invention obtains initial image segmentation interaction information input by the user, performs segmentation constraint signal conversion processing on the initial image segmentation interaction information to obtain target image segmentation information; performs feature fusion on the region constraint mask of the target image segmentation information to obtain a fused feature map; performs preliminary segmentation on the fused feature map to obtain an initial mask; segments the initial mask according to secondary interaction information to obtain an intermediate mask; and performs edge fusion processing on the intermediate mask to obtain the target mask. This improves the efficiency, accuracy, and multimodal understanding capability of intelligent mask generation. Attached Figure Description
[0017] Figure 1 This is an overall schematic diagram of the intelligent mask generation method for interactive image segmentation according to the present invention; Figure 2 This is a flowchart illustrating the intelligent mask generation method for interactive image segmentation according to the present invention. Figure 3 This is a schematic diagram of the system structure of the interactive image segmentation intelligent mask generation method of the present invention; Figure 4 This is a schematic diagram of the intelligent mask generation device for interactive image segmentation according to the present invention. Detailed Implementation
[0018] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.
[0019] like Figure 1 As shown, embodiments of the present invention propose an interactive image segmentation intelligent mask generation method, comprising: Step S1: Obtain the initial image segmentation interaction information input by the user; Step S2: Perform segmentation constraint signal conversion processing on the initial image segmentation interaction information to obtain target image segmentation information; Step S3: Perform feature fusion on the constraint mask of the target image segmentation information region to obtain a fused feature map; Step S4: Perform preliminary segmentation on the fused feature map to obtain an initial mask; Step S5: Segment the initial mask according to the secondary interaction information to obtain the intermediate mask; Step S6: Perform edge blending processing on the intermediate mask to obtain the target mask.
[0020] In this embodiment, as Figure 1 As shown, in the intelligent mask generation method for interactive image segmentation, firstly, the initial image segmentation interaction information input by the user is obtained. The initial image segmentation interaction information may include a variety of user input instructions, specifically including at least one of text instructions (natural language description of the target), point selection marks (foreground / background points), and bounding box data (approximate range of the target). Then, the initial image segmentation interaction information is processed by segmentation constraint signal transformation according to different types to obtain target image segmentation information; feature fusion is performed on the region constraint mask of the target image segmentation information to obtain a fused feature map; the fused feature map is initially segmented based on visual features to obtain an initial mask; the initial mask is further segmented according to the secondary interaction information to obtain an intermediate mask; the intermediate mask is edge-blending processed to obtain the target mask. Through multimodal semantic parsing and dynamic segmentation adjustment, an accurate target mask is generated, achieving accurate embedding of image content.
[0021] In an optional embodiment of the present invention, step S2 involves performing segmentation constraint signal conversion processing on the initial image segmentation interaction information to obtain target image segmentation information, including: Step S21: Extract semantic tags and attribute weights from the text instructions in the initial image segmentation interaction information to obtain the target text instructions; Step S22: Perform weighted diffusion on the point selection marks in the initial image segmentation interaction information to obtain the target point selection marks; Step S23: Perform a confidence check on the border data in the initial image segmentation interaction information to obtain the target border data.
[0022] In this embodiment, firstly, semantic tags and attribute weights are extracted based on the text instruction to obtain the target text instruction; the text instruction description is parsed to output semantic tag embeddings and attribute weights; attribute words (such as "blue" and "shirt") in the text are identified, attribute confidence scores are output, and attribute weights are calculated using the attribute confidence scores. Specifically... ,in For attribute weights, For attribute confidence, This is the attribute enhancement coefficient; Then, semantic labels are bound to attribute weights to generate semantic constraint pairs. Where L is the target category, such as "person", the target text instruction is obtained; for example, for the text instruction "person wearing a blue shirt" entered by the user, attributes such as "blue" and "shirt" are extracted and assigned weights. The weight of "blue" can be set to 0.5 and the weight of "shirt" can be set to 0.8. When segmenting, text attributes are matched first to improve the accuracy of attribute matching. The selected point markers are weighted and diffused to obtain the target selected point markers. Through weighted diffusion, the constraint range of a single point selection is expanded from the "pixel level" to the "local region level," better reflecting the user's true intent regarding the "foreground / background region." The user's selected coordinates are used as the basis for the target selection. For example, select type as t , t =1 represents the foreground. t Using -1 as the background, a weighted diffusion map is generated to simulate the local influence range of the selected points. The formula for generating the weight map is:
[0023] in, This represents a weighted graph, where t represents the selection type. Indicates the diffusion radius; diffusion radius It can adaptively adjust according to image resolution; for example, when the image resolution is greater than 2K. A value of 8 is used when the image resolution is less than 2K. The value is 3; By applying weighted diffusion to the selected markers, segmentation errors caused by "pixel deviation" in single-point selection can be avoided. For example, when selecting the edge of a hair strand, the local area is considered as the foreground.
[0024] The border data is subjected to a confidence check to obtain the target border data; based on the width and height values in the border data, a region confidence score is generated; specifically, the formula for calculating the region confidence score is:
[0025] in, Indicates the regional confidence level. Represents the distance from pixel (x, y) to the center of the bounding box. Normalized distance, w represents the width of the bounding box, and h represents the height of the bounding box; Calculate the average confidence score of pixels within the bounding box, and simultaneously evaluate the confidence score distribution in neighboring regions outside the bounding box. If the average confidence score within the bounding box is below a threshold, and there are high-confidence pixels (e.g., exceeding 0.8) in a certain region outside the bounding box, then it is determined that this region may belong to the true boundary of the target. Then, using the high-confidence pixels as seed points, under the condition of satisfying the confidence decay constraint (e.g., the confidence score decreases by no more than 0.1 for each pixel extended), the target boundary is gradually extended towards the high-confidence region.
[0026] The constraint strength at the edge of the border is lower than that in the center area to avoid segmentation deviation caused by the border being drawn off-center. Through the region confidence check, even if the border is drawn off-center, if it does not completely enclose the target, the segmentation can still be corrected to the true boundary of the target.
[0027] In an optional embodiment of the present invention, step S3 involves performing feature fusion on the constraint mask of the target image segmentation information region to obtain a fused feature map, including: Step S31: Obtain feature weights based on the importance of each feature in the target image segmentation information; Step S32: Obtain the fused feature map based on the feature weights.
[0028] In this embodiment, the features in the target image segmentation information include: visual features, input semantic features, and localization features. The formula for calculating the weight of the visual features is as follows:
[0029] in, MI represents the mutual information between visual features, representing the visual feature weights. Represents visual feature maps. Indicates semantic tags, Represents bounding box features; The formula for calculating semantic feature weights is:
[0030] in, Represents semantic feature weights, Indicates semantic attribute weights; The formula for calculating the location feature weight is:
[0031] in, Indicates the weight of the localization feature. Represents the weights of visual features. Indicates semantic feature weights; Then, based on the feature weights, the fused feature map is obtained; the feature fusion formula is:
[0032] in, Represents the fused feature map. Represents visual feature maps. Indicates semantic tags, Represents bounding box features, Indicates the weight of the localization feature. Represents the weights of visual features. Represents semantic feature weights, This represents an H×W dimension matrix of all ones. Indicates broadcast operation; Feature weights enable feature fusion to better fit the input scenario, laying a high-precision foundation for subsequent segmentation.
[0033] In an optional embodiment of the present invention, step S4, which involves preliminary segmentation of the fused feature map to obtain an initial mask, includes: Step S41: Locate the bounding box in the fused feature map to obtain the guide mask; Step S42: Combine the guide mask with image visual features to obtain an initial mask.
[0034] In this embodiment, bounding box localization is performed on the fused feature map, and a guiding mask is obtained based on the bounding box; target features are located in the fused feature map to obtain bounding box coordinates. B = Regress ( ) Where B represents the bounding box, and Regress represents the bounding box regression function. Represents the fused feature map; Based on this bounding box, a region is defined, with pixel values within the region set to 1 and pixel values outside the region set to 0, thereby constructing a binarized guiding mask. This will serve as a guiding principle for subsequent processing.
[0035] Then, the guiding mask is combined element-by-element with the visual feature map to obtain the initial mask; the expression for the initial mask is:
[0036] in, Indicates the initial mask. Represents visual feature maps. Indicates the boot mask; By segmenting the image within the bounding box, unnecessary calculations of irrelevant regions can be avoided, thus improving efficiency.
[0037] In an optional embodiment of the present invention, step S5, segmenting the initial mask according to the secondary interaction information to obtain an intermediate mask, includes: Step S51: Based on the secondary interaction information, perform a deviation evaluation on the initial mask to obtain the deviation evaluation result; Step S52: Obtain the deviation heatmap based on the deviation evaluation results; Step S53: Adjust the initial mask according to the deviation heatmap to obtain the intermediate mask.
[0038] In this embodiment, firstly, based on the secondary interaction information, a deviation assessment is performed on the initial mask to obtain the deviation assessment result; then, through the calculation of the deviation index, the consistency between the initial mask and the interactive input is comprehensively evaluated; the selection deviation rate refers to the proportion of misclassified selected pixels to the total selected pixels, and the formula for calculating the selection deviation rate is:
[0039] in, Indicates the selection deviation rate. Represents the selected pixel set. This indicates that the type label has been selected. Indicates the initial mask; Border overlap refers to the intersection-union ratio of the mask and the border. The formula for calculating border overlap is:
[0040] in, Indicates the degree of border overlap. B represents the initial mask; Semantic matching degree refers to the cosine similarity between mask features and text semantic features. The formula for calculating semantic matching degree is:
[0041] in, Indicates semantic matching degree. Indicates the initial mask Global average pooling characteristics, Represents semantic tags.
[0042] Then, the overall deviation is calculated using the point selection deviation rate, border overlap, and semantic matching degree. The formula for calculating the overall deviation is:
[0043] Where D represents the overall deviation, Indicates the selection deviation rate. Indicates the weight of the selection deviation rate. Indicates the degree of border overlap. Indicates the weight of border overlap. Indicates semantic matching degree. Indicates the semantic matching degree weight; Then, based on the deviation assessment results (i.e., the overall deviation), a deviation heatmap is obtained; For selected areas with high selection deviation rates, border edge areas with low border overlap, and semantic mismatch areas with low semantic matching, a high deviation value (0~1) is assigned based on the comprehensive deviation to obtain a deviation heatmap. Finally, the initial mask is adjusted based on the deviation heatmap to obtain an intermediate mask; the deviation heatmap is then used as an attention bias and superimposed on the cross-layer attention matrix to obtain an updated attention matrix, thereby adjusting the initial mask to obtain the intermediate mask; specifically, the updated attention matrix is as follows:
[0044] in, This represents the updated attention matrix. Represents the basic attention matrix. For heatmap weights, This represents a deviation heatmap. Indicates the number of attention channels; The initial mask is adjusted, and the deviation area is optimized in a fine-grained manner to reduce the calculation time and improve the segmentation accuracy of the deviation area.
[0045] In an optional embodiment of the present invention, step S6, performing edge blending processing on the intermediate mask to obtain the target mask, includes: Step S61: Obtain the edge gradient based on the horizontal and vertical gradients; Step S62: Perform edge blending processing on the intermediate mask according to the edge gradient to obtain the target mask.
[0046] In this embodiment, multi-scale edge fusion is used to process the intermediate mask using gradients at multiple scales, enabling the simultaneous capture of coarse outlines (such as human body outlines) and fine edges (such as hair strands and fabric textures). Edge gradients are obtained based on horizontal and vertical gradients. Specifically, the horizontal template is:
[0047] The horizontal template is:
[0048] The horizontal gradient is obtained through convolution operations. and vertical gradient Then the edge gradient is:
[0049] Weights are assigned based on the magnitude of the edge gradient; the larger the gradient (the sharper the edge), the higher the weight. The formula for calculating the gradient weight is:
[0050] in, Represents gradient weights. For edge map gradient magnitude, This is the maximum value of all gradients; Multi-scale edge fusion can improve the detection rate of fine edges (such as hair strands), enhance the integrity of coarse contours, and avoid the phenomenon of losing details or adding noise in single-scale detection.
[0051] In an optional embodiment of the present invention, step S6, which involves edge blending of the intermediate mask to obtain the target mask, further includes: Step S63: Based on the semantic hierarchy of the target text instruction, the intermediate mask is split into layers to obtain the target mask.
[0052] In this embodiment, when outputting the target mask, the mask format is automatically adjusted according to the target editing tool. If it is Photoshop, a PSD format mask with an alpha channel is output; if it is Figma, an SVG path format mask is output; if it is a deep learning model, a binary mask is output. Semantic segmentation is performed on the text (e.g., "a person wearing a blue shirt, with a red tie and a silver watch") to obtain a sub-target set O = {O1 = shirt, O2 = tie, O3 = watch}. For each sub-target... O i It segments sub-masks from the final mask through semantic feature matching. Each sub-mask is an independent layer, labeled with the name of the sub-target, and supports individual adjustments by the user, improving the user's editing efficiency for multiple sub-targets. At the same time, it automatically extracts several foreground images for the user to customize and combine the set of objects to be extracted.
[0053] The solution of this invention is based on a multimodal fusion architecture, which realizes accurate association between text semantics and visual features and supports end-to-end guidance of "semantic description-target segmentation". It integrates multiple interaction methods such as text, point selection, and borders, and transforms different inputs into unified segmentation constraints through standardized modules, thereby improving the flexibility of user intent expression. It introduces a dynamic feedback adjustment mechanism, which automatically triggers secondary segmentation of key areas based on the deviation analysis between the initial segmentation results and user interaction, thereby improving the segmentation accuracy in complex scenarios. It designs an edge refinement process for fine-grained targets (such as hair strands and transparent objects) to ensure that the mask edges closely match the real contours of the targets.
[0054] The above-mentioned solution of the present invention significantly improves the segmentation accuracy through two segmentation steps: in complex background scenes (such as when the similarity between the target and the background is >80%), the average intersection-union ratio (IoU) between the mask and the real outline of the target reaches 92%, which is 15-20% higher than the existing model; the segmentation accuracy of fine-grained targets (such as hair strands) is improved by more than 30%.
[0055] By comparing with secondary interaction information, interaction efficiency is improved: through multimodal guidance, the average number of user interactions is reduced to 2-3 times (traditional tools require 5-8 times), and the time to generate accurate masks is shortened to 1-2 seconds.
[0056] Enhanced semantic understanding capabilities: Supports text instruction segmentation with attributes and location, with an instruction execution accuracy of up to 90%, far exceeding tools that only support visual interaction (accuracy <65%).
[0057] The mask has wider applicability: the generated mask has smooth edges and complete details, and can be directly used in professional scenarios such as image content embedding and special effects compositing, reducing manual retouching workload by more than 90%.
[0058] like Figure 4 As shown, embodiments of the present invention also provide an interactive image segmentation intelligent mask generation device 40, comprising: The acquisition module 41 is used to acquire the initial image segmentation interaction information input by the user; Processing module 42 is used to perform segmentation constraint signal conversion processing on the initial image segmentation interaction information to obtain target image segmentation information; perform feature fusion on the region constraint mask of the target image segmentation information to obtain a fused feature map; perform preliminary segmentation on the fused feature map to obtain an initial mask; segment the initial mask according to the secondary interaction information to obtain an intermediate mask; and perform edge fusion processing on the intermediate mask to obtain a target mask.
[0059] Optionally, the initial image segmentation interaction information is subjected to segmentation constraint signal transformation processing to obtain target image segmentation information, including: Based on the text instruction, semantic tags and attribute weights are extracted to obtain the target text instruction; The selected points are weighted and diffused to obtain the target selected points; The confidence level of the border data is checked to obtain the target border data.
[0060] Optionally, feature fusion is performed on the constraint mask of the target image segmentation information region to obtain a fused feature map, including: The feature weights are obtained based on the importance of each feature in the target image segmentation information; Based on the feature weights, a fused feature map is obtained.
[0061] Optionally, the fused feature map is initially segmented to obtain an initial mask, including: The bounding box of the fused feature map is located to obtain the guiding mask; The guiding mask is combined with the image visual features to obtain the initial mask.
[0062] Optionally, the initial mask is segmented based on the secondary interaction information to obtain an intermediate mask, including: Based on the secondary interaction information, the initial mask is evaluated for deviation, and the deviation evaluation result is obtained; Based on the deviation assessment results, a deviation heatmap is obtained; The initial mask is adjusted based on the deviation heatmap to obtain the intermediate mask.
[0063] Optionally, edge blending processing is performed on the intermediate mask to obtain the target mask, including: The edge gradient is obtained based on the horizontal and vertical gradients; The intermediate mask is edge-blending processed according to the edge gradient to obtain the target mask.
[0064] Optionally, performing edge blending processing on the intermediate mask to obtain the target mask further includes: Based on the semantic hierarchy of the target text instruction, the intermediate mask is split into layers to obtain the target mask.
[0065] It should be noted that all implementation methods in the above method embodiments are applicable to the embodiments of this device and can achieve the same technical effect.
[0066] Embodiments of the present invention also provide a computing device, including: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the interactive image segmentation intelligent mask generation method of the present invention. All implementations in the above method embodiments are applicable to the embodiments of this computing device and can achieve the same technical effects.
[0067] Embodiments of the present invention also provide a computer-readable storage medium storing a program that, when executed by a processor, implements the intelligent mask generation method for interactive image segmentation described in this invention. All implementations in the above method embodiments are applicable to the embodiments of this computer-readable storage medium and achieve the same technical effects.
[0068] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0069] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0070] In the embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0071] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0072] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0073] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0074] Furthermore, it should be noted that in the apparatus and method of the present invention, it is obvious that the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered equivalent solutions of the present invention. Moreover, the steps performing the above series of processes can naturally be executed in the order described, but are not necessarily required to be executed in chronological order; some steps can be executed in parallel or independently of each other. Those skilled in the art will understand that all or any step or component of the method and apparatus of the present invention can be implemented in any computing device (including processors, storage media, etc.) or network of computing devices, in hardware, firmware, software, or a combination thereof. This is something that those skilled in the art can achieve by using their basic programming skills after reading the description of the present invention.
[0075] Therefore, the object of the present invention can also be achieved by running a program or a set of programs on any computing device. The computing device can be a known general-purpose device. Therefore, the object of the present invention can also be achieved simply by providing a program product containing program code for implementing the method or apparatus. That is, such a program product also constitutes the present invention, and the storage medium storing such a program product also constitutes the present invention. Obviously, the storage medium can be any known storage medium or any storage medium developed in the future. It should also be noted that in the apparatus and method of the present invention, it is obvious that the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered equivalent to the present invention. Furthermore, the steps for performing the above series of processes can naturally be performed in the order described, but are not necessarily required to be performed in chronological order. Some steps can be performed in parallel or independently of each other.
[0076] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A smart mask generation method for interactive image segmentation, characterized in that, include: Obtain initial image segmentation interaction information input by the user; The initial image segmentation interaction information is subjected to segmentation constraint signal conversion processing to obtain the target image segmentation information; The target image segmentation information region constraint mask is fused to obtain a fused feature map; The fused feature map is initially segmented to obtain an initial mask; The initial mask is segmented based on the secondary interaction information to obtain the intermediate mask; The intermediate mask is then subjected to edge blending processing to obtain the target mask.
2. The intelligent mask generation method for interactive image segmentation according to claim 1, characterized in that, The initial image segmentation interaction information is subjected to segmentation constraint signal transformation processing to obtain target image segmentation information, including: Based on the text instruction, semantic tags and attribute weights are extracted to obtain the target text instruction; The selected points are weighted and diffused to obtain the target selected points; The confidence level of the border data is checked to obtain the target border data.
3. The intelligent mask generation method for interactive image segmentation according to claim 1, characterized in that, The target image segmentation information region constraint mask is subjected to feature fusion to obtain a fused feature map, including: The feature weights are obtained based on the importance of each feature in the target image segmentation information; Based on the feature weights, a fused feature map is obtained.
4. The intelligent mask generation method for interactive image segmentation according to claim 1, characterized in that, The fused feature map is initially segmented to obtain an initial mask, including: The bounding box of the fused feature map is located to obtain the guiding mask; The guiding mask is combined with the image visual features to obtain the initial mask.
5. The intelligent mask generation method for interactive image segmentation according to claim 1, characterized in that, The initial mask is segmented based on the secondary interaction information to obtain an intermediate mask, including: Based on the secondary interaction information, the initial mask is evaluated for deviation, and the deviation evaluation result is obtained; Based on the deviation assessment results, a deviation heatmap is obtained; The initial mask is adjusted based on the deviation heatmap to obtain the intermediate mask.
6. The intelligent mask generation method for interactive image segmentation according to claim 1, characterized in that, The intermediate mask is subjected to edge blending processing to obtain the target mask, including: The edge gradient is obtained based on the horizontal and vertical gradients; The intermediate mask is edge-blending processed according to the edge gradient to obtain the target mask.
7. The intelligent mask generation method for interactive image segmentation according to claim 1, characterized in that, The process of performing edge blending on the intermediate mask to obtain the target mask also includes: Based on the semantic hierarchy of the target text instruction, the intermediate mask is split into layers to obtain the target mask.
8. An interactive image segmentation intelligent mask generation device, characterized in that, include: The acquisition module is used to acquire the initial image segmentation interaction information input by the user; The processing module is used to perform segmentation constraint signal conversion processing on the initial image segmentation interaction information to obtain target image segmentation information; The target image segmentation information region constraint mask is fused to obtain a fused feature map; the fused feature map is then preliminarily segmented to obtain an initial mask. The initial mask is segmented based on the secondary interaction information to obtain an intermediate mask; the intermediate mask is then subjected to edge blending processing to obtain the target mask.
9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.