An image segmentation method, device, apparatus and medium

CN117292129BActive Publication Date: 2026-09-25SHANGHAI FUYA INTELLIGENT TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311224285.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-21
Publication Date
2026-09-25
Estimated Expiration
2043-09-21

AI Technical Summary

Technical Problem

[0004]然而该标注方法具有以下显著缺点:首先,传统人工标注对于像素区域较小的目标或形状较为不规则的目标,需要耗费较长的时间逐一进行标注,标注效率较低,标注成本较高;且对于目标边界像素的划定只能采用多条直线进行分割,会导致部分非直线边界像素的标注错误,标签的分割精度受到客观条件限制;其次,在当前深度学习模型的发展趋势下,模型结构日益复杂及模型参数量也在爆炸式增长,因此模型训练所需的数据量也在变得非常庞大(目前部分大模型训练数据中掩码数量已为亿级),传统的人工标注方法已经无法满足快速增长的模型训练数据规模的需求

Benefits of technology

[0021]本发明实施例的技术方案,通过获取待分割图像及筛选颜色条件,确定待分割图像中每个像素坐标点所对应的数字阵列信息;根据预设特征颜色筛选区间对数字阵列信息进行筛选,确定待分割图像中满足颜色筛选条件的提词候选点集;根据提词候选点集、预设提词分割模型及预设像素占比阈值,确定满足筛选颜色条件的目标对象;根据目标对象对待分割图像进行分割。通过预设特征颜色筛选区间对待分割图像进行筛选,得到提词候选点,进而通过预设提词分割模型得到目标分割掩码,再经过预设像素占比阈值进行过滤,得到目标对象。实现了对要分割的目标对象的自动确定,保证了目标对象确定的准确性,提高了图像分割的效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117292129B_ABST
    Figure CN117292129B_ABST
Patent Text Reader

Abstract

The application discloses an image segmentation method, device, equipment and medium. The method comprises the following steps: acquiring an image to be segmented and screening color conditions, and determining digital array information corresponding to each pixel coordinate point in the image to be segmented; screening the digital array information according to a preset characteristic color screening interval, and determining a word suggestion candidate point set in the image to be segmented that meets the color screening conditions; determining a target object that meets the screening color conditions according to the word suggestion candidate point set, a preset word segmentation model and a preset pixel proportion threshold; and segmenting the image to be segmented according to the target object. The image to be segmented is screened through the preset characteristic color screening interval to obtain a word suggestion candidate point, then a target segmentation mask is obtained through the preset word segmentation model, and the target object is obtained through filtering of the preset pixel proportion threshold. The automatic determination of the target object to be segmented is realized, the accuracy of the target object determination is ensured, and the efficiency of the image segmentation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to an image segmentation method, apparatus, device, and medium. Background Technology

[0002] In the field of computer vision, the segmentation of objects or backgrounds in images (pixel-level classification) has always been an important research direction, mainly due to its wide application in various scenarios. For example, in the field of drone inspection, image segmentation of abnormal targets can accurately locate alarm areas; in augmented reality scenarios, image segmentation of actual targets can be used to calculate and obtain more accurate relative poses, thereby obtaining a better and more realistic visual experience; and so on.

[0003] Currently, the main data source for image segmentation tasks using deep learning models is manual polygon annotation using annotation tools.

[0004] However, this annotation method has the following significant drawbacks: First, traditional manual annotation requires a long time to annotate targets with small pixel areas or irregular shapes one by one, resulting in low annotation efficiency and high annotation costs. Moreover, the delineation of target boundary pixels can only be done by using multiple straight lines, which can lead to incorrect annotation of some non-straight boundary pixels, and the segmentation accuracy of the labels is limited by objective conditions. Second, with the current development trend of deep learning models, the model structure is becoming increasingly complex and the number of model parameters is also growing explosively. Therefore, the amount of data required for model training is also becoming very large (currently, the number of masks in the training data of some large models has reached hundreds of millions). Traditional manual annotation methods can no longer meet the needs of the rapidly growing scale of model training data. Summary of the Invention

[0005] This invention provides an image segmentation method, apparatus, device, and medium to achieve automatic determination of the target object for image segmentation.

[0006] According to a first aspect of the present invention, an image segmentation method is provided, comprising:

[0007] Obtain the image to be segmented and the color filtering conditions, and determine the digital array information corresponding to each pixel coordinate point in the image to be segmented;

[0008] The digital array information is filtered according to a preset feature color filtering range to determine the set of candidate points for word retrieval in the image to be segmented that meet the color filtering conditions.

[0009] Based on the candidate point set, the preset word prompting segmentation model, and the preset pixel ratio threshold, the target object that meets the filtering color condition is determined;

[0010] The image to be segmented is segmented based on the target object.

[0011] According to a second aspect of the present invention, an image segmentation apparatus is provided, comprising:

[0012] The information acquisition module is used to acquire the image to be segmented and the color filtering conditions, and to determine the digital array information corresponding to each pixel coordinate point in the image to be segmented;

[0013] The first determining module is used to filter the digital array information according to a preset feature color filtering range to determine the set of candidate points for word selection in the image to be segmented that meet the color filtering conditions.

[0014] The second determining module is used to determine the target object that meets the filtering color conditions based on the word prompting candidate point set, the preset word prompting segmentation model and the preset pixel ratio threshold.

[0015] The image segmentation module is used to segment the image to be segmented based on the target object.

[0016] According to a third aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0017] At least one processor; and

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the image segmentation method according to any embodiment of the present invention.

[0020] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the image segmentation method according to any embodiment of the present invention.

[0021] The technical solution of this invention involves acquiring the image to be segmented and setting color selection criteria, determining the digital array information corresponding to each pixel coordinate point in the image to be segmented, filtering the digital array information according to a preset feature color selection interval, determining a set of candidate points in the image to be segmented that meet the color selection criteria, determining the target object that meets the color selection criteria based on the candidate point set, a preset word-selection segmentation model, and a preset pixel proportion threshold, and then segmenting the image to be segmented based on the target object. By filtering the image to be segmented through a preset feature color selection interval to obtain candidate points, obtaining the target segmentation mask through a preset word-selection segmentation model, and then filtering it through a preset pixel proportion threshold to obtain the target object, the invention achieves automatic determination of the target object to be segmented, ensuring the accuracy of target object determination and improving the efficiency of image segmentation.

[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart of an image segmentation method provided according to Embodiment 1 of the present invention;

[0025] Figure 2 This is a flowchart of an image segmentation method provided according to Embodiment 2 of the present invention;

[0026] Figure 3 This is a flowchart of an image segmentation method provided according to Embodiment 2 of the present invention;

[0027] Figure 4 This is a schematic diagram of a word-prompting candidate point set in an image segmentation method according to Embodiment 2 of the present invention;

[0028] Figure 5 This is an example diagram of a target segmentation mask in an image segmentation method according to Embodiment 2 of the present invention;

[0029] Figure 6 This is an example diagram of the target segmentation mask set in an image segmentation method according to Embodiment 2 of the present invention;

[0030] Figure 7This is an example diagram of the target object in an image segmentation method provided according to Embodiment 2 of the present invention;

[0031] Figure 8 This is a schematic diagram of the structure of an image segmentation device according to Embodiment 3 of the present invention;

[0032] Figure 9 This is a schematic diagram of the structure of an electronic device that implements an embodiment of the present invention. Detailed Implementation

[0033] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0035] Example 1

[0036] Figure 1 The flowchart of an image segmentation method provided in Embodiment 1 of the present invention is applicable to the automatic annotation of image segmentation objects. This method can be executed by an image segmentation device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:

[0037] S110. Obtain the image to be segmented and the color filtering conditions, and determine the digital array information corresponding to each pixel coordinate point in the image to be segmented.

[0038] In this embodiment, the image to be segmented can be understood as the image from which the object of interest needs to be segmented. The color filtering condition can be understood as the color condition corresponding to the object of interest to be segmented; for example, if the object of interest is the green scaffolding covering the exterior of a building, then the color filtering condition would be green. Pixel coordinates can be understood as the pixel coordinates corresponding to each point in the image's pixel coordinate system. Digital array information can be understood as representing the image in the form of a three-channel digital array.

[0039] Specifically, the processor can obtain the image to be segmented through uploading or other means. Users can select objects of interest, such as by clicking with a mouse. The processor can obtain the filtering color conditions corresponding to the objects of interest. The processor can directly determine the digital array information corresponding to each pixel coordinate point in the RGB format of the image to be segmented. It can also convert the image to be segmented into the required format and obtain the required digital array information.

[0040] For example, the processor can read data from an image to be segmented. The read data is an H×W×C digital array in RGB (red, green, and blue color channels) format (H is the number of pixels in the height direction of the image; W is the number of pixels in the width direction of the image; C is the R, G, and B channels). Alternatively, the processor can convert the image data to other formats, such as converting the digital array to HSV (Hue, Saturation, Value) color gamut. The converted data format is also an H×W×C digital array (H is the number of pixels in the height direction of the image; W is the number of pixels in the width direction of the image; C is the H, S, and V channels). The following conversion method can be used:

[0041] HSV color gamut conversion formula

[0042] (R,G,B)=(R,G,B) / 255.0

[0043] Max = maX(R,G,B)

[0044] Min = min(R, G, B)

[0045]

[0046] S = ((Max - Min) * 2255)

[0047] V = Max * 255

[0048] Where R is the value of the red color channel in the image data to be segmented (range [0, 255]), G is the value of the green color channel in the image data to be segmented (range [0, 255]), B is the value of the blue color channel in the image data to be segmented (range [0, 255]), Max is the maximum value among R, G, and B, Min is the minimum value among R, G, and B, H is in the range [0, 180], S is in the range [0, 255], and V is in the range [0, 255].

[0049] S120. Filter the digital array information according to the preset feature color filtering range to determine the set of candidate points in the image to be segmented that meet the color filtering conditions.

[0050] In this embodiment, the preset feature color filtering range can be understood as a color filtering range corresponding to the color filtering conditions, used to filter out colors that match the color filtering conditions. The candidate point set can be understood as a set consisting of all the filtered points that match the color filtering conditions.

[0051] Specifically, the processor can filter the values ​​of each channel in the digital array information according to the preset feature color filtering interval, determine whether the value is within the preset feature color filtering interval, and take each point in the image to be segmented that is within the preset feature color filtering interval as a set of candidate points that meet the color filtering conditions.

[0052] S130. Based on the candidate point set for word prompting, the preset word prompting segmentation model, and the preset pixel ratio threshold, determine the target objects that meet the color selection criteria.

[0053] In this embodiment, the preset word-prompting segmentation model can be understood as a model used to extract the objects corresponding to the word-prompting candidate points, such as the Segment Anything Model (SAM). The preset pixel proportion threshold can be understood as a proportion threshold set to filter whether a target object meets the filtering color conditions. The target object can be understood as the target object in the image to be segmented that meets the filtering color conditions.

[0054] Specifically, the processor can input the candidate point set into a preset word-prompting segmentation model to determine the overall object to which the candidate points belong. Since the overall objects determined by the preset word-prompting segmentation model may have deviations, only a small portion of the colors meet the color selection criteria, but the overall objects they constitute do not belong to the target objects that meet the color selection criteria. Therefore, this part of the error objects can be found by setting a preset pixel proportion threshold. The proportion of all pixels included in each overall object and the proportion of pixels belonging to each overall object in the candidate point set can be compared with the preset pixel proportion threshold to determine whether each overall object is a target object that meets the color selection criteria, thus realizing the division of target objects in the image to be segmented.

[0055] S140. Segment the image to be segmented according to the target object.

[0056] Specifically, the processor can segment the image to be segmented according to the target object using a preset method.

[0057] The technical solution of this invention involves acquiring the image to be segmented and setting color selection criteria, determining the digital array information corresponding to each pixel coordinate point in the image to be segmented, filtering the digital array information according to a preset feature color selection interval, determining a set of candidate points in the image to be segmented that meet the color selection criteria, determining the target object that meets the color selection criteria based on the candidate point set, a preset word-selection segmentation model, and a preset pixel proportion threshold, and then segmenting the image to be segmented based on the target object. By filtering the image to be segmented through a preset feature color selection interval to obtain candidate points, obtaining the target segmentation mask through a preset word-selection segmentation model, and then filtering it through a preset pixel proportion threshold to obtain the target object, the invention achieves automatic determination of the target object to be segmented, ensuring the accuracy of target object determination and improving the efficiency of image segmentation.

[0058] Example 2

[0059] Figure 2 This is a flowchart of an image segmentation method provided in Embodiment 2 of the present invention. This embodiment is a further refinement based on the above embodiments, such as... Figure 2 As shown, the method includes:

[0060] S210. Obtain the image to be segmented and the color filtering conditions, and determine the digital array information corresponding to each pixel coordinate point in the image to be segmented.

[0061] S220. For each pixel coordinate point in the image to be segmented, determine the corresponding digital array in the digital array information.

[0062] Specifically, the processor can determine the digital array corresponding to each pixel coordinate point in the digital array information for each pixel coordinate point in the image to be segmented.

[0063] S230. Compare the channel values ​​of each channel in the digital array with the preset feature color filtering range.

[0064] In this embodiment, each channel can be understood as a channel in a digital array that represents different data.

[0065] Specifically, the processor can compare the channel values ​​of each channel in the digital array with the preset feature color filtering range in turn to determine whether the channel values ​​of each pixel coordinate point in the image to be segmented are within the preset feature color filtering range.

[0066] For example, taking a green screen as an example, in RGB data format, the channel values ​​of the preset characteristic color filtering range for green can be set to R<=0.5*max(R), G>=0.5*max(G), B<=0.5*max(B); or in HSV format image data, the channel values ​​of the preset characteristic color filtering range for green can be set to 30<=H<=90, 50<=S<=255, 50<=V<=255.

[0067] S240. When the values ​​of each channel belong to the preset feature color filtering range, the pixel coordinates corresponding to the digital array are used as candidate points for word prompting.

[0068] Specifically, when the values ​​of each channel all fall within the preset feature color filtering range, the processor can use the pixel coordinates corresponding to the digital array as candidate points for word selection.

[0069] S250. Based on each candidate point, determine the set of candidate points in the image to be segmented that meet the color selection criteria.

[0070] Specifically, the processor can treat all the word-prompting candidate points in the image to be segmented as a set, forming a set of word-prompting candidate points that meet the color filtering criteria.

[0071] S260. Determine the target segmentation mask set of the image to be segmented based on the candidate point set and the preset segmentation model.

[0072] In this embodiment, the target segmentation mask set can be understood as a collection of multiple target segmentation masks. A target segmentation mask can be understood as a mask used to represent all points belonging to an object.

[0073] Specifically, the processor can input any number of prompting candidate points from the prompting candidate point set into a preset prompting segmentation model to determine the target segmentation mask of the complete object to which the prompting candidate point belongs. The selected prompting candidate points are then deleted from the prompting candidate point set. Then, from the remaining prompting candidate points, prompting candidate points are randomly selected to determine the target segmentation mask, until all prompting candidate points are selected or the set minimum threshold is reached. The obtained target segmentation mask is then used as the target segmentation mask set of the image to be segmented.

[0074] Furthermore, based on the above embodiments, the step of determining the target segmentation mask set of the image to be segmented according to the candidate point set and the preset segmentation model can be optimized as follows:

[0075] a1. Randomly select a set number of prompting candidate points from the prompting candidate point set as target prompting candidate points.

[0076] In this embodiment, the target word prompting candidate point can be understood as the word prompting candidate point used as input into the model.

[0077] Specifically, the processor can randomly select a set number of prompting candidate points from the prompting candidate point set as target prompting candidate points.

[0078] b1. Based on each target word-prompting candidate point and the preset word-prompting segmentation model, determine the target segmentation mask corresponding to the intermediate target in the image to be segmented.

[0079] In this embodiment, the intermediate target can be understood as the object segmented by the preset word segmentation model.

[0080] Specifically, the processor can perform word segmentation using a preset word segmentation model for each target word segmentation candidate point, such as using the open-source word segmentation model: Segment Anything Model (SAM) to perform word segmentation and obtain the corresponding target segmentation mask.

[0081] For example, a yellow crane is placed in front of a green curtain. Since the structure of the yellow crane has gaps, the green curtain is in the gaps. Therefore, the prompting candidate point A includes the green curtain in the gap of the yellow crane. After the prompting candidate point A is segmented by the preset prompting segmentation model, the intermediate target to which the prompting candidate point A belongs is the yellow crane.

[0082] The step of determining the target segmentation mask corresponding to the intermediate target in the image to be segmented based on each target word-prompting candidate point and the preset word-prompting segmentation model can be optimized as follows:

[0083] b11. Perform position encoding on the pixel coordinates corresponding to each target word suggestion candidate point to obtain position encoding information.

[0084] In this embodiment, the location encoding information can be understood as encoding information used to convert the location into a location-reflecting information.

[0085] Specifically, the processor can use a preset word segmentation model to perform position encoding on the pixel coordinates corresponding to each target word candidate point. The word (token) type input to the model is pixel coordinates. The N coordinate points (w,h) are position encoded into an N×256 array (prompt toekn) to obtain the position encoding information.

[0086] b12. Input the image to be segmented into the preset visual model to obtain image encoding information.

[0087] In this embodiment, the preset visual model can be understood as a model used to obtain image encoding information, such as the ViT model. Image encoding information can be understood as encoding information used to represent an image with a small number of bits.

[0088] Specifically, the processor can input the image to be segmented into a preset visual model to obtain image encoding information (256×64×64).

[0089] b13. Determine the classification coding information corresponding to each pixel coordinate point.

[0090] In this embodiment, the classification coding information can be understood as the coding information that classifies the coordinates of each pixel into foreground or background.

[0091] Specifically, the processor can learn from the preset word segmentation model to determine whether the coordinates of each pixel belong to the foreground or the background (output token, an N×256 array), thus obtaining classification encoding information.

[0092] For example, taking the aforementioned yellow crane as an example, the yellow crane body is the foreground, while the green curtain between the cranes is the background.

[0093] b14. Input the location encoding information, image encoding information and classification encoding information into the preset word segmentation model to determine the target segmentation mask corresponding to the intermediate target of the image to be segmented.

[0094] Specifically, the processor can simultaneously send the classification information, the aforementioned location encoding information, and the image encoding information into the model decoding module (which includes a multi-layer self-attention module, a cross-attention module, and a fully connected network module) to obtain the target segmentation mask with the highest output score from the model.

[0095] c1. Filter out the candidate points to be filtered out corresponding to the target segmentation mask in the candidate point set to obtain the updated candidate point set.

[0096] The target segmentation mask includes the pixel coordinates of the foreground target belonging to the foreground target and the pixel coordinates of the background target belonging to the background target.

[0097] In this embodiment, the foreground target can be understood as the target that serves as the main body, and the pixel coordinates of the foreground target can be understood as all pixel coordinates belonging to the foreground target in the target segmentation mask. The background target can be understood as the target that serves as the background in the target segmentation mask, and the pixel coordinates of the background target can be understood as all pixel coordinates belonging to the background target.

[0098] In this embodiment, the candidate points to be filtered out can be understood as the candidate points corresponding to the foreground target in the target segmentation mask.

[0099] Specifically, the processor can filter the candidate point set. Based on the target segmentation mask obtained in the above steps, the candidate points (pixel coordinates) corresponding to the foreground target pixel coordinates in the target segmentation mask (set of pixel coordinates) are removed from the candidate point set, and only the candidate point set other than the foreground target mask is retained, thus obtaining the updated candidate point set.

[0100] d1. Based on the updated prompting candidate point set, return to the execution of the target segmentation mask determination step until the number of prompting candidate points in the updated prompting candidate point set is less than the set end threshold, and then determine the target segmentation mask set according to each target segmentation mask.

[0101] In this embodiment, setting the end threshold can be understood as a threshold used to end the target segmentation mask determination step.

[0102] Specifically, the processor can return to the determination steps of the target segmentation mask based on the updated prompting candidate point set, that is, steps a1 to c1 above, until the number of prompting candidate points in the updated prompting candidate point set is less than the set end threshold, and then each target segmentation mask is used as the target segmentation mask set of the image to be segmented.

[0103] S270. Based on the target segmentation mask set and the word prompting candidate point set, determine the proportion of feature color pixels corresponding to each intermediate target.

[0104] In this embodiment, the proportion of feature color pixels can be understood as the proportion of the feature color in the middle target out of all the colors in the middle target.

[0105] Furthermore, based on the above-mentioned strengths, the step of determining the proportion of feature color pixels corresponding to each intermediate target according to the target segmentation mask set and the word suggestion candidate point set can be further optimized as follows:

[0106] a2. Determine the number of foreground coordinate points in the target segmentation mask.

[0107] In this embodiment, the number of foreground coordinate points can be understood as the number of all coordinate points belonging to the foreground target.

[0108] Specifically, the processor can count the number of all foreground target pixel coordinates belonging to the foreground target in the target segmentation mask to obtain the number of foreground coordinates.

[0109] b2. Determine the set of color coordinate points belonging to the foreground target from the candidate point set, and determine the number of color coordinate points in the set.

[0110] It is important to know that the prompting candidate point set only includes prompting candidate points that meet the filtering color conditions, while the target segmentation mask includes all points belonging to the intermediate target, which may include some points that do not meet the filtering color conditions.

[0111] In this embodiment, the number of color coordinate points can be understood as the total number of pixel coordinate points that meet the color selection criteria.

[0112] Specifically, the processor can search for the set of color coordinate points belonging to the foreground target in the candidate point set and count the number of color coordinate points.

[0113] c2. Based on the number of color coordinate points and the number of foreground coordinate points, determine the proportion of feature color pixels of the intermediate target corresponding to the target segmentation mask.

[0114] Specifically, the processor can divide the number of color coordinate points by the number of foreground coordinate points to obtain the proportion of feature color pixels of the middle target corresponding to the target segmentation mask.

[0115] S280. When the proportion of feature color pixels is greater than the preset pixel proportion threshold, the corresponding intermediate target is taken as the target object that meets the color filtering conditions.

[0116] Specifically, the processor can compare the proportion of feature color pixels with a preset pixel proportion threshold. When the proportion of feature color pixels is greater than the preset pixel proportion threshold, the corresponding intermediate target is selected as the target object that meets the color filtering conditions. Intermediate targets that are less than or equal to the threshold are discarded.

[0117] For example, taking the yellow crane in the above embodiment as an example, since the main body of the yellow crane is yellow, that is, the foreground target in the target segmentation mask is the yellow crane and the background target is the green curtain, then all the word prompting candidate points corresponding to the green curtain are counted as the number of color coordinate points, and the number of foreground coordinate points corresponding to the yellow crane is the number of foreground coordinate points. Since the main body is the yellow crane, the proportion of the number of color coordinate points to the number of foreground coordinate points is very low, which is less than the preset pixel proportion threshold, so the yellow crane is discarded.

[0118] S290. Segment the image to be segmented according to the target object.

[0119] The technical solution of this invention filters the image to be segmented by a preset feature color filtering range to obtain candidate points for word extraction. Then, a preset word extraction segmentation model is used to extract words from each candidate point to obtain a high-quality target segmentation mask. The candidate point set within the target segmentation mask range is then filtered out. This iterative process of word extraction segmentation and candidate point filtering avoids repeated determination based on a single candidate point, improving word extraction speed and thus image segmentation speed. The process continues until the number of remaining candidate points is less than a set termination threshold, at which point a target segmentation mask set is obtained. This mask is then filtered by a preset pixel percentage threshold to obtain the target object in the image to be segmented. This achieves automatic determination of the target object to be segmented, ensuring the accuracy of target object determination, improving image segmentation efficiency, and avoiding the low efficiency and labeling errors of manual annotation.

[0120] To facilitate understanding of this solution, an example will be used for description. Figure 3 This is an example flowchart of an image segmentation method provided in Embodiment 2 of the present invention. Figure 3 As shown, the steps of this method are as follows:

[0121] S310. Obtain the image to be segmented and the color filtering conditions, and determine the digital array information corresponding to each pixel coordinate point in the image to be segmented;

[0122] S320. The digital array information is filtered by a preset feature color filtering range to obtain a set of candidate points that meet the filtering color conditions.

[0123] S330. Select one word-prompting candidate point from the word-prompting candidate point set and input it into the preset word-prompting segmentation model to obtain the target segmentation mask;

[0124] S340. Filter the candidate point set for word prompting according to the target segmentation mask to obtain the filtered candidate point set for word prompting.

[0125] S350. Is the number of candidate points in the filtered candidate point set less than the set end threshold? If yes, proceed to step S360; if no, proceed to step S330.

[0126] S360. Obtain the target segmentation mask set corresponding to the image to be segmented;

[0127] S370. Traverse the target segmentation mask in the target segmentation mask set, determine the proportion of the number of color coordinate points of the target segmentation mask in the word prompting candidate point set to the proportion of the number of foreground coordinate points of the foreground target pixel coordinate points in the target segmentation mask, and obtain the proportion of feature color pixels corresponding to the target segmentation mask.

[0128] S380. Take the target segmentation mask whose feature color pixel ratio is greater than the preset pixel ratio threshold as the target object, and discard the target segmentation mask whose feature color pixel ratio is less than or equal to the preset pixel ratio threshold.

[0129] S390. Obtain all target objects in the image to be segmented that meet the color filtering conditions, and segment the image to be segmented according to the target objects.

[0130] For example, to facilitate understanding of this solution, a specific example is used: to extract the green scaffolding curtain from the image to be segmented, the filter color condition is green. Figure 4 This is a schematic diagram of a candidate point set for word prompting in an image segmentation method according to Embodiment 2 of the present invention. The preset feature color filtering interval is green. The digital array information is filtered through this green filtering interval to obtain a candidate point set that meets the filtering color conditions, as shown below. Figure 4 As shown in the white area, due to changes in lighting conditions and the fact that some areas of the screen layout are not green, the white areas identified in the image will appear mottled. Figure 5 This is an example diagram of a target segmentation mask in an image segmentation method provided in Embodiment 2 of the present invention, as shown below. Figure 5 As shown, the star-shaped markers represent the candidate points for the input to the preset word segmentation model, and the white part is the target segmentation mask generated by the preset word segmentation model based on the candidate points. It can be seen that the target segmentation mask after processing by the preset word segmentation model includes the overall intermediate object to which the green curtain belongs. Figure 6 This is an example diagram of the target segmentation mask set in an image segmentation method provided in Embodiment 2 of the present invention, as shown below. Figure 6 As shown, it can be seen that all intermediate objects include the green curtain, including non-green cranes. Since the gaps between the non-green cranes are covered by the green curtain, the non-green cranes are also identified as intermediate objects when determining the target segmentation mask. Figure 7 This is an example diagram of the target object in an image segmentation method provided in Embodiment 2 of the present invention, such as... Figure 7 As shown, since the non-green crane is entirely non-green, with only some areas having a green background, the proportion of characteristic color pixels of the identified non-green crane is less than a preset pixel proportion threshold. Therefore, intermediate objects not belonging to the green background are filtered out using the characteristic color pixel proportion and the preset pixel proportion threshold, resulting in the following... Figure 7 The final target object shown.

[0131] Example 3

[0132] Figure 8 This is a schematic diagram of the structure of an image segmentation device provided in Embodiment 3 of the present invention. Figure 8As shown, the device includes: an information acquisition module 41, a first determination module 42, a second determination module 43, and an image segmentation module 44. Among them,

[0133] Information acquisition module 41 is used to acquire the image to be segmented and the color filtering conditions, and to determine the digital array information corresponding to each pixel coordinate point in the image to be segmented;

[0134] The first determining module 42 is used to filter the digital array information according to the preset feature color filtering range, and determine the set of word selection candidate points in the image to be segmented that meet the color filtering conditions.

[0135] The second determining module 43 is used to determine the target object that meets the filtering color conditions based on the word prompting candidate point set, the preset word prompting segmentation model and the preset pixel ratio threshold.

[0136] Image segmentation module 44 is used to segment the image to be segmented according to the target object.

[0137] The technical solution of this invention involves acquiring the image to be segmented and setting color selection criteria, determining the digital array information corresponding to each pixel coordinate point in the image to be segmented, filtering the digital array information according to a preset feature color selection interval, determining a set of candidate points in the image to be segmented that meet the color selection criteria, determining the target object that meets the color selection criteria based on the candidate point set, a preset word-selection segmentation model, and a preset pixel proportion threshold, and then segmenting the image to be segmented based on the target object. By filtering the image to be segmented through a preset feature color selection interval to obtain candidate points, obtaining the target segmentation mask through a preset word-selection segmentation model, and then filtering it through a preset pixel proportion threshold to obtain the target object, the invention achieves automatic determination of the target object to be segmented, ensuring the accuracy of target object determination and improving the efficiency of image segmentation.

[0138] Furthermore, the first determining module 42 is specifically used for:

[0139] For each pixel coordinate point in the image to be segmented, the corresponding digital array is determined in the digital array information;

[0140] The channel values ​​of each channel in the digital array are compared with the preset feature color filtering range;

[0141] When the values ​​of each channel belong to the preset feature color filtering range, the pixel coordinates corresponding to the digital array are used as candidate points for word prompting.

[0142] Based on each of the aforementioned candidate points, a set of candidate points for prompting words that meet the color selection criteria is determined in the image to be segmented.

[0143] Furthermore, the second determining module 43 includes:

[0144] The first determining unit is used to determine the target segmentation mask set of the image to be segmented based on the candidate point set and the preset segmentation model.

[0145] The second determining unit is used to determine the proportion of feature color pixels corresponding to each intermediate target based on the target segmentation mask set and the word prompting candidate point set.

[0146] The third determining unit is used to determine the corresponding intermediate target as the target object that satisfies the filtering color condition when the proportion of the feature color pixels is greater than the preset pixel proportion threshold.

[0147] Furthermore, the first determining unit includes:

[0148] The first determining subunit is used to randomly select a set number of prompting candidate points from the prompting candidate point set as target prompting candidate points;

[0149] The second determining subunit is used to determine the target segmentation mask corresponding to the intermediate target of the image to be segmented based on each of the target word-prompting candidate points and the preset word-prompting segmentation model.

[0150] The third determining subunit is used to filter out the candidate points to be filtered out corresponding to the foreground target pixel coordinates in the candidate point set, so as to obtain the updated candidate point set.

[0151] The fourth determining subunit is used to return the determination step of the target segmentation mask based on the updated word prompting candidate point set, until the number of word prompting candidate points in the updated word prompting candidate point set is less than a set end threshold, and then determine the target segmentation mask set according to each target segmentation mask.

[0152] Specifically, the second determining subunit is used for:

[0153] Position encoding is performed on the pixel coordinates corresponding to each of the target word suggestion candidates to obtain position encoding information;

[0154] The image to be segmented is input into the preset visual model to obtain image encoding information;

[0155] Determine the classification coding information corresponding to each pixel coordinate point;

[0156] The location encoding information, the image encoding information, and the classification encoding information are input into the preset word segmentation model to determine the target segmentation mask corresponding to the intermediate target of the image to be segmented.

[0157] The target segmentation mask includes the pixel coordinates of the foreground target belonging to the foreground target and the pixel coordinates of the background target belonging to the background target.

[0158] Furthermore, the second unit is specifically used for:

[0159] Determine the number of foreground coordinate points in the target segmentation mask;

[0160] In the candidate prompting point set, determine the set of color coordinate points belonging to the foreground target, and determine the number of color coordinate points in the set of color coordinate points;

[0161] Based on the number of color coordinate points and the number of foreground coordinate points, the proportion of feature color pixels of the intermediate target corresponding to the target segmentation mask is determined.

[0162] The image segmentation apparatus provided in the embodiments of the present invention can execute the image segmentation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0163] Example 4

[0164] Figure 9 A schematic diagram of an electronic device 50 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0165] like Figure 9 As shown, the electronic device 50 includes at least one processor 51 and a memory, such as a read-only memory (ROM) 52 and a random access memory (RAM) 53, communicatively connected to the at least one processor 51. The memory stores computer programs executable by the at least one processor. The processor 51 can perform various appropriate actions and processes based on the computer program stored in the ROM 52 or loaded into the RAM 53 from storage unit 58. The RAM 53 can also store various programs and data required for the operation of the electronic device 50. The processor 51, ROM 52, and RAM 53 are interconnected via a bus 54. An input / output (I / O) interface 55 is also connected to the bus 54.

[0166] Multiple components in electronic device 50 are connected to I / O interface 55, including: input unit 56, such as keyboard, mouse, etc.; output unit 57, such as various types of monitors, speakers, etc.; storage unit 58, such as disk, optical disk, etc.; and communication unit 59, such as network card, modem, wireless transceiver, etc. Communication unit 59 allows electronic device 50 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0167] Processor 51 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 51 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 51 performs the various methods and processes described above, such as image segmentation methods.

[0168] In some embodiments, the image segmentation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 58. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 50 via ROM 52 and / or communication unit 59. When the computer program is loaded into RAM 53 and executed by processor 51, one or more steps of the image segmentation method described above may be performed. Alternatively, in other embodiments, processor 51 may be configured to perform the image segmentation method by any other suitable means (e.g., by means of firmware).

[0169] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0170] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0171] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0172] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0173] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0174] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0175] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0176] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. An image segmentation method, characterized in that, include: Obtain the image to be segmented and the color filtering conditions, and determine the digital array information corresponding to each pixel coordinate point in the image to be segmented; The digital array information is filtered according to a preset feature color filtering range to determine the set of candidate points for word retrieval in the image to be segmented that meet the color filtering conditions. Based on the candidate point set, the preset word prompting segmentation model, and the preset pixel ratio threshold, the target object that meets the filtering color condition is determined; The image to be segmented is segmented according to the target object; The step of determining the target object in the image to be segmented that meets the filtering color conditions based on the candidate point set, the preset word-prompting segmentation model, and the preset pixel proportion threshold includes: Based on the candidate point set and the preset segmentation model, the target segmentation mask set of the image to be segmented is determined; Based on the target segmentation mask set and the word prompting candidate point set, determine the proportion of feature color pixels corresponding to each intermediate target; When the proportion of the feature color pixels is greater than the preset pixel proportion threshold, the corresponding intermediate target is used as the target object that meets the filtering color conditions.

2. The method according to claim 1, characterized in that, The step of filtering the digital array information according to a preset feature color filtering range to determine the set of candidate points for word retrieval in the image to be segmented that meet the color filtering conditions includes: For each pixel coordinate point in the image to be segmented, the corresponding digital array is determined in the digital array information; The channel values ​​of each channel in the digital array are compared with the preset feature color filtering range; When the values ​​of each channel belong to the preset feature color filtering range, the pixel coordinates corresponding to the digital array are used as candidate points for word prompting. Based on each of the aforementioned candidate points, a set of candidate points for prompting words that meet the color selection criteria is determined in the image to be segmented.

3. The method according to claim 1, characterized in that, The target segmentation mask includes foreground target pixel coordinates belonging to foreground targets and background target pixel coordinates belonging to background targets. Correspondingly, determining the target segmentation mask set of the image to be segmented based on the candidate word-prompting point set and the preset word-prompting segmentation model includes: A set number of prompting candidate points are randomly selected from the prompting candidate point set as target prompting candidate points; Based on each of the target word-prompting candidate points and the preset word-prompting segmentation model, determine the target segmentation mask corresponding to the intermediate target of the image to be segmented; The candidate points to be filtered out corresponding to the foreground target pixel coordinates are filtered out from the candidate point set to obtain the updated candidate point set. The step of determining the target segmentation mask is returned based on the updated prompting candidate point set, until the number of prompting candidate points in the updated prompting candidate point set is less than a set end threshold, and the target segmentation mask set is determined according to each target segmentation mask.

4. The method according to claim 3, characterized in that, The step of determining the target segmentation mask corresponding to the intermediate target in the image to be segmented based on each of the target word-prompting candidate points and the preset word-prompting segmentation model includes: Position encoding is performed on the pixel coordinates corresponding to each of the target word suggestion candidates to obtain position encoding information; The image to be segmented is input into a preset visual model to obtain image encoding information; Determine the classification coding information corresponding to each pixel coordinate point; The location encoding information, the image encoding information, and the classification encoding information are input into the preset word segmentation model to determine the target segmentation mask corresponding to the intermediate target of the image to be segmented.

5. The method according to claim 3, characterized in that, The step of determining the proportion of feature color pixels corresponding to each intermediate target based on the target segmentation mask set and the word prompting candidate point set includes: Determine the number of foreground coordinate points in the target segmentation mask; In the candidate prompting point set, determine the set of color coordinate points belonging to the foreground target, and determine the number of color coordinate points in the set of color coordinate points; Based on the number of color coordinate points and the number of foreground coordinate points, the proportion of feature color pixels of the intermediate target corresponding to the target segmentation mask is determined.

6. An image segmentation apparatus, characterized in that, include: The information acquisition module is used to acquire the image to be segmented and the color filtering conditions, and to determine the digital array information corresponding to each pixel coordinate point in the image to be segmented; The first determining module is used to filter the digital array information according to a preset feature color filtering range to determine the set of candidate points for word selection in the image to be segmented that meet the color filtering conditions. The second determining module is used to determine the target object that meets the filtering color conditions based on the word prompting candidate point set, the preset word prompting segmentation model and the preset pixel ratio threshold. An image segmentation module is used to segment the image to be segmented based on the target object; The second determining module includes: The first determining unit is used to determine the target segmentation mask set of the image to be segmented based on the candidate point set and the preset segmentation model. The second determining unit is used to determine the proportion of feature color pixels corresponding to each intermediate target based on the target segmentation mask set and the word prompting candidate point set. The third determining unit is used to determine the corresponding intermediate target as the target object that satisfies the filtering color condition when the proportion of the feature color pixels is greater than the preset pixel proportion threshold.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the image segmentation method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the image segmentation method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Image segmentation method, apparatus and device, and computer readable storage medium

    CN114972367A

  • Image target object auxiliary labeling method based on point labeling

    CN116630977A