Image segmentation method, device, equipment and storage medium

Through vector field prediction model and aggregation technology, an image segmentation method is realized. The generated superpixel blocks carry semantic information, improve the segmentation speed, and solve the problem of superpixels without semantic information and inefficiency in the prior art.

CN114037716BActive Publication Date: 2025-05-16DOUYIN VISION CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111322032.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-09
Publication Date
2025-05-16
Estimated Expiration
2041-11-09

AI Technical Summary

Technical Problem

The superpixels generated by the existing image segmentation technology do not contain semantic information and are inefficient, so the deep learning method is time-consuming.

Method used

By inputting the image to be segmented to set the vector field prediction model, the vector field of each pixel point is obtained, the pixel point chain and root node are determined, the root node is aggregated using a preset size box, and the sub-regions are aggregated according to the segmentation boundary to obtain multiple second sub-regions.

Benefits of technology

The generated superpixel block carries semantic information, which improves image segmentation speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114037716B_ABST
    Figure CN114037716B_ABST
Patent Text Reader

Abstract

The disclosed embodiments disclose an image segmentation method, apparatus, device and storage medium. The image to be segmented is input into a set vector field prediction model to obtain the vector field of each pixel in the image to be segmented; the pixel point chain where each pixel point is located and the root node of the pixel point chain are determined according to the vector field; the root node is aggregated using a frame of a preset size to obtain an aggregated pixel point chain; wherein the pixels in a pixel point chain form a first sub-region; the first sub-region is aggregated according to the segmentation boundary of the first sub-region to obtain multiple second sub-regions, and a segmented image is obtained. The image segmentation method provided by the disclosed embodiments realizes image segmentation based on the vector field of pixel points, so that the generated super-pixel block carries semantic information and can improve the image segmentation speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of image processing technology, and in particular to an image segmentation method, apparatus, device and storage medium. Background Art

[0002] Image segmentation technology is to divide an image into several image blocks. It can be applied to image editing to automatically segment images uploaded by users to facilitate subsequent editing.

[0003] Existing image segmentation technologies, such as the simple linear iterative clustering (SLIC) algorithm, use the underlying color features of the image to segment the image. The generated superpixels do not carry semantic information and require parameter adjustment, which is inefficient. Other image segmentation technologies that use deep learning are time-consuming. Summary of the invention

[0004] The embodiments of the present disclosure provide an image segmentation method, apparatus, device and storage medium to achieve image segmentation, so that the generated superpixel blocks carry semantic information and the image segmentation speed can be improved.

[0005] In a first aspect, an embodiment of the present disclosure provides an image segmentation method, comprising:

[0006] Inputting the image to be segmented into a set vector field prediction model to obtain the vector field of each pixel point in the image to be segmented;

[0007] Determine the pixel point chain where each pixel point is located and the root node of the pixel point chain according to the vector field; wherein the pixel point chain is composed of a plurality of pixel points that meet a preset condition, and the root node is a pixel point in the pixel point chain;

[0008] Aggregating the root nodes using a frame of a preset size to obtain an aggregated pixel point chain; wherein the pixels in one pixel point chain constitute a first sub-region;

[0009] The first sub-regions are aggregated according to the segmentation boundaries of the first sub-regions to obtain a plurality of second sub-regions, thereby obtaining a segmented image.

[0010] In a second aspect, the present disclosure also provides an image segmentation device, including:

[0011] A vector field acquisition module is used to input the image to be segmented into a set vector field prediction model to obtain the vector field of each pixel point in the image to be segmented;

[0012] A root node determination module, used to determine the pixel point chain where each pixel point is located and the root node of the pixel point chain according to the vector field; wherein the pixel point chain is composed of a plurality of pixel points that meet preset conditions, and the root node is a pixel point in the pixel point chain;

[0013] A root node aggregation module, used to aggregate the root node using a frame of a preset size to obtain an aggregated pixel point chain; wherein the pixels in one pixel point chain constitute a first sub-region;

[0014] The segmented image acquisition module is used to aggregate the first sub-regions according to the segmentation boundaries of the first sub-regions to obtain multiple second sub-regions and obtain a segmented image.

[0015] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:

[0016] one or more processing devices;

[0017] A storage device for storing one or more programs;

[0018] When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement the image segmentation method as described in the embodiment of the present disclosure.

[0019] In a fourth aspect, the embodiments of the present disclosure further provide a computer-readable medium on which a computer program is stored, and when the program is executed by a processing device, the image segmentation method as described in the embodiments of the present disclosure is implemented.

[0020] The disclosed embodiments disclose an image segmentation method, apparatus, device and storage medium. The image to be segmented is input into a set vector field prediction model to obtain the vector field of each pixel in the image to be segmented; the pixel point chain where each pixel point is located and the root node of the pixel point chain are determined according to the vector field; the root node is aggregated using a frame of a preset size to obtain an aggregated pixel point chain; wherein the pixels in a pixel point chain form a first sub-region; the first sub-region is aggregated according to the segmentation boundary of the first sub-region to obtain multiple second sub-regions, and a segmented image is obtained. The image segmentation method provided by the disclosed embodiments realizes image segmentation based on the vector field of pixel points, so that the generated super-pixel block carries semantic information and can improve the image segmentation speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a flow chart of an image segmentation method in an embodiment of the present disclosure;

[0022] Figure 2 is a structural schematic diagram of a vector field prediction model in an embodiment of the present disclosure;

[0023] Figure 3 is an example diagram of a pixel point vector field in an embodiment of the present disclosure;

[0024] Figure 4 is an example diagram for visualizing a vector field in an embodiment of the present disclosure;

[0025] Figure 5 is an example diagram of a vector field in an embodiment of the present disclosure;

[0026] Figure 6 is an example diagram of image segmentation in an embodiment of the present disclosure;

[0027] Figure 7 is an example diagram of aggregating sub-regions based on segmentation boundaries in an embodiment of the present disclosure;

[0028] Figure 8 is an example diagram of merging sub-regions in an embodiment of the present disclosure;

[0029] Fig. 9 is a structural schematic diagram of an image segmentation device in an embodiment of the present disclosure;

[0030] Fig.10 It is a structural diagram of an electronic device in an embodiment of the present disclosure;. DETAILED DESCRIPTION

[0031] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0032] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0033] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0034] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0035] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0036] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0037] Figure 1 This is a flow chart of an image segmentation method provided by an embodiment of the present disclosure. This embodiment is applicable to the case of segmenting an image. The method can be executed by an image segmentation device, which can be composed of hardware and / or software and can generally be integrated in a device with image segmentation function, which can be an electronic device such as a server, a mobile terminal or a server cluster. Figure 1 As shown, the method specifically comprises the following steps:

[0038] Step 110: input the image to be segmented into a set vector field prediction model to obtain the vector field of each pixel in the image to be segmented.

[0039] The image to be segmented can be any color or grayscale image. The vector field prediction model can be a deep learning model. In order to deploy the model on a mobile terminal, the model needs to have a small amount of calculation, high computation efficiency and simplicity. In the embodiment of the present disclosure, the traditional convolutional network is replaced by a deep separable convolutional network.

[0040] Optional, Figure 2 Schematic diagram of a vector field prediction model in an embodiment of the present disclosure. Figure 2 As shown, the vector field prediction model is set to include: a channel exchange network, a channel segmentation network and a depth-separable convolutional network.

[0041] Among them, the deep separable convolutional network can be easily deployed on the mobile terminal due to its advantages of small structure and small amount of calculation. In this embodiment, since the deep separable convolutional network is included in the set vector field prediction model, the image to be segmented is input into the set vector field prediction model, and the vector field of each pixel point can be determined with a small amount of calculation, thereby improving the efficiency of image segmentation according to the vector field.

[0042] Among them, the deep separable convolutional network includes a first channel convolutional subnetwork, a deep convolutional subnetwork, a second channel convolutional subnetwork and a channel merging layer; the channel exchange network, the channel splitting network, the first channel convolutional subnetwork, the deep convolutional subnetwork, the second channel convolutional subnetwork and the channel merging layer are connected in sequence; and the output of the channel splitting network is jump-connected to the input of the channel merging layer. Among them, the deep convolutional subnetwork can improve the feature extraction capability of the set vector field prediction model, and the channel convolutional subnetwork and the deep convolutional subnetwork have the advantages of small structure and small computational complexity. In this embodiment, since the deep convolutional subnetwork has a high feature extraction capability, the image to be segmented is input into the set vector field prediction model, and the features related to each pixel point and the vector field can be accurately extracted and processed, thereby improving the accuracy of the determined vector field.

[0043] like Figure 2 As shown, the first channel convolution subnetwork includes the first channel convolution layer, the nonlinear activation layer and the linear transformation layer; the depth convolution subnetwork includes the depth convolution layer (Depthwise Convolution), the nonlinear activation layer and the linear transformation layer; the second channel convolution subnetwork includes the second channel convolution layer (Pointwise Convolution), the nonlinear activation layer and the linear transformation layer; the depth convolution layer consists of multiple parallel convolution kernels.

[0044] Among them, the first channel convolution layer and the second channel convolution layer can both be composed of 1×1 convolution kernels. The deep convolution layer can be composed of 3×3 convolution kernels, and the 3×3 convolution kernel is composed of three parallel convolution kernels. The sizes of the three parallel convolution kernels are divided into 3×3, 3×1 and 1×3. The 3×3 convolution kernel is implemented by three parallel convolution kernels, which can improve the calculation speed of the model. The channel exchange network can be implemented by channel shuffle, the nonlinear activation layer can be implemented by the linear rectification function (Rectified Linear Unit, ReLU), and the linear transformation layer can be implemented by the batch normalization (BatchNormalization, BN) algorithm. The vector field prediction model provided in this embodiment has low working time consumption and can be applied to mobile terminals with high time consumption requirements.

[0045] In the disclosed embodiment, the training process of setting the vector field prediction model may be: collecting a large amount of image data, firstly performing panoramic segmentation annotation or image edge annotation on the image data to obtain the segmented image, and then using the distance transformation function (distanceTransform) to calculate the direction vector of each pixel point pointing to the nearest segmentation edge, that is, the vector field of each pixel point, to obtain training samples, and based on the training samples, the set vector field prediction model is trained. The loss function that can be used during training may be a mean square error (MSE) function or a mean absolute error (MAE) function. This embodiment does not limit the choice of the loss function.

[0046] The vector field can represent the direction vector of the pixel pointing to the nearest segmentation edge. Assume that the length of the image is H and the width is W, and each pixel has a value representing the direction of the current pixel (range 0-360 degrees), that is, the vector field. For example, Figure 3 is an example diagram of a pixel vector field in an embodiment of the present disclosure. Figure 3 As shown, it is a 3*3 pixel block, and each pixel has a direction vector representing the direction of the current pixel. The set vector field prediction model provided in this embodiment can convert the RGB image into vector field information. The model will output 2 channels, one channel represents the horizontal direction cos(θ), and the other channel represents the vertical direction sin(θ). θ can be obtained according to the cos(θ) and sin(θ) of each pixel. Figure 4 is an example diagram for visualizing a vector field. Figure 4 As shown, different θ values ​​correspond to different grayscale values ​​(if the image is in color, different θ values ​​correspond to different color values).

[0047] Step 120, determining the pixel point chain where each pixel point is located and the root node of the pixel point chain according to the vector field.

[0048] The pixel point chain is composed of a plurality of pixel points that meet the preset conditions, and the root node is a pixel point in the pixel point chain. In this embodiment, the root node is the pixel point with the highest level in the pixel point chain. The preset condition may be: the angle between the vector fields of two adjacent pixel points in the pixel point chain is less than the first set threshold. Specifically, after determining the vector field of each pixel point in the image to be segmented, each pixel point is first traversed. For the traversed pixel point, the parent node is determined from the eight pixel points adjacent to the pixel point according to the vector field, and the adjacent pixel point with the smallest vector field angle is used as the parent node of the current pixel point. For the parent node, its parent node is calculated in the above manner until the angle between the vector field of the parent node of the Nth level and the parent node of the N+1th level is greater than the first set threshold, then the parent node of the Nth level is used as the root node, and the current pixel point, the parent nodes of each level and the root node constitute the pixel point chain.

[0049] Optionally, the method of determining the pixel point chain where each pixel point is located and the root node of the pixel point chain according to the vector field can be: traverse each pixel point in the image to be segmented, take the traversed pixel point as the current pixel point, obtain a set number of pixel points adjacent to the current pixel point and whose positions meet preset conditions as candidate pixel points; determine the angle between the vector field of the candidate pixel point and the vector field of the current pixel point; determine the candidate pixel point corresponding to the minimum angle as the parent node, and determine the parent node as the current pixel point; return to execute the operation of obtaining a set number of pixel points adjacent to the current pixel point and whose positions meet preset conditions as candidate pixel points, until the minimum angle is greater than the first set threshold, then determine the current node as the root node, and obtain the pixel point chain where the traversed pixel point is located; continue to traverse the next pixel point that has not formed a pixel point chain, until all pixels in the image to be segmented form a pixel point chain, and obtain multiple root nodes.

[0050] The preset condition is to be located in the starting direction of the vector field of the pixel point. The starting direction of the vector field can be understood as the direction from which the vector is emitted. Figure 5 is an example diagram of a vector field in this embodiment, such as Figure 5 As shown in the figure, the starting direction of the vector field of the pixel numbered "5" is the lower right corner, and the pixels in the lower right corner and adjacent to pixel 5 are "6, 8, 9", that is, for pixel 5, pixels 6, 8 and 9 are candidate pixels. Similarly, the starting direction of the vector field of the pixel numbered "4" is the upper right corner, and the pixels in the upper right corner and adjacent to pixel 4 are "1, 2, 5", that is, for pixel 4, pixels 1, 2 and 5 are candidate pixels.

[0051] The method for determining the vector field angle may be directly calculated according to the method for calculating the vector angle. In this embodiment, since the vector field is represented by an angle θ, the method for determining the vector field angle may be directly to make a difference of the angle corresponding to the vector field.

[0052] Specifically, after obtaining the parent node of the current node, if the angle between the current node and the parent node vector field is less than the first set threshold, continue to obtain the parent node of the parent node in the same manner until the root node is obtained. Continue to traverse the pixel points that have not formed a pixel point chain until all the pixel points in the image to be segmented form a pixel point chain, thereby obtaining multiple root nodes. The pixel point chain corresponding to each root node constitutes a segmentation area. From the above method, it can be seen that the physical meaning of the root node is that the vector fields of the pixel points in the pixel point chain all start from here, and the root node is the center point of the segmentation area. In this embodiment, the pixel point chain is determined by the angle of the vector field of the pixel point, and the preliminary clustering of the pixel points by the vector field is realized, which can improve the accuracy of pixel point clustering.

[0053] Step 130: Aggregate the root nodes using a frame of a preset size to obtain an aggregated pixel point chain.

[0054] The pixels in a pixel chain form a first sub-region. Figure 6 is an example diagram of image segmentation in this embodiment, such as Figure 6 As shown in , this figure is the result of segmenting the image according to the pixel chain, and different areas have different grayscales (if it is a color image, different areas have different colors). Figure 6 It can be seen that the granularity of each sub-region of the segmentation result is too small, which is not conducive to subsequent processing. Therefore, it is necessary to further aggregate the pixel point chain.

[0055] The frame of the preset size may be a rectangular frame or a circular frame, etc. Taking the rectangular frame as an example, the size of the frame may be 3*3 or 5*5. Specifically, the pixel point chains corresponding to at least two root nodes falling into the frame are aggregated into the same pixel point chain. In this example, aggregating the root nodes can be understood as aggregating the first sub-region. The advantage of doing so is to increase the granularity of the segmented sub-regions.

[0056] Optionally, the root nodes are aggregated using a frame of a preset size, and the aggregated pixel point chain is obtained by: translating the frame of the preset size in the image to be segmented according to a set step size; and aggregating the pixel point chains of at least two root nodes falling within the frame of the preset size into the same pixel point chain.

[0057] The step size includes a horizontal step size and a vertical step size. If the frame of the preset size is translated horizontally on the image to be segmented, it is translated according to the horizontal step size. If the frame of the preset size is translated vertically on the image to be segmented, it is translated according to the vertical step size. The horizontal step size is less than or equal to the width W of the frame, and the vertical step size is less than or equal to the length H of the frame. The technical solution of this embodiment clusters the root nodes of the pixel point chain to increase the granularity of the segmented area.

[0058] Step 140 , aggregating the first sub-regions according to the segmentation boundaries of the first sub-regions to obtain a plurality of second sub-regions, thereby obtaining a segmented image.

[0059] In this embodiment, after the root node is aggregated, there are still sub-regions with smaller granularity, so the first sub-region needs to be further aggregated. The vector field of each pixel point in the segmentation boundary is used to determine whether the segmentation boundary is a "strong" boundary or a "weak" boundary. If it is a "weak" boundary, the segmentation boundary can be removed, that is, the first sub-region corresponding to the "weak" boundary is aggregated into one region to obtain the second sub-region. The technical solution of this embodiment aggregates the sub-regions based on the strength of the segmentation boundary, which can further increase the granularity of the segmentation region.

[0060] Optionally, the first sub-regions are aggregated according to the segmentation boundaries of the first sub-regions to obtain multiple second sub-regions by: for each segmentation boundary, determining the boundary strength of the segmentation boundary; if the boundary strength is less than a second set threshold, aggregating the two first sub-regions corresponding to the segmentation boundary to obtain the second sub-region.

[0061] Among them, the method for determining the boundary strength of the segmentation boundary can be: for each pixel point of the segmentation boundary, calculate the angle between the vector field of the pixel point and the vector field of the Nth level parent node; where N≥1; calculate the average value of the angles corresponding to all pixel points included in the segmentation boundary; and determine the average value as the boundary strength.

[0062] Among them, the first-level parent node of the pixel point is the parent node of the pixel point, the second-level parent node is the parent node of the parent node of the pixel point, and so on, the N-th level parent node can be obtained. For example: N can be 3. Specifically, the process of determining the N-th level parent node of the pixel point can refer to the above embodiment, which will not be repeated here. The method for determining the vector field angle can be directly calculated according to the calculation method of the vector angle. In this embodiment, since the vector field is represented by an angle θ, the method for determining the vector field angle can directly make a difference in the angle corresponding to the vector field.

[0063] Specifically, for a certain segmentation boundary, the angle between each pixel point of the segmentation boundary and the vector field of its N-th level parent node is calculated, and then the average value of the corresponding angles of all pixel points of the segmentation boundary is calculated, and the average value is used as the boundary strength of the segmentation boundary. If the boundary strength is less than the second set threshold, it indicates that the segmentation boundary is a "weak" boundary, and the segmentation boundary is removed, that is, the two first sub-regions corresponding to the segmentation boundary are aggregated into one sub-region. Exemplary, Figure 7 is an example diagram of sub-region aggregation based on segmentation boundaries in this embodiment. Figure 7As shown, the left side is the segmentation map before aggregation, and the right side is the segmentation map after aggregation. It can be seen from the figure that the granularity of the sub-region in the right side is significantly larger than that of the sub-region in the left side. In this embodiment, the boundary strength of the segmentation boundary is determined according to the angle between the vector field of the pixel point on the segmentation boundary and the vector field of its Nth level parent node, which can improve the accuracy of determining the boundary strength.

[0064] In the disclosed embodiment, a minimum area parameter may be set, that is, the size of each area cannot be smaller than the minimum area parameter. Therefore, based on the minimum area parameter, further sub-area aggregation needs to be performed on the segmented image.

[0065] Optionally, after aggregating the first sub-region according to the segmentation boundary of the first sub-region to obtain multiple second sub-regions, the following steps are also included: extracting sub-regions in which the proportion of first pixels is less than a third set threshold among the multiple second sub-regions, and determining them as sub-regions to be merged; obtaining the common segmentation boundary between the sub-region to be merged and the adjacent second sub-region; determining the adjacent second sub-region corresponding to the segmentation boundary with the largest proportion of second pixels in the common segmentation boundary as the target sub-region; merging the region to be merged with the target region to obtain a new second sub-region.

[0066] The first pixel ratio is the ratio of the pixels included in the second sub-region to the total pixels included in the image to be segmented. The second pixel ratio is the ratio of the pixels included in the common segmentation boundary to the pixels included in all common segmentation boundaries.

[0067] Among them, there are one or more adjacent sub-regions of the sub-region to be merged, and there is a segmentation boundary between the sub-region to be merged and each adjacent sub-region, and the segmentation boundary is the common segmentation boundary of the two sub-regions. In this embodiment, the sub-region to be merged is merged into the region with the largest proportion of adjacent pixels to the region. Exemplarily, assuming that the region to be merged has three adjacent regions, namely sub-region a, sub-region b and sub-region c, and the corresponding common segmentation boundaries are common segmentation boundary A, common segmentation boundary B and common segmentation boundary C, wherein the second pixel point proportion of common segmentation boundary A is 50%, the second pixel point proportion of common segmentation boundary B is 20%, and the second pixel point proportion of common segmentation boundary C is 30%, then the region to be merged is merged into sub-region a. In this embodiment, it is necessary to detect all second sub-regions until the proportion of the first pixel points of the sub-regions contained in the segmented image is greater than the third set threshold. The advantage of this is that the segmented sub-regions can carry semantic information and the granularity of region segmentation can be further improved. Exemplarily, Figure 8 FIG. 2 is an example diagram of merging sub-regions in this embodiment. Figure 8 As shown, the left picture is the segmentation picture before merging, and the right picture is the segmentation picture after merging. Figure 4By comparing the left figure, we can see that Figure 8 The sub-regions in the right middle image are the segmented regions of the sky, water, and buildings on both sides.

[0068] The technical solution of the disclosed embodiment is to input the image to be segmented into a set vector field prediction model to obtain the vector field of each pixel in the image to be segmented; determine the pixel point chain where each pixel point is located and the root node of the pixel point chain according to the vector field; aggregate the root node using a frame of a preset size to obtain an aggregated pixel point chain; wherein the pixels in a pixel point chain form a first sub-region; aggregate the first sub-region according to the segmentation boundary of the first sub-region to obtain multiple second sub-regions, and obtain a segmented image. The image segmentation method provided by the disclosed embodiment realizes image segmentation based on the vector field of pixel points, so that the generated super pixel block carries semantic information and can improve the image segmentation speed.

[0069] Fig. 9 is a structural schematic diagram of an image segmentation device provided by an embodiment of the present disclosure, such as Fig. 9 As shown, the device comprises:

[0070] The vector field acquisition module 210 is used to input the image to be segmented into a set vector field prediction model to obtain the vector field of each pixel point in the image to be segmented;

[0071] A root node determination module 220 is used to determine the pixel point chain where each pixel point is located and the root node of the pixel point chain according to the vector field; wherein the pixel point chain is composed of a plurality of pixel points that meet a preset condition, and the root node is a pixel point in the pixel point chain;

[0072] A root node aggregation module 230 is used to aggregate the root nodes using a frame of a preset size to obtain an aggregated pixel point chain; wherein the pixels in a pixel point chain constitute a first sub-region;

[0073] The segmented image acquisition module 240 is used to aggregate the first sub-regions according to the segmentation boundaries of the first sub-regions to obtain a plurality of second sub-regions and obtain a segmented image.

[0074] Optionally, setting the vector field prediction model includes: a channel exchange network, a channel segmentation network and a depth separable convolutional network;

[0075] Among them, the depth separable convolutional network includes a first channel convolutional subnetwork, a deep convolutional subnetwork, a second channel convolutional subnetwork and a channel merging layer;

[0076] The channel exchange network, the channel splitting network, the first channel convolution subnetwork, the deep convolution subnetwork, the second channel convolution subnetwork and the channel merging layer are connected in sequence; and the output of the channel splitting network is jump-connected to the input of the channel merging layer;

[0077] The first channel convolution subnetwork includes the first channel convolution layer, the nonlinear activation layer and the linear transformation layer; the deep convolution subnetwork includes the deep convolution layer, the nonlinear activation layer and the linear transformation layer; the second channel convolution subnetwork includes the second channel convolution layer, the nonlinear activation layer and the linear transformation layer; the deep convolution layer consists of multiple parallel convolution kernels.

[0078] Optionally, the root node determination module 220 is further configured to:

[0079] Traverse each pixel point in the image to be segmented, take the traversed pixel point as the current pixel point, and obtain a set number of pixel points adjacent to the current pixel point and whose positions meet the preset conditions as candidate pixel points; wherein the preset condition is located in the starting direction of the pixel point vector field;

[0080] Determine the angle between the vector field of the candidate pixel and the vector field of the current pixel;

[0081] The candidate pixel point corresponding to the minimum angle is determined as the parent node, and the parent node is determined as the current pixel point;

[0082] Return to execute the operation of obtaining a set number of pixel points adjacent to the current pixel point and whose positions meet the preset conditions as candidate pixel points, until the minimum angle is greater than the first set threshold, then determine the current node as the root node, and obtain the pixel point chain where the traversed pixel point is located;

[0083] Continue to traverse the next pixel point that has not formed a pixel point chain until all the pixel points in the image to be segmented form a pixel point chain, and obtain multiple root nodes.

[0084] Optionally, the root node aggregation module 230 is further configured to:

[0085] Translate a frame of a preset size in the image to be segmented according to a set step size;

[0086] The pixel point chains where at least two root nodes fall within a frame of a preset size are aggregated into the same pixel point chain.

[0087] Optionally, the segmented image acquisition module 240 is further used for:

[0088] For each segmentation boundary, determining a boundary strength of the segmentation boundary;

[0089] If the boundary strength is less than the second set threshold, the two first sub-regions corresponding to the segmentation boundary are aggregated to obtain a second sub-region.

[0090] Optionally, the segmented image acquisition module 240 is further used for:

[0091] For each pixel point on the segmentation boundary, calculate the angle between the vector field of the pixel point and the vector field of the Nth level parent node; where N ≥ 1;

[0092] Calculate the average value of the angles corresponding to all pixels included in the segmentation boundary;

[0093] The mean value was determined as the boundary intensity.

[0094] Optionally, it further includes: a sub-region merging module, which is used to:

[0095] Extracting sub-regions in which the proportion of first pixels is less than a third set threshold value from among the multiple second sub-regions, and determining them as sub-regions to be merged; wherein the proportion of first pixels is the ratio of pixels included in the second sub-region to the total pixels included in the image to be segmented;

[0096] Obtaining a common segmentation boundary between the sub-region to be merged and the adjacent second sub-region;

[0097] Determine the adjacent second sub-region corresponding to the segmentation boundary with the largest second pixel point ratio in the common segmentation boundary as the target sub-region; wherein the second pixel point ratio is the ratio of the pixel points included in the common segmentation boundary to the pixel points included in all common segmentation boundaries;

[0098] The region to be merged is merged with the target region to obtain a new second sub-region.

[0099] The above device can execute the methods provided by all the above embodiments of the present disclosure, and has the corresponding functional modules and beneficial effects of executing the above methods. For technical details not described in detail in this embodiment, please refer to the methods provided by all the above embodiments of the present disclosure.

[0100] Reference below Fig.10 , which shows a schematic diagram of the structure of an electronic device 300 suitable for implementing the embodiment of the present disclosure. The electronic device in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc., or various forms of servers, such as independent servers or server clusters. Fig.10 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0101] like Fig.10As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only storage device (ROM) 302 or a program loaded from a storage device 308 to a random access storage device (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0102] Typically, the following devices may be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 309. The communication devices 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Although Fig.10 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0103] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing a method for recommending words. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 309, or installed from a storage device 308, or installed from a ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.

[0104] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0105] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0106] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0107] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: inputs the image to be segmented into a set vector field prediction model to obtain the vector field of each pixel point in the image to be segmented; determines the pixel point chain where each pixel point is located and the root node of the pixel point chain according to the vector field; aggregates the root node using a frame of a preset size to obtain an aggregated pixel point chain; wherein the pixels in one of the pixel point chains constitute a first sub-region; aggregates the first sub-region according to the segmentation boundary of the first sub-region to obtain multiple second sub-regions, and obtains a segmented image.

[0108] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0109] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0110] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit does not, in some cases, constitute a limitation on the unit itself.

[0111] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0112] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0113] According to one or more embodiments of the present disclosure, the present disclosure discloses an image segmentation method, including:

[0114] Inputting the image to be segmented into a set vector field prediction model to obtain the vector field of each pixel point in the image to be segmented;

[0115] Determine the pixel point chain where each pixel point is located and the root node of the pixel point chain according to the vector field; wherein the pixel point chain is composed of a plurality of pixel points that meet a preset condition, and the root node is a pixel point in the pixel point chain;

[0116] Aggregating the root nodes using a frame of a preset size to obtain an aggregated pixel point chain; wherein the pixels in one pixel point chain constitute a first sub-region;

[0117] The first sub-regions are aggregated according to the segmentation boundaries of the first sub-regions to obtain a plurality of second sub-regions, thereby obtaining a segmented image.

[0118] Furthermore, the set vector field prediction model includes: a channel exchange network, a channel segmentation network and a depth separable convolutional network;

[0119] Wherein, the depth separable convolutional network includes a first channel convolutional subnetwork, a depth convolutional subnetwork, a second channel convolutional subnetwork and a channel merging layer;

[0120] The channel exchange network, the channel splitting network, the first channel convolution subnetwork, the deep convolution subnetwork, the second channel convolution subnetwork and the channel merging layer are connected in sequence; and the output of the channel splitting network is jump-connected to the input of the channel merging layer;

[0121] The first channel convolution subnetwork includes a first channel convolution layer, a nonlinear activation layer and a linear transformation layer; the deep convolution subnetwork includes a deep convolution layer, a nonlinear activation layer and a linear transformation layer; the second channel convolution subnetwork includes a second channel convolution layer, a nonlinear activation layer and a linear transformation layer; the deep convolution layer is composed of multiple parallel convolution kernels.

[0122] Further, determining the pixel point chain where each pixel point is located and the root node of the pixel point chain according to the vector field includes:

[0123] Traversing each pixel point in the image to be segmented, taking the traversed pixel point as the current pixel point, and obtaining a set number of pixel points adjacent to the current pixel point and whose positions meet a preset condition as candidate pixel points; wherein the preset condition is located in the starting direction of the pixel point vector field;

[0124] Determine the angle between the vector field of the candidate pixel point and the vector field of the current pixel point;

[0125] Determine the candidate pixel point corresponding to the minimum angle as the parent node, and determine the parent node as the current pixel point;

[0126] Return to executing the operation of obtaining a set number of pixel points adjacent to the current pixel point and whose positions meet a preset condition as candidate pixel points, until the minimum angle is greater than a first set threshold, then determining the current node as a root node, and obtaining a pixel point chain where the traversed pixel point is located;

[0127] Continue to traverse the next pixel point that has not formed a pixel point chain until all the pixel points in the image to be segmented form a pixel point chain, and obtain multiple root nodes.

[0128] Furthermore, the root nodes are aggregated using a frame of a preset size to obtain an aggregated pixel point chain, including:

[0129] Translating the frame of the preset size in the image to be segmented according to a set step size;

[0130] The pixel point chains where at least two root nodes fall within the frame of the preset size are located are aggregated into the same pixel point chain.

[0131] Furthermore, the first sub-regions are aggregated according to the segmentation boundaries of the first sub-regions to obtain a plurality of second sub-regions, including:

[0132] For each segmentation boundary, determining a boundary strength of the segmentation boundary;

[0133] If the boundary strength is less than a second set threshold, the two first sub-regions corresponding to the segmentation boundary are aggregated to obtain a second sub-region.

[0134] Further, determining the boundary strength of the segmentation boundary includes:

[0135] For each pixel point of the segmentation boundary, calculate the angle between the vector field of the pixel point and the vector field of the Nth level parent node; wherein N≥1;

[0136] Calculate the average value of the angles corresponding to all the pixel points included in the segmentation boundary;

[0137] The average value was determined as the boundary intensity.

[0138] Furthermore, after the first sub-regions are aggregated according to the segmentation boundaries of the first sub-regions to obtain a plurality of second sub-regions, the method further includes:

[0139] Extracting a sub-region in which the first pixel ratio is less than a third set threshold value from the plurality of second sub-regions, and determining the sub-region as the sub-region to be merged; wherein the first pixel ratio is the ratio of the pixel points included in the second sub-region to the total pixel points included in the image to be segmented;

[0140] Acquire a common segmentation boundary between the sub-region to be merged and an adjacent second sub-region;

[0141] Determine the adjacent second sub-region corresponding to the segmentation boundary with the largest second pixel point ratio in the common segmentation boundary as the target sub-region; wherein the second pixel point ratio is the ratio of the pixel points included in the common segmentation boundary to the pixel points included in all common segmentation boundaries;

[0142] The region to be merged is merged with the target region to obtain a new second sub-region.

[0143] Note that the above are only preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art will understand that the present disclosure is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present disclosure. Therefore, although the present disclosure is described in more detail through the above embodiments, the present disclosure is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present disclosure, and the scope of the present disclosure is determined by the scope of the attached claims.

Claims

1. An image segmentation method, characterized in that: include: Inputting the image to be segmented into a set vector field prediction model to obtain the vector field of each pixel point in the image to be segmented; Determine the pixel point chain where each pixel point is located and the root node of the pixel point chain according to the vector field; wherein the pixel point chain is composed of a plurality of pixel points that meet a preset condition, and the root node is a pixel point in the pixel point chain; Aggregating the root nodes using a frame of a preset size to obtain an aggregated pixel point chain; wherein the pixels in one pixel point chain constitute a first sub-region; Aggregating the first sub-regions according to the segmentation boundaries of the first sub-regions to obtain a plurality of second sub-regions, thereby obtaining a segmented image; The first sub-regions are aggregated according to the segmentation boundaries of the first sub-regions to obtain a plurality of second sub-regions, and a segmented image is obtained, including: Aggregating the first sub-regions based on the boundary strength of the segmentation boundary to obtain a plurality of second sub-regions, thereby obtaining a segmented image; wherein the boundary strength of the segmentation boundary is determined by the angle between the vector field of each pixel point in the segmentation boundary and the vector field of its parent node in the pixel point chain; Determining the boundary strength of the segmentation boundary includes: For each pixel point of the segmentation boundary, calculate the angle between the vector field of the pixel point and the vector field of the Nth level parent node; wherein N≥1; Calculate the average value of the angles corresponding to all the pixel points included in the segmentation boundary; determining the average value as the boundary strength; The parent node is a corresponding candidate pixel point that satisfies the minimum angle of the pixel vector field among a set number of candidate pixels adjacent to the current pixel point in the image to be segmented and located in the starting direction of the current pixel vector field.

2. The method according to claim 1, characterized in that The set vector field prediction model includes: a channel exchange network, a channel segmentation network and a depth separable convolutional network.

3. The method according to claim 2, characterized in that The depth separable convolutional network includes a first channel convolutional subnetwork, a depth convolutional subnetwork, a second channel convolutional subnetwork and a channel merging layer; The channel exchange network, the channel splitting network, the first channel convolution subnetwork, the deep convolution subnetwork, the second channel convolution subnetwork and the channel merging layer are connected in sequence; and the output of the channel splitting network is jump-connected to the input of the channel merging layer.

4. The method according to claim 3, characterized in that The first channel convolution subnetwork includes a first channel convolution layer, a nonlinear activation layer and a linear transformation layer; the deep convolution subnetwork includes a deep convolution layer, a nonlinear activation layer and a linear transformation layer; the second channel convolution subnetwork includes a second channel convolution layer, a nonlinear activation layer and a linear transformation layer; the deep convolution layer is composed of multiple parallel convolution kernels.

5. The method according to claim 1, characterized in that Determining the pixel point chain where each pixel point is located and the root node of the pixel point chain according to the vector field includes: Traversing each pixel point in the image to be segmented, taking the traversed pixel point as the current pixel point, and obtaining a set number of pixel points adjacent to the current pixel point and whose positions meet a preset condition as candidate pixel points; wherein the preset condition is located in the starting direction of the pixel point vector field; Determine the angle between the vector field of the candidate pixel point and the vector field of the current pixel point; Determine the candidate pixel point corresponding to the minimum angle as the parent node, and determine the parent node as the current pixel point; Return to executing the operation of obtaining a set number of pixel points adjacent to the current pixel point and whose positions meet the preset conditions as candidate pixel points, until the minimum angle is greater than the first set threshold, then determine the current node as the root node, and obtain the pixel point chain where the traversed pixel point is located; Continue to traverse the next pixel point that has not formed a pixel point chain until all the pixel points in the image to be segmented form a pixel point chain, and obtain multiple root nodes.

6. The method according to claim 1, characterized in that The root node is aggregated using a frame of a preset size to obtain an aggregated pixel point chain, including: Translating the frame of the preset size in the image to be segmented according to a set step size; The pixel point chains where at least two root nodes fall within the frame of the preset size are located are aggregated into the same pixel point chain.

7. The method according to claim 5, characterized in that The first sub-regions are aggregated according to the segmentation boundaries of the first sub-regions to obtain a plurality of second sub-regions, including: For each segmentation boundary, determining a boundary strength of the segmentation boundary; If the boundary strength is less than a second set threshold, the two first sub-regions corresponding to the segmentation boundary are aggregated to obtain a second sub-region.

8. The method according to claim 1, characterized in that After the first sub-regions are aggregated according to the segmentation boundaries of the first sub-regions to obtain a plurality of second sub-regions, the method further includes: Extracting a sub-region in which the first pixel ratio is less than a third set threshold value from the plurality of second sub-regions, and determining the sub-region as the sub-region to be merged; wherein the first pixel ratio is the ratio of the pixel points included in the second sub-region to the total pixel points included in the image to be segmented; Acquire a common segmentation boundary between the sub-region to be merged and an adjacent second sub-region; Determine the adjacent second sub-region corresponding to the segmentation boundary with the largest second pixel point ratio in the common segmentation boundary as the target sub-region; wherein the second pixel point ratio is the ratio of the pixel points included in the common segmentation boundary to the pixel points included in all common segmentation boundaries; The sub-region to be merged is merged with the target sub-region to obtain a new second sub-region.

9. An image segmentation device, characterized in that: include: A vector field acquisition module is used to input the image to be segmented into a set vector field prediction model to obtain the vector field of each pixel point in the image to be segmented; A root node determination module, used to determine the pixel point chain where each pixel point is located and the root node of the pixel point chain according to the vector field; wherein the pixel point chain is composed of a plurality of pixel points that meet preset conditions, and the root node is a pixel point in the pixel point chain; A root node aggregation module, used to aggregate the root node using a frame of a preset size to obtain an aggregated pixel point chain; wherein the pixels in one pixel point chain constitute a first sub-region; a segmented image acquisition module, configured to aggregate the first sub-regions according to the segmentation boundaries of the first sub-regions to obtain a plurality of second sub-regions and obtain a segmented image; The segmented image acquisition module is further used to: aggregate the first sub-regions based on the boundary strength of the segmentation boundary to obtain multiple second sub-regions and obtain a segmented image; wherein the boundary strength of the segmentation boundary is determined by the angle between the vector field of each pixel point in the segmentation boundary and the vector field of its parent node in the pixel point chain; The segmented image acquisition module is also used for: For each pixel point of the segmentation boundary, calculate the angle between the vector field of the pixel point and the vector field of the Nth level parent node; wherein N≥1; Calculate the average value of the angles corresponding to all the pixel points included in the segmentation boundary; determining the average value as the boundary strength; The parent node is a corresponding candidate pixel point that satisfies the minimum angle of the pixel vector field among a set number of candidate pixels adjacent to the current pixel point in the image to be segmented and located in the starting direction of the current pixel vector field.

10. An electronic device, characterized in that: The electronic device comprises: one or more processing devices; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement the image segmentation method as described in any one of claims 1-8.

11. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processing device, the image segmentation method as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Direction superpixel-based rapid image segmentation method

    CN110992379A