Method, apparatus, device, and medium for dividing an image region

By redividing the mask area of the object to be detected in the specified image, the mask area is optimized to avoid character crosstalk, and role consistency in multi-character image processing is improved, and more realistic image generation is achieved.

CN119206226BActive Publication Date: 2025-07-25BEIJING SHENGSHU TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411344973.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-25
Publication Date
2025-07-25
Estimated Expiration
2044-09-25

AI Technical Summary

Technical Problem

In multi-role image processing, the prior art has problems of character information crosstalk and reduced consistency, especially when the detection box areas overlap, resulting in poor character consistency of generated images.

Method used

Through object detection and image segmentation, each object to be detected in the specified image is redivided, and the mask division is optimized, so that the target mask area of each object to be detected is reasonably expanded based on the initial mask area, and redivided using the detection frame and the segmentation mask to avoid role crosstalk and improve role consistency.

Benefits of technology

When generating images containing multiple characters, the consistency between the generated images and character information is improved, and the problem of crosstalk and consistency of character information is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119206226B_ABST
    Figure CN119206226B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a method, apparatus, device, and medium for dividing an image region, relating to the technical field of image processing. The method includes: performing object detection on a specified image to obtain the detection frame regions of each object to be detected in the specified image; the specified image includes at least one object to be detected; performing image segmentation on the specified image to obtain the initial mask regions of each object to be detected in the specified image; and respectively expanding the initial mask regions of each object to be detected based on the detection frame regions of each object to be detected to obtain the target mask regions of each object to be detected. Embodiments of the present disclosure can optimize the mask division of the objects to be detected in the specified image, so that the target mask regions of each object to be detected are reasonably expanded. Furthermore, based on the optimized target mask regions, the consistency between the generated image and the character information can be improved when generating an image containing multiple characters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a method, apparatus, device, and medium for dividing an image region. Background Art

[0002] In computer image processing, a mask is a technique used to select, filter, or hide specific regions in an image. During the image processing process, a specific image, graphic, or object can be used to cover and block all or part of the image to be processed, thereby controlling the region or process of image processing. Among them, the specific image or object used for covering and blocking is called a mask or template.

[0003] For example, given a layout image containing a person, the purpose of image processing is to combine the specified character information with the person information in the layout image to generate an image containing the specified character information and the person information in the layout image. During the image processing process, human body detection can be performed on the layout image to obtain the rectangular region of each human body in the layout image. However, using the rectangular region for character replacement to generate a new image will cause crosstalk of character information in the overlapping region. Especially when dealing with multi-person images, the processing effect of the image is not ideal. Summary of the Invention

[0004] In view of the above problems in the prior art, embodiments of the present disclosure provide a method, apparatus, device, and medium for dividing an image region. The embodiments of the present disclosure perform region re-division on each object to be detected in a specified image through object detection and image segmentation, optimize the mask division of the object to be detected in the specified image, so that the target mask region of each object to be detected is reasonably expanded based on the initial mask region, and then based on the optimized target mask region, the consistency between the generated image and the character information can be improved when generating an image containing multiple characters.

[0005] To achieve the above object, the first aspect of the present disclosure provides a method for dividing an image region, including:

[0006] Performing object detection on a specified image to obtain the detection frame region of each object to be detected in the specified image; wherein, the specified image includes at least one object to be detected;

[0007] Performing image segmentation on the specified image to obtain the initial mask region of each object to be detected in the specified image;

[0008] Based on the detection frame region of each object to be detected, respectively expanding the initial mask region of each object to be detected to obtain the target mask region of each object to be detected.

[0009] As a possible implementation of the first aspect, based on the detection box regions of each object to be detected, the initial mask regions of each object to be detected are respectively expanded to obtain the target mask regions of each object to be detected, including:

[0010] For each pixel point in the detection box regions of at least one object to be detected, determine the distance between the pixel point and the object identifier of each object to be detected;

[0011] According to the distance, determine the initial mask region of the object to be detected to which the pixel point corresponding to the distance belongs, so as to expand the initial mask region of each object to be detected and obtain the target mask region of each object to be detected.

[0012] As a possible implementation of the first aspect, for each pixel point in the detection box regions of at least one object to be detected, determining the distance between the pixel point and the object identifier of each object to be detected includes:

[0013] For each object to be detected, mark the pixel points within the initial mask region of the object to be detected as the object identifier of the object to be detected;

[0014] For each pixel point in the detection box regions of at least one object to be detected, determine the distance between the pixel point and the object identifier of each object to be detected.

[0015] As a possible implementation of the first aspect, according to the distance, determining the initial mask region of the object to be detected to which the pixel point corresponding to the distance belongs includes:

[0016] For each pixel point in the detection box regions of at least one object to be detected, determine the object identifier closest to the pixel point, and determine the initial mask region of the object to be detected corresponding to the closest object identifier as the initial mask region of the object to be detected to which the pixel point belongs.

[0017] As a possible implementation of the first aspect, for each object to be detected, marking the pixel points within the initial mask region of the object to be detected as the object identifier of the object to be detected includes:

[0018] Sequentially mark the pixel points within the initial mask regions of at least one object to be detected as the object identifiers of their corresponding objects to be detected; wherein,

[0019] In response to the existence of overlap in the initial mask regions of at least one object to be detected, the later-marked object identifier covers the previously marked object identifier.

[0020] As a possible implementation of the first aspect, based on the detection box regions of each object to be detected, the initial mask regions of each object to be detected are respectively expanded to obtain the target mask regions of each object to be detected, including:

[0021] In response to the detection box regions of at least one object to be detected overlapping, determine the union of the detection box regions of at least one object to be detected as the region to be marked;

[0022] Based on the region to be marked, the initial mask regions of each object to be detected are respectively expanded to obtain the target mask regions of each object to be detected.

[0023] As a possible implementation of the first aspect, based on the detection box regions of each object to be detected, the initial mask regions of each object to be detected are respectively expanded to obtain the target mask regions of each object to be detected, including:

[0024] According to the specified image, the detection box regions and the initial mask regions of each object to be detected, construct a marking map, where the size of the marking map is the same as the size of the specified image, and the detection box regions and the initial mask regions of each object to be detected are included in the marking map;

[0025] Based on the detection box regions of each object to be detected in the marking map, the initial mask regions of the objects to be detected in the marking map are respectively expanded to obtain the target mask regions of each object to be detected.

[0026] As a possible implementation of the first aspect, the pixel values of the pixel points in the marking map are the first preset value.

[0027] A second aspect of the present disclosure provides an apparatus for dividing an image region, including:

[0028] A detection unit for performing object detection on a specified image to obtain the detection box regions of each object to be detected in the specified image; where at least one object to be detected is included in the specified image;

[0029] A segmentation unit for performing image segmentation on the specified image to obtain the initial mask regions of each object to be detected in the specified image;

[0030] An expansion unit for respectively expanding the initial mask regions of each object to be detected based on the detection box regions of each object to be detected to obtain the target mask regions of each object to be detected.

[0031] As a possible implementation of the second aspect, the expansion unit includes:

[0032] The first determination subunit is configured to determine, for each pixel point in the detection frame area of at least one object to be detected, the distance between the pixel point and the object identifier of each object to be detected;

[0033] The second determination subunit is configured to determine, according to the distance, the initial mask area of the object to be detected to which the pixel point corresponding to the distance belongs, so as to expand the initial mask area of each object to be detected to obtain the target mask area of each object to be detected.

[0034] As a possible implementation manner of the second aspect, the first determination subunit is configured to:

[0035] For each object to be detected, mark the pixel points in the initial mask area of the object to be detected as the object identifier of the object to be detected;

[0036] For each pixel point in the detection frame area of at least one object to be detected, determine the distance between the pixel point and the object identifier of each object to be detected.

[0037] As a possible implementation manner of the second aspect, the second determination subunit is configured to:

[0038] For each pixel point in the detection frame area of at least one object to be detected, determine the object identifier closest to the pixel point, and determine the initial mask area of the object to be detected corresponding to the closest object identifier as the initial mask area of the object to be detected to which the pixel point belongs.

[0039] As a possible implementation manner of the second aspect, the first determination subunit is configured to:

[0040] Sequentially mark the pixel points in the initial mask area of at least one object to be detected as the object identifier of the corresponding object to be detected; wherein,

[0041] In response to the existence of overlap in the initial mask areas of at least one object to be detected, the later-marked object identifier covers the previously marked object identifier.

[0042] As a possible implementation manner of the second aspect, the expansion unit is configured to:

[0043] In response to the existence of overlap in the detection frame areas of at least one object to be detected, determine the union of the detection frame areas of at least one object to be detected as the area to be marked;

[0044] Based on the area to be marked, expand the initial mask area of each object to be detected respectively to obtain the target mask area of each object to be detected.

[0045] As a possible implementation manner of the second aspect, the expansion unit is configured to:

[0046] Construct a labeling map based on the specified image, the detection box regions of each object to be detected, and the initial mask regions, where the size of the labeling map is the same as that of the specified image, and the labeling map includes the detection box regions and the initial mask regions of each object to be detected;

[0047] Based on the detection box regions of each object to be detected in the labeling map, expand the initial mask regions of the objects to be detected in the labeling map respectively to obtain the target mask regions of each object to be detected.

[0048] As a possible implementation of the second aspect, the pixel values of the pixel points in the labeling map are the first preset value.

[0049] The third aspect of the present disclosure provides an electronic device, including:

[0050] A memory for storing a computer program product;

[0051] A processor for executing the computer program product stored in the memory, and when the computer program product is executed, implementing the method according to any one of the above first aspects.

[0052] The fourth aspect of the present disclosure provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, implementing the method according to any one of the above first aspects.

[0053] The fifth aspect of the present disclosure provides a computer program product, including computer program instructions, and when the computer program instructions are executed by a processor, implementing the method according to any one of the above first aspects. The technical solutions of the present disclosure will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings

[0054] The drawings forming a part of the specification depict embodiments of the present disclosure and, together with the description, are used to explain the principles of the present disclosure.

[0055] Referring to the drawings, the present disclosure can be more clearly understood according to the following detailed description, where:

[0056] Figure 1 is a flowchart of an embodiment of the method for dividing image regions of the present disclosure;

[0057] Figure 2 is a schematic diagram of a human segmentation mask of an embodiment of the method for dividing image regions of the present disclosure;

[0058] Figure 3 is a schematic diagram of mask division of an embodiment of the method for dividing image regions of the present disclosure;

[0059] Figure 4Flowchart of an embodiment of the method for dividing an image area of the present disclosure;

[0060] Figure 5 Flowchart of an embodiment of the method for dividing an image area of the present disclosure;

[0061] Figure 6 Flowchart of an embodiment of the method for dividing an image area of the present disclosure;

[0062] Figure 7 Flowchart of an embodiment of the method for dividing an image area of the present disclosure;

[0063] Figure 8 Flowchart of an embodiment of the method for dividing an image area of the present disclosure;

[0064] Figure 9 Flowchart of an embodiment of the method for dividing an image area of the present disclosure;

[0065] Figure 10 Flowchart of an embodiment of the method for dividing an image area of the present disclosure;

[0066] Figure 11 Schematic diagram of the generation effect of an embodiment of the apparatus for dividing an image area of the present disclosure;

[0067] Figure 12 Schematic diagram of the structure of an embodiment of the apparatus for dividing an image area of the present disclosure;

[0068] Figure 13 Schematic diagram of the structure of an embodiment of the apparatus for dividing an image area of the present disclosure;

[0069] Figure 14 Block diagram of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners

[0070] Although the embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.

[0071] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "said", and "the" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0072] It should be understood that although the terms "first", "second", "third", etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, the meaning of "a plurality of" is two or more, unless otherwise specifically defined.

[0073] First, the existing methods will be introduced below, and then the technical solutions of the present disclosure will be introduced in detail.

[0074] In an example of image processing, given an arbitrary image containing a person, for example, the image is a layout diagram for guiding the action postures of a character. The purpose of image processing is to combine the specified character information with the person information in the layout diagram to generate an image containing the specified character information and the person information in the layout diagram. During the process of generating images containing the specified character information multiple times, in order to ensure that the image of the same character is consistent and the effect is realistic, it is necessary to accurately identify the area corresponding to the character information in the layout diagram.

[0075] In one solution of the prior art, a rectangular human detection frame is used as the mask area for each person. Specifically, an image object detection algorithm is used to perform human detection on the given layout diagram to obtain the rectangular human detection frame for each person in the layout diagram, and the obtained rectangular human detection frame is used as the mask area for each person.

[0076] In the above solution, if the layout diagram to be processed involves a multi-person image, for the case where the rectangular detection frames of multiple persons do not overlap or have very small overlap, the interference between the information of multiple persons is relatively small. However, if there is a large area of overlap in the rectangular area, then the overlapping area belongs to multiple persons at the same time, and in this way, the person information in the overlapping area will be mutually crosstalked, which will cause image distortion and seriously affect the person consistency of the generated image.

[0077] In another solution of the prior art, an image segmentation algorithm is used to segment the persons in the layout diagram to obtain the human segmentation mask for each person in the layout diagram, and the obtained human segmentation mask is used as the mask area for each person. By adopting this solution, even if two persons in the layout diagram are very close to each other, there will be no overlapping situation, thus avoiding the crosstalk of person information and improving the person consistency of the generated image to a certain extent.

[0078] However, the mask region obtained in the above solution is strictly limited to the human region in the layout diagram. When there are significant differences in body shape, clothing, hairstyle, etc. between the human in the layout diagram and the human in the specified character information, the generated image will maintain the outline of the human in the layout diagram. In this case, the layout diagram imposes too many restrictions on the generated image, and the human image cannot be changed according to the character requirements in the generated image, resulting in the human in the generated image not conforming to the human image in the specified character information, thus reducing the human consistency of the image.

[0079] In summary, the prior art has the following defects: when dealing with multi-character images, it causes problems of character information crosstalk and reduced character consistency.

[0080] Based on the above technical problems existing in the prior art, the present disclosure provides a method, apparatus, device, and medium for dividing image regions. When dealing with multi-character image processing, on the one hand, when there is an overlapping area in the detection frame regions between characters, there is a problem of mutual crosstalk of character information in the overlapping area, which in turn affects the character consistency of the generated image. On the other hand, if the segmentation mask is directly used as the target mask region for image replacement, the final image will be strictly restricted by the character features, resulting in the inability to change the hairstyle, body posture, etc., and will also lead to a reduction in the consistency between the generated image and the character information. In summary, how to reasonably divide the image mask region is crucial for the character consistency effect of the generated image.

[0081] In view of this, the embodiments of the present disclosure re-divide the region of each object to be detected in the specified image through object detection and image segmentation, optimize the mask division of the object to be detected in the specified image, so that the target mask region of each object to be detected is reasonably expanded on the basis of the initial mask region. Furthermore, based on the optimized target mask region, when generating an image containing multiple characters, the consistency between the generated image and the character information can be improved, thus solving the technical problems of character information crosstalk and reduced character consistency mentioned in the prior art.

[0082] Figure 1 This is a flowchart of an embodiment of the method for dividing image regions of the present disclosure. As Figure 1 shown, it may specifically include:

[0083] Step S110, perform object detection on the specified image to obtain the detection frame region of each object to be detected in the specified image; wherein, the specified image includes at least one object to be detected. Among them, the object to be detected may include target objects such as humans, animals, or other objects.

[0084] Under normal circumstances, a specified image may include multiple characters. The processing tasks for the specified image may include character replacement, information filtering or masking, feature extraction, etc. for multiple characters. For example, when replacing the characters in the specified image with pre-specified character information, a reasonable division of the image mask area for the specified image can make the image processing effect better and make the generated image more realistic. Among them, the character information may include the character's figure, body proportion, clothing style, and can represent the character's appearance features, personality traits, etc.

[0085] In step S110, first, use an image object detection algorithm to perform object detection on the specified image. Image object detection refers to using theories and methods in the fields of image processing and pattern recognition to locate objects of interest in an image and give the bounding box of each object. The area enclosed by the bounding box is the detection box area. The detection box is a rectangular box used to calibrate the position and size of each object in the image. Image object detection algorithms may include RCNN (Regions with Convolutional Neural Networks), YOLO (You Only Look Once) algorithm, etc.

[0086] In one example, the image object detection algorithm includes the following steps:

[0087] (1) Preprocessing: Perform operations such as image denoising, image enhancement, and color space conversion on the image to be detected.

[0088] (2) Window sliding: Slide a window of a fixed size in the image to be detected, and use the sub-image in the window as the candidate area.

[0089] (3) Feature extraction: Use a specific algorithm to extract features from the candidate area.

[0090] (4) Feature selection: Select representative features from the feature vectors to reduce the dimension of the extracted features.

[0091] (5) Feature classification: Use a specific classifier to classify the features and determine whether the candidate area contains the target and its category.

[0092] (6) Post-processing: Merge the intersecting candidate areas determined to be of the same category, calculate the bounding box of each target, and complete the target detection.

[0093] Such as Figure 1As shown in the figure, in step S120, the specified image is segmented to obtain the initial mask region of each object to be detected in the specified image. Specifically, first, the specified image is segmented to obtain the segmentation mask of each object to be detected in the specified image, and then, based on the segmentation mask of the object to be detected, the initial mask region of the object to be detected is determined.

[0094] Image segmentation is the process of separating objects from the background in an image. The result of segmentation is the boundary contour line between the objects and the background extracted from the image. The difference between image segmentation and image object detection is that: the result of image segmentation is the contour line that fits the object boundary. See the appendix Figure 2 As shown, this contour line is usually the curve where the boundary between the object and the background lies. While the result of image object detection is usually a rectangular box.

[0095] In step S120, an image segmentation algorithm is used to segment the specified image. Among them, the image segmentation algorithm can include Mask R-CNN (Mask Region-based Convolutional Neural Network), DeepLab, etc. Among them, DeepLab is a series of semantic segmentation algorithms based on convolutional neural networks. It realizes accurate recognition and segmentation of different objects and the background by performing pixel-level classification and annotation on the image.

[0096] In one example, the image segmentation algorithm includes the following steps:

[0097] (1) Image preprocessing: This step involves operations such as grayscale conversion, filtering, and binarization on the input image to improve the accuracy of image segmentation. These operations help reduce image noise, enhance the contrast between the target and the background, and prepare for the subsequent segmentation steps.

[0098] (2) Edge detection: Use edge detection algorithms (such as Roberts algorithm, Prewitt algorithm, Sobel algorithm, etc.) to detect the edges of the image and obtain the edge map. These algorithms can identify the edge information in the image and provide a basis for subsequent connected component analysis.

[0099] (3) Connected component analysis: Perform connected component analysis on the edge map and divide the connected components as the regions of the image. The purpose of this step is to segment the image into several specific regions with unique properties.

[0100] (4) Region merging: Merge the segmented regions to eliminate small regions and improve the accuracy of segmentation. This step helps to further optimize the segmentation result and ensure that the segmented regions are more in line with the actual requirements.

[0101] Through the above steps, the image segmentation algorithm can divide an image into several specific regions with unique properties, thereby extracting the target object of interest.

[0102] See Figure 2 , after performing image segmentation on a specified image, segmentation masks of multiple objects can be obtained. The segmentation mask includes multiple pixel points in the specified image that belong to the corresponding object. The edge of the segmentation mask can be a contour curve that includes the object in the specified image.

[0103] After performing image segmentation on a specified image to obtain the segmentation masks of each object to be detected, then based on the segmentation masks of the objects to be detected, determine the initial mask regions of the objects to be detected. For each object to be detected, respectively mark the region range of the segmentation mask of each object to be detected as the initial mask region of the corresponding object to be detected. See Figure 3 for an example of Figure 3 In Figure 3 , the objects to be detected are two human bodies. According to the specified image, first obtain the human detection box regions of the two human bodies respectively, that is Figure 3 the box regions of the human bodies in

[0104] Specifically, for each object to be detected in the specified image, the set of pixel points in the segmentation mask of the object to be detected is used as the initial mask region of the object to be detected. Among them, the contour line of the segmentation mask fits the boundary of the object to be detected, so the pixel points in the segmentation mask all belong to the image range corresponding to the object to be detected. Based on this, determine the initial mask region according to the segmentation mask, and further expand the range on the basis of the initial mask region to finally obtain a reasonable target mask region. In subsequent steps, through the marking process of pixel points, the target mask region of the object to be detected is finally determined. The above pixel-based image processing method realizes high-precision image processing by directly operating on each pixel point of the image.

[0105] As Figure 1 shown, in step S130, based on the detection box region of each object to be detected, expand the initial mask region of each object to be detected respectively to obtain the target mask region of each object to be detected. Specifically, within the range of the detection box region, expand the initial mask region of the object to be detected to obtain the target mask region of the object to be detected.

[0106] In this step, the mask region of each object to be detected in the specified image is re-divided by using the detection box region obtained in step S110 and the initial mask region obtained in step S120. For example, the initial mask region is moderately expanded, and during the expansion process, the range of the detection box region can be referred to for expansion within the range of the detection box region.

[0107] In one example, for the overlapping region of the detection box regions between multiple characters, the pixel points within the overlapping region are simultaneously within the detection boxes of character A and character B. For the pixel points within the overlapping region, the distance from the pixel point to the initial mask regions of the above two characters can be compared, or the pixel value of the pixel point can be compared with the pixel values of the edge lines of the initial mask regions of the above two characters and / or other image feature values. If the distance from the pixel point to the initial mask region of character A is closer, or the pixel value and / or other image feature values of the pixel point are closer to character A, then the pixel points in the overlapping region are expanded into the initial mask region of character A.

[0108] In summary, the embodiments of the present disclosure use the detection box and the segmentation mask to re-divide the region of each object to be detected in the specified image, which can optimize the mask division of the objects to be detected in the specified image, so that the target mask region of each object to be detected is reasonably expanded on the basis of the segmentation mask, thereby improving the character consistency between the generated image and the character information without introducing character crosstalk. Furthermore, based on the optimized mask region, the consistency between the generated image and the character information can be improved when generating an image containing multiple characters.

[0109] As Figure 4 shown, in one implementation manner, Figure 1 in step S130, based on the detection box region of each object to be detected, the initial mask region of each object to be detected is expanded respectively to obtain the target mask region of each object to be detected, including:

[0110] Step S410, for each pixel point in the detection box region of at least one object to be detected, determine the distance between the pixel point and the object identifier of each object to be detected;

[0111] Step S420, according to the distance, determine the initial mask region of the object to be detected to which the pixel point corresponding to the distance belongs, so as to expand the initial mask region of each object to be detected to obtain the target mask region of each object to be detected.

[0112] In one example, the object to be detected includes character A and character B. For the overlapping area of the detection box regions between multiple characters, the pixel points within the overlapping area are simultaneously within the detection boxes of character A and character B. For the pixel points within the overlapping area, calculate the first distance from each pixel point to each object identifier of character A, where the minimum value among the first distances is LA; and calculate the second distance from each pixel point to each object identifier of character B, where the minimum value among the second distances is LB. If LA is less than LB, determine that the object to be detected to which the pixel point belongs is character A, and expand the pixel point into the initial mask region of character A to obtain the target mask region of character A.

[0113] In yet another example, the object to be detected includes character A and character B. For the pixel points within the overlapping area, calculate the distance LA from each pixel point to the initial mask region of character A, and calculate the distance LB from each pixel point to the initial mask region of character B. If LA is less than LB, determine that the object to be detected to which the pixel point belongs is character A, and expand the pixel point into the initial mask region of character A to obtain the target mask region of character A. Among them, the shortest distance from the pixel point to the edge line of the initial mask region can be used as the distance from the pixel point to the initial mask region.

[0114] As Figure 5 shown, in one implementation, Figure 4 in step S410, for each pixel point in the detection box region of at least one object to be detected, determining the distance between the pixel point and the object identifier of each object to be detected includes:

[0115] Step S510, for each object to be detected, mark the pixel points within the initial mask region of the object to be detected as the object identifier of the object to be detected;

[0116] Step S520, for each pixel point in the detection box region of at least one object to be detected, determine the distance between the pixel point and the object identifier of each object to be detected.

[0117] In one example, set the object identifier corresponding to the first object to be detected in the specified image to 1, the object identifier corresponding to the second object to be detected to 2, …… the object identifier corresponding to the Nth object to be detected to N. Then, all the pixel points within the initial mask region of the first object to be detected are marked as 1, all the pixel points within the initial mask region of the second object to be detected are marked as 2, and so on until all the initial mask regions of the objects to be detected in the specified image are marked. Refer to Figure 3 the example of

[0118] Refer to Figure 3, for each pixel point in the total box area of the person, that is, for each pixel point in the detection box area of at least one object to be detected, determine the distance between the pixel point and the object identifier of each object to be detected. See Figure 3 , on the basis of having marked object identifiers 1 and 2, for each pixel point in the total box area, respectively determine the distance between the pixel point and each object identifier, that is, respectively calculate the distance values between the pixel point and all the pixel points with object identifiers 1 and 2.

[0119] In one implementation, Figure 4 in step S420 of , according to the distance, determine the initial mask area of the object to be detected to which the pixel point corresponding to the distance belongs, including:

[0120] For each pixel point in the detection box area of at least one object to be detected, determine the object identifier closest to the pixel point, and determine the initial mask area of the object to be detected corresponding to the closest object identifier as the initial mask area of the object to be detected to which the pixel point belongs.

[0121] Specifically, if the pixel point itself is located within the initial mask area of a certain object to be detected, then the pixel point itself has been marked with the object identifier of the object to be detected. In this case, the object identifier closest to the pixel point is itself, and the distance is zero. Therefore, after the processing of step S420, the pixel point still belongs to the initial mask area of the original object to be detected. That is to say, the pixel points originally with object identifier 1 still belong to the initial mask area of the original object to be detected after the processing of step S420, and the identifier is still 1; similarly, the pixel points originally with object identifier 2 still belong to the initial mask area of the original object to be detected after the processing of step S420, and the identifier is still 2.

[0122] Specifically, if the position of the pixel point is not within the initial mask area of a certain object to be detected, then calculate the object identifier closest to the pixel point; the initial mask area corresponding to the closest object identifier is the initial mask area to which the pixel point belongs. Assume that there are only 2 objects to be detected in the specified image, and their object identifiers are 1 and 2 respectively. In step S410, the distance values between the pixel point and all the pixel points with object identifiers 1 and 2 have been calculated. In step S420, take the minimum value among these distance values. If the distance value between the pixel point and a pixel point with object identifier 1 is the smallest, then the initial mask area corresponding to object identifier 1 is the initial mask area to which the pixel point belongs, and expand the pixel point into the initial mask area corresponding to object identifier 1.

[0123] In one implementation, Figure 5In step S510 thereof, for each object to be detected, the pixel points within the initial mask region of the object to be detected are marked with the object identifier of the object to be detected, including:

[0124] Sequentially mark the pixel points within the initial mask region of at least one object to be detected with the object identifier of the corresponding object to be detected; wherein,

[0125] In response to the initial mask regions of at least one object to be detected overlapping, the subsequently marked object identifier overwrites the previously marked object identifier.

[0126] In one example, first mark all the pixel points within the initial mask region of the first object to be detected as 1, then mark all the pixel points within the initial mask region of the second object to be detected as 2. If the initial mask regions of these two objects to be detected overlap, then the mark 2 overwrites the mark 1, that is, the subsequently marked object identifier overwrites the previously marked object identifier. And so on until all the initial mask regions of the objects to be detected in the specified image are marked.

[0127] As Figure 6 shown, in one implementation manner, Figure 1 In step S130 thereof, based on the detection frame region of each object to be detected, the initial mask region of each object to be detected is respectively expanded to obtain the target mask region of each object to be detected, including:

[0128] Step S210, in response to the detection frame regions of at least one object to be detected overlapping, determine the union of the detection frame regions of at least one object to be detected as the region to be marked;

[0129] Step S215, based on the region to be marked, respectively expand the initial mask region of each object to be detected to obtain the target mask region of each object to be detected.

[0130] Taking the replacement of multiple human characters as an example, to improve the consistency between the generated image and the specified character information, the segmentation mask of the human body can be appropriately expanded. And considering avoiding information crosstalk between multiple human characters, the expansion range should not be too large. The detection frame of the object to be detected in the specified image is the human body detection frame. During the expansion process, the region range of the human body detection frame can be referred to, and the maximum expansion range does not exceed the region range of the human body detection frame. The union of all the human body detection frames in the specified image is the maximum expansion range of the mask region corresponding to each human body. In step S210, the union of all the human body detection frames in the specified image is used as the region to be marked. In subsequent steps, each pixel point in the region to be marked is respectively expanded into the target mask region corresponding to each human body.

[0131] Refer to again Figure 3For an example, according to the specified image, first obtain the human detection frames of two persons respectively, that is Figure 3 the box regions of the persons in it. The union of the box regions of the two persons is the total box region of the persons, that is, the region to be marked.

[0132] Such as Figure 7 shown, in one implementation, Figure 6 step S215 in it may specifically include:

[0133] Step S220, determine the re-marked region according to the region to be marked and the segmentation mask of at least one object to be detected in the specified image; wherein, the union of the detection frame regions of all objects to be detected is the region to be marked, and the re-marked region is the region in the region to be marked except for the mask regions of all objects to be detected, that is to say, the region to be marked includes the re-marked region and the initial mask region of each object to be detected;

[0134] In the region to be marked, the region outside the initial mask region of each object to be detected is determined as the re-marked region. Specifically, the pixel points in the initial mask region can be marked first in the region to be marked, and then the unmarked region is determined as the re-marked region.

[0135] Step S230, based on the re-marked region, expand the initial mask region of the object to be detected to obtain the target mask region corresponding to the object to be detected. Specifically, calculate the distance from each pixel point in the re-marked region to the object identifier of each object to be detected, and then determine the initial mask region of the object to be detected to which the pixel point in the re-marked region belongs according to the distance corresponding to each pixel point in the re-marked region.

[0136] For each pixel point in the re-marked region, it needs to be expanded into the mask region corresponding to each object to be detected. It can be expanded according to the distance between the pixel point and each object to be detected, and each pixel point in the re-marked region is expanded into the mask region corresponding to the object to be detected with the closest distance to it.

[0137] In the above embodiments, the maximum region range for expanding the mask region of the object to be detected in the specified image is determined based on the region to be marked, and the pixel points in the re-marked region are expanded into the mask region of the object based on the segmentation mask. Based on the above scheme, the mask region division is more reasonable, and better image processing effects can be achieved based on the reasonably divided mask.

[0138] In one implementation, Figure 7 step S220 in it, determine the re-marked region according to the region to be marked and the segmentation mask of at least one object to be detected in the specified image, includes:

[0139] Mark the pixel points within the initial mask region with the mask identifier corresponding to the object to be detected;

[0140] Take the pixel points in the region to be marked that are not marked as the mask identifier as the points to be marked; determine the set of all points to be marked as the region to be re-marked.

[0141] Still taking the replacement of multiple characters as an example, according to the segmentation mask of each human body in the specified image, mark the initial mask region of each human body in the region to be marked respectively, which specifically includes: for each human body in the specified image, take the set of pixel points within the segmentation mask of each human body as the initial mask region of each human body respectively; and mark the pixel points within the segmentation mask of each human body with the mask identifier corresponding to each human body.

[0142] Specifically, take the union of all human body detection boxes in the specified image as the maximum region range that the characters can expand. In one example, mark all pixel points within the regions of all human body detection boxes in the specified image as 255. All pixel points marked as 255 constitute the union of all human body detection boxes, that is, the region to be marked. See Figure 3 In the example of, mark all pixel points in the total box region of the character as 255.

[0143] Mark on the marking map successively according to the segmentation masks of each human body in the specified image, and mark the pixel points within the segmentation mask of each human body with the mask identifier corresponding to each human body. For example, the pre-set mask flag value can be used as the mask identifier. In one example, set the mask flag value corresponding to the first human body as 1, the mask flag value corresponding to the second human body as 2, until the mask flag value corresponding to the Nth human body is N. Then mark all pixel points within the segmentation mask of the first human body as 1, mark all pixel points within the segmentation mask of the second human body as 2, and so on until all segmentation masks of human bodies in the specified image are marked. The set of pixel points with the same mask flag value constitutes the initial mask region of the corresponding human body. See Figure 3 In the example of, in the total box region of the character, the marks of the pixel points within the seg regions of the two characters can be marked as 1 and 2 respectively.

[0144] After the pixel points in the initial mask region are marked, continue to mark the unmarked pixel points in the region to be marked. As described above, in the specified image, first mark all pixel points within all human detection frame regions as 255, that is, all pixel points marked as 255 form the region to be marked. Then, the pixel points within each human segmentation mask in the region to be marked are sequentially marked as the corresponding mask flag values. Then, the pixel points in the region to be marked that are not marked as mask flag values are the points to be marked, that is, all pixel points marked as 255 are the points to be marked. After that, re-mark all the points marked as 255 in the specified image, and determine the set of all points to be marked as the re-marked region.

[0145] In subsequent steps, by re-marking the pixel points in the re-marked region and marking them as the mask flag value of a certain human body, the mask region of the human body is reasonably expanded and there is no intersection between the mask regions of each human body, so as to achieve the effect of improving the consistency of the characters and having no crosstalk.

[0146] In one implementation, Figure 7 In step S230 of, based on the re-marked region, expand the initial mask region of the object to be detected to obtain the target mask region corresponding to the object to be detected, including:

[0147] According to the distances from the pixel points in the re-marked region to the initial mask regions of each object to be detected, expand the initial mask region of each object to be detected to obtain the expanded mask region of each object to be detected;

[0148] Take the expanded mask region of the object to be detected as the target mask region corresponding to the object to be detected.

[0149] For each pixel point in the re-marked region, it needs to be expanded into the mask region corresponding to each object to be detected. Specifically, the distances from the pixel points in the re-marked region to the initial mask regions of each object to be detected can be calculated, and the pixel point is expanded into the initial mask region with the shortest distance. Among them, the shortest distance from the pixel point to the edge line of the initial mask region can be used as the distance from the pixel point to the initial mask region. After all the pixel points in the re-marked region are expanded, the expanded mask region of each object to be detected is finally obtained. See Figure 3 In the marked region, all the pixel points in the previous entire region to be marked have been marked. In the marked region, the dark image region and the light image region respectively represent the expanded mask regions of two human bodies.

[0150] See Figure 3, the extended mask regions of the two characters are respectively extracted from the marked regions to obtain the final Mask region (mask region) of the characters. Based on the distance from the pixel points to the initial mask regions of each object to be detected, the finally obtained mask region reasonably extends the initial mask regions, which can prevent the mutual interference of character information in the overlapping region of the human detection frames, and can also avoid the reduction of character consistency caused by the excessive limitation of the human segmentation mask, and can achieve a realistic and natural image effect.

[0151] Figure 8 This is a flowchart of an embodiment of the method for dividing the image mask region of the present disclosure. As Figure 8 shown, in one implementation manner, according to the distance from each pixel point in the relabeled region to the initial mask region of each object to be detected, the initial mask region of each object to be detected is extended to obtain the extended mask region of each object to be detected, including:

[0152] Step S310, for all the points to be labeled, calculate the initial mask region closest to the point to be labeled in the region to be labeled respectively;

[0153] Step S320, label the point to be labeled with the mask identifier corresponding to the initial mask region closest to it; label the pixel points within the initial mask region with the mask identifier corresponding to the object to be detected;

[0154] Step S330, according to the mask identifiers of all the pixel points in the region to be labeled, obtain the extended mask region of each object to be detected.

[0155] For all the points to be labeled, calculate the distance between the point to be labeled and the initial mask region of each object to be detected in the specified image respectively. For example, there are N humans in the specified image. If the shortest distance from the point to be labeled to the edge line of the initial mask region of the Mth human is the smallest, then the Mth human is determined as the initial mask region closest to the point to be labeled. Label the point to be labeled with the mask flag value M corresponding to the Mth human, that is, extend the point to be labeled into the mask region of the Mth human. After all the pixel points in the region to be labeled are relabeled, the set of pixel points with the same mask flag value constitutes the extended mask region of the corresponding human. Refer to Figure 3 the example in. The labels of the pixel points in the seg regions of the two characters are labeled as 1 and 2 respectively, and then the other points still labeled as 255 in the specified image are relabeled to obtain the extended mask region. After all the pixel points are labeled, the Figure 3 marked region in

[0156] In the above embodiments, through distance calculation and marking processing for each point to be marked, the points to be marked are reasonably divided into the corresponding mask regions of each object to be detected, and then the image processing effect is improved based on the reasonable mask regions.

[0157] In one implementation, Figure 8 in step S310 of, calculating the initial mask region closest to the point to be marked in the region to be marked includes:

[0158] Calculating the distance values from the point to be marked to all the pixel points in the initial mask region of each object to be detected in the region to be marked respectively;

[0159] Determining the initial mask region corresponding to the minimum value among all the distance values as the initial mask region closest to the point to be marked.

[0160] In the example of character replacement, for all the points to be marked, calculate the distance values from all the pixel points in the initial mask region of each human body to the point to be marked respectively. Take the pixel point corresponding to the minimum value among all the distance values, which is the pixel point closest to the point to be marked in all the initial mask regions. The initial mask region of the human body where this pixel point is located is the initial mask region of the human body closest to the point to be marked. And, re-mark the point to be marked with the mask flag value of this pixel point.

[0161] The above embodiments calculate the distance by traversing all the initial mask regions, so that the points to be marked are finally divided into the mask region with the closest distance, and then the image processing effect is improved based on the reasonably divided mask regions.

[0162] In one implementation, based on the detection frame region of each object to be detected, expand the initial mask region of each object to be detected respectively to obtain the target mask region of each object to be detected, including:

[0163] Construct a marking map according to the specified image, the detection frame region and the initial mask region of each object to be detected, where the size of the marking map is the same as the size of the specified image, and the marking map includes the detection frame region and the initial mask region of each object to be detected;

[0164] Based on the detection frame region of each object to be detected in the marking map, expand the initial mask region of the object to be detected in the marking map respectively to obtain the target mask region of each object to be detected.

[0165] In one implementation, the pixel value of the pixel point in the marking map is the first preset value. For example, the first preset value can be set to 0.

[0166] Specifically, still taking the replacement of multiple character roles as an example, the union of all human detection boxes in the specified image is used as the maximum expandable area range of the characters. An image with the same size as the specified image and all pixel points marked as 0 can be constructed as a marking map. First, all pixel points within the areas of all human detection boxes in the marking map are marked as 255. All pixel points marked as 255 constitute the union of all human detection boxes, that is, the area to be marked. Refer to Figure 3 For the example in, all pixel points in the total area of the character's box are marked as 255. Through this embodiment, there is no interference information in the marking map and the pixel values of the pixel points in the marking map are uniformly marked, which is convenient for quickly calculating the target mask area of each object to be detected.

[0167] According to each human body segmentation mask in the specified image, marks are made on the marking map in sequence, and the pixel points within the segmentation mask of each human body are respectively marked with the mask identifier corresponding to each human body. The pre-set mask flag value can be used as the mask identifier. For example, the mask flag value corresponding to the first human body is set to 1, the mask flag value corresponding to the second human body is set to 2, until the mask flag value corresponding to the Nth human body is N. Then all pixel points within the first human body segmentation mask are marked as 1, all pixel points within the second human body segmentation mask are marked as 2, and so on until all human body segmentation masks in the specified image are marked. The set of pixel points with the same mask flag value constitutes the initial mask area of the corresponding human body. Refer to Figure 3 For the example in, in the total area of the character's box, the pixel points within the seg areas of the two characters can be marked as 1 and 2 respectively.

[0168] After the pixel points within the initial mask area are marked, continue to mark the pixel points that have not been marked in the area to be marked. As described above, in the above-mentioned marking map, first, all pixel points within the areas of all human detection boxes are marked as 255, that is, all pixel points marked as 255 constitute the area to be marked. Then, the pixel points within the segmentation masks of each human body in the area to be marked are sequentially marked with the corresponding mask flag values. Then, the pixel points that have not been marked as mask flag values in the area to be marked are the points to be marked, that is, all pixel points marked as 255 are the points to be marked. After that, all points marked as 255 in the marking map are re-marked, and the set of all points to be marked is determined as the re-marked area. Finally, according to the distances from the pixel points in the re-marked area to the initial mask areas of each object to be detected, the initial mask areas of each object to be detected are expanded to obtain the target mask areas corresponding to each object to be detected. The specific implementation steps for realizing the expansion in the marking map are similar to the way of realizing the expansion on the specified image. For the specific steps, beneficial effects or technical problems to be solved, reference can be made to the descriptions in the above-mentioned embodiments or the description in the summary of the invention, which will not be elaborated here one by one.

[0169] In summary, taking the replacement of multiple characters as an example, an embodiment of the method for dividing the image area disclosed in the present disclosure may include the following steps:

[0170] S1: Perform human detection and human segmentation on the specified image to obtain a human detection frame and a human segmentation mask.

[0171] S2: Construct an image with the same size as the specified image and all pixels marked as 0 as a marking map, and mark the pixels within the human detection frame in the marking map as 255.

[0172] S3: According to the coordinates of the human segmentation mask, sequentially mark the pixels within the human segmentation mask area as corresponding human marking points. The first human is marked as 1, the second human is marked as 2, and so on. The marked number is the mask identifier.

[0173] S4: Take each pixel marked as 255 in the marking map as a point to be marked, and re-mark the point to be marked. That is, calculate the distance from each human marking point to the point to be marked, find the human marking point closest to the point to be marked, and re-mark the current point to be marked with the mask identifier of this human marking point.

[0174] S5: After all points to be marked are marked, starting from 1 in the marking order, sequentially extract the mask area of each human from the marking map as the mask area used for the final character replacement.

[0175] Using the mask area divided by the embodiment disclosed in the present disclosure, the image can be further processed to generate an image that meets the requirements. As Figure 9 shown, an example of image generation may specifically include the following steps:

[0176] Step S100: Receive a specified image and an image generation prompt, where the image generation prompt includes the role information corresponding to each object to be detected in the specified image;

[0177] Steps S110 to S130: Use the method for dividing the image area to obtain the target mask area of each object to be detected in the specified image;

[0178] Step S150: Use the role information corresponding to each object to be detected to separately replace the target mask area of each object to be detected, and generate an object replacement image corresponding to the specified image.

[0179] Figure 10 Shows a flow example diagram for generating a character replacement image. In an example of image generation, the input information and output information of the character replacement task are respectively:

[0180] Input information: A layout diagram, designated role information for each role, and corresponding text. Among them, the layout diagram is used to guide the action postures of the roles. The designated role information includes information that can represent the characteristics of the role's appearance, personality, etc., specifically including the character model, body proportion, clothing style, etc. of the role, and the role information can be shown using a role diagram. The corresponding text is used to indicate the action postures or dressing styles of the roles, etc.

[0181] Output information: A diagram in which multiple roles are generated according to the postures of the layout diagram and the instructions of the text.

[0182] See Figure 9 and Figure 10 , first, take the layout diagram as the designated image, perform human detection and human segmentation on the layout diagram to obtain a human detection box and a human segmentation mask respectively. According to the image region division method provided by the present disclosure, construct a labeling diagram based on the human detection box and the human segmentation mask. Automatically divide the human regions in the labeling diagram to obtain the mask regions corresponding to each human body in the layout diagram. Then extract the mask regions of the characters, use the mask regions for character replacement, and finally generate the character replacement effect diagram. For the implementation details of the image region division, reference can be made to Figures 1 to 8 the relevant descriptions of the embodiments in

[0183] In the above example, the role information can be input during the character replacement stage. The input role information can include the image of the role, the description of the role, etc. These role information and the Mask regions of the characters are in one-to-one correspondence, that is, for each Mask, there is a corresponding role information input. There can be various ways of character replacement. For example, the IPAdapter tool can be used for character replacement.

[0184] In the above example, the IPAdapter tool itself does not use the mask limit to generate images. In the generated image, it is necessary to maintain the expected layout structure in the layout diagram. Based on the embodiments of the present disclosure, when generating the image, the mask can be used as guiding information to replace the mask regions corresponding to the human bodies in the layout diagram, rather than replacing the entire image. Specifically, the information of the mask regions corresponding to the human bodies divided in the layout diagram can be added to the input information of the character replacement task. This information serves as guiding information and plays a supervisory role, indicating to replace the image within the reasonable regions divided to avoid information crosstalk and reduced character consistency caused by overly large or small regions.

[0185] Figure 11 This is a schematic diagram of the generation effect of an embodiment of the image region division method of the present disclosure. In Figure 11In the example, replace the two corresponding characters in the layout diagram with reference to the character images in the two provided character diagrams. Among them, one of the characters in the character diagram is wearing red clothes and has long curly hair. Figure 11 The effects of the generated images of the prior art solution and the solution of the present disclosure are compared. It can be clearly seen that the hairstyle of the character wearing red clothes in the generated image of the prior art solution is short hair. It can be seen that the generated image of the prior art solution cannot change the character image according to the role requirements, and the characters in the generated image do not conform to the character images in the character diagram, and the character consistency of the image is poor. In contrast, the hairstyle of the character wearing red clothes in the generated image of the present disclosure is long curly hair. It can be seen that the present disclosure reasonably optimizes the mask division of the human body area in the layout diagram, so that the mask area corresponding to each human body is reasonably enlarged, thereby improving the character consistency between the generated image and the character diagram.

[0186] Such as Figure 12 As shown, the present disclosure also provides an embodiment of a corresponding device for dividing an image area. For the beneficial effects or technical problems solved by this device, reference can be made to the descriptions in the methods corresponding to each device, or to the descriptions in the summary of the invention, which will not be elaborated here one by one.

[0187] In the embodiment of the device for dividing an image area, the device includes:

[0188] A detection unit 100, configured to: perform object detection on a specified image to obtain the detection frame area of each object to be detected in the specified image; wherein, the specified image includes at least one object to be detected;

[0189] A segmentation unit 200, configured to: perform image segmentation on the specified image to obtain the initial mask area of each object to be detected in the specified image;

[0190] An expansion unit 300, configured to: based on the detection frame area of each object to be detected, respectively expand the initial mask area of each object to be detected to obtain the target mask area of each object to be detected.

[0191] Such as Figure 13 As shown, in one implementation manner, the expansion unit 300 includes:

[0192] A first determination subunit 310, configured to determine the distance between each pixel point of the detection frame area of at least one object to be detected and the object identifier of each object to be detected;

[0193] A second determination subunit 320, configured to determine the initial mask area of the object to be detected to which the pixel point corresponding to the distance belongs according to the distance, so as to expand the initial mask area of each object to be detected to obtain the target mask area of each object to be detected.

[0194] In one implementation, the first determination subunit 310 is configured to:

[0195] For each object to be detected, mark the pixel points within the initial mask region of the object to be detected with the object identifier of the object to be detected;

[0196] For each pixel point in the detection frame region of at least one object to be detected, determine the distance between the pixel point and the object identifier of each object to be detected.

[0197] In one implementation, the second determination subunit 320 is configured to:

[0198] For each pixel point in the detection frame region of at least one object to be detected, determine the object identifier closest to the pixel point, and determine the initial mask region of the object to be detected corresponding to the closest object identifier as the initial mask region of the object to be detected to which the pixel point belongs.

[0199] In one implementation, the first determination subunit 310 is configured to:

[0200] Sequentially mark the pixel points within the initial mask region of at least one object to be detected with their corresponding object identifiers of the objects to be detected; wherein,

[0201] In response to the existence of overlap in the initial mask regions of at least one object to be detected, the later-marked object identifier overwrites the previously marked object identifier.

[0202] In one implementation, the expansion unit 300 is configured to:

[0203] In response to the existence of overlap in the detection frame regions of at least one object to be detected, determine the union of the detection frame regions of at least one object to be detected as the region to be marked;

[0204] Based on the region to be marked, expand the initial mask region of each object to be detected respectively to obtain the target mask region of each object to be detected.

[0205] In one implementation, the expansion unit 300 is configured to:

[0206] Construct a marking map according to the specified image and the detection frame region and initial mask region of each object to be detected, wherein the size of the marking map is the same as the size of the specified image, and the marking map includes the detection frame region and initial mask region of each object to be detected;

[0207] Based on the detection frame region of each object to be detected in the marking map, expand the initial mask region of the object to be detected in the marking map respectively to obtain the target mask region of each object to be detected.

[0208] In one embodiment, the pixel value of the pixel points in the marked graph is set to a first preset value.

[0209] Next, an electronic device according to an embodiment of the present disclosure will be described with reference to Figure 14 The electronic device may be either the first device or the second device, or both, or a stand-alone device independent of them. The stand-alone device may communicate with the first device and the second device to receive the input signals collected from them.

[0210] Figure 14 The block diagram of an electronic device according to an embodiment of the present disclosure is illustrated.

[0211] As Figure 14 shown, the electronic device includes one or more processors and a memory.

[0212] The processor may be a central processing unit (CPU) or other form of processing unit having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.

[0213] The memory may store one or more computer program products. The memory may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program products may be stored on the computer-readable storage media, and the processor may run the computer program products to implement the methods of the various embodiments of the present disclosure described above and / or other desired functions.

[0214] In one example, the electronic device may further include: an input device and an output device, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0215] In addition, the input device may further include, for example, a keyboard, a mouse, etc.

[0216] The output device may output various information to the outside, including the determined distance information, direction information, etc. The output device may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0217] Of course, for simplicity, Figure 14 only some of the components related to the present disclosure in the electronic device are shown in and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device may further include any other appropriate components.

[0218] In addition to the above methods and devices, embodiments of the present disclosure may also be computer program products, which include computer program instructions that, when run by a processor, cause the processor to execute the steps in the methods according to various embodiments of the present disclosure described in the above part of this specification.

[0219] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0220] Furthermore, embodiments of the present disclosure may also be computer-readable storage media, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the methods according to various embodiments of the present disclosure described in the above part of this specification.

[0221] The computer-readable storage media may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0222] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above specific details are only for illustrative and facilitating understanding purposes, rather than limitations, and the above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0223] In each embodiment described in this specification, a progressive approach is adopted. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For system embodiments, since they basically correspond to method embodiments, the description is relatively simple. For relevant parts, reference can be made to the corresponding descriptions in the method embodiments.

[0224] The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with each other.

[0225] The methods and apparatuses of this disclosure can be implemented in many ways. For example, the methods and apparatuses of this disclosure can be implemented through software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of the steps for the methods is only for illustration. The steps of the methods of this disclosure are not limited to the specific order described above, unless otherwise specifically stated in other ways. In addition, in some embodiments, this disclosure can also be implemented as a program recorded on a recording medium, and these programs include machine-readable instructions for implementing the methods according to this disclosure. Therefore, this disclosure also covers the recording medium storing the programs for executing the methods according to this disclosure.

[0226] It should also be noted that in the apparatuses, equipment, and methods of this disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this disclosure.

[0227] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects are very obvious to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but to the broadest scope consistent with the principles and novel features disclosed herein.

[0228] The foregoing description has been presented for purposes of illustration and description. In addition, this description is not intended to limit embodiments of the present disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.

Claims

1. A method for dividing an image area, characterized in that Including: Performing object detection on a specified image to obtain a detection box region for each object to be detected in the specified image; wherein, a plurality of objects to be detected are included in the specified image; Performing image segmentation on the specified image to obtain an initial mask region for each object to be detected in the specified image; Based on the detection box region of each object to be detected, respectively expanding the initial mask region of each object to be detected to obtain a target mask region for each object to be detected; the step of based on the detection box region of each object to be detected, respectively expanding the initial mask region of each object to be detected to obtain a target mask region for each object to be detected includes: For each pixel point in the detection box regions of the plurality of objects to be detected, determining the distance between the pixel point and the object identifier of each object to be detected; According to the distance, determining the initial mask region of the object to be detected to which the pixel point corresponding to the distance belongs, so as to expand the initial mask region of each object to be detected to obtain a target mask region for each object to be detected.

2. The method according to claim 1, wherein The step of for each pixel point in the detection box regions of the plurality of objects to be detected, determining the distance between the pixel point and the object identifier of each object to be detected includes: For each object to be detected, marking the pixel points within the initial mask region of the object to be detected as the object identifier of the object to be detected; For each pixel point in the detection box regions of the plurality of objects to be detected, determining the distance between the pixel point and the object identifier of each object to be detected.

3. The method according to claim 2, wherein The step of according to the distance, determining the initial mask region of the object to be detected to which the pixel point corresponding to the distance belongs includes: For each pixel point in the detection box regions of the plurality of objects to be detected, determining the object identifier closest to the pixel point, and determining the initial mask region of the object to be detected corresponding to the closest object identifier as the initial mask region of the object to be detected to which the pixel point belongs.

4. The method according to claim 2, wherein The step of for each object to be detected, marking the pixel points within the initial mask region of the object to be detected as the object identifier of the object to be detected includes: Sequentially marking the pixel points within the initial mask regions of the plurality of objects to be detected as the object identifiers of the corresponding objects to be detected; wherein, In response to an overlap existing in the initial mask regions of the plurality of objects to be detected, the later marked object identifier overwrites the previously marked object identifier.

5. The method according to any one of claims 1 - 4, wherein: In response to an overlap existing in the detection box regions of the plurality of objects to be detected, determining the union of the detection box regions of the plurality of objects to be detected as the region to be marked; Based on the region to be marked, respectively expanding the initial mask region of each object to be detected to obtain a target mask region for each object to be detected.

6. The method according to any one of claims 1 - 4, wherein: Construct a labeled map according to the specified image, the detection box regions of each object to be detected, and the initial mask regions, where the size of the labeled map is the same as the size of the specified image, and the labeled map includes the detection box regions and the initial mask regions of each object to be detected; Based on the detection box regions of each object to be detected in the labeled map, expand the initial mask regions of the objects to be detected in the labeled map respectively to obtain the target mask region of each object to be detected.

7. The method according to claim 6, wherein The pixel value of the pixel points in the labeled map is the first preset value.

8. An apparatus for dividing an image region, characterized in that, It includes: A detection unit for performing object detection on a specified image to obtain the detection box regions of each object to be detected in the specified image; where there are multiple objects to be detected in the specified image; A segmentation unit for performing image segmentation on the specified image to obtain the initial mask regions of each object to be detected in the specified image; An expansion unit for expanding the initial mask regions of each object to be detected respectively based on the detection box regions of each object to be detected to obtain the target mask region of each object to be detected; The expansion unit includes: A first determination subunit for determining, for each pixel point in the detection box regions of the multiple objects to be detected, the distance between the pixel point and the object identifier of each object to be detected; A second determination subunit for determining, according to the distance, the initial mask region of the object to be detected to which the pixel point corresponding to the distance belongs, so as to expand the initial mask region of each object to be detected to obtain the target mask region of each object to be detected.

9. An electronic device, characterized in that, It includes: A memory for storing a computer program product; A processor for executing the computer program product stored in the memory, and when the computer program product is executed, implementing the method according to any one of claims 1-7 above.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1-7 above is implemented.

Citation Information

Patent Citations

  • Target detection positioning confidence determination method and device, electronic equipment and storage medium

    CN112668573A