Object segmentation method, device, equipment and storage medium

By semantic recognition and color value clustering of images, the target mask map is generated, which solves the problems of missing segmentation and miss segmentation in sky segmentation, and improves the accuracy of segmentation.

CN114494298BActive Publication Date: 2025-07-11BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210107771.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-28
Publication Date
2025-07-11
Estimated Expiration
2042-01-28

AI Technical Summary

Technical Problem

The prior art has problems of missed segmentation and missed segmentation in sky segmentation, especially the method based on color information fails to images with small color aberrations.

Method used

The initial mask map is obtained by semantic recognition of the image to be segmented, and the initial target object area is determined based on the initial mask map, and the pixel points in the area are clustered in color values, multiple color classifications are obtained, and the confidence of the mask map is adjusted using the difference map, and finally the target mask map is generated for segmentation.

Benefits of technology

Effectively prevent missed segmentation of objects, improving the accuracy and accuracy of sky segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114494298B_ABST
    Figure CN114494298B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose an object segmentation method, apparatus, device, and storage medium. The method includes: performing semantic recognition on a target object in an image to be segmented to obtain an initial mask image; determining an initial target object region in the image to be segmented based on the initial mask image; performing clustering processing on pixel points in the initial target object region according to color values to obtain N color classifications of the target object; obtaining N difference images according to the N color classifications and the image to be segmented; determining a target mask image according to the N difference images and the initial mask image; and segmenting the target object in the image to be segmented based on the target mask image. The object segmentation method provided by the embodiments of the present disclosure determines a target mask image according to the difference images and the initial mask image, and thus segments the target object based on the target mask image, which can achieve the segmentation of the object in the image, prevent the missed segmentation of the object, and improve the accuracy of object segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of image processing technologies, and in particular, to an object segmentation method, apparatus, device, and storage medium. Background Art

[0002] Currently, there are the following two implementation methods for sky segmentation: One is to use a deep learning algorithm of a convolutional neural network to segment the sky. In this way, there is a situation of partial missing segmentation in the middle of the segmented mask image; the other is a traditional algorithm based on color information. This method relies on the color of the sky for segmentation and may have the phenomenon of missegmentation. For images with small color differences, it may also fail. Summary of the Invention

[0003] Embodiments of the present disclosure provide an object segmentation method, apparatus, device, and storage medium to achieve the segmentation of an object in an image, prevent missing segmentation of the object, and improve the accuracy of object segmentation.

[0004] In a first aspect, embodiments of the present disclosure provide an object segmentation method, including:

[0005] Performing semantic recognition on a target object in an image to be segmented to obtain an initial mask image;

[0006] Determining an initial target object region in the image to be segmented based on the initial mask image;

[0007] Performing clustering processing on pixel points in the initial target object region according to color values to obtain N color classifications of the target object; where N is a positive integer greater than or equal to 1;

[0008] Obtaining N difference maps according to the N color classifications and the image to be segmented;

[0009] Determining a target mask image according to the N difference maps and the initial mask image;

[0010] Segmenting the target object in the image to be segmented based on the target mask image.

[0011] In a second aspect, embodiments of the present disclosure further provide an object segmentation apparatus, including:

[0012] An initial mask image acquisition module, configured to perform semantic recognition on a target object in an image to be segmented to obtain an initial mask image;

[0013] An initial target object region determination module, configured to determine an initial target object region in the image to be segmented based on the initial mask image;

[0014] A clustering module, configured to perform clustering processing on pixel points in the initial target object region according to color values to obtain N color classifications of the target object, where N is a positive integer greater than or equal to 1;

[0015] A difference map obtaining module, configured to obtain N difference maps according to the N color classifications and the image to be segmented;

[0016] A target mask map obtaining module, configured to determine a target mask map according to the N difference maps and the initial mask map;

[0017] An image segmentation module, configured to segment the target object in the image to be segmented based on the target mask map.

[0018] In a third aspect, an embodiment of the present disclosure further provides an electronic device, where the electronic device includes:

[0019] One or more processing devices;

[0020] A storage device, configured to store one or more programs;

[0021] When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement the object segmentation method as described in the embodiments of the present disclosure.

[0022] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, the object segmentation method as described in the embodiments of the present disclosure is implemented.

[0023] Embodiments of the present disclosure disclose an object segmentation method, apparatus, device, and storage medium. Semantic recognition is performed on a target object in an image to be segmented to obtain an initial mask map; an initial target object region in the image to be segmented is determined based on the initial mask map; clustering processing is performed on pixel points in the initial target object region according to color values to obtain N color classifications of the target object; N difference maps are obtained according to the N color classifications and the image to be segmented; a target mask map is determined according to the N difference maps and the initial mask map; and the image to be segmented is segmented based on the target mask map. The object segmentation method provided by the embodiments of the present disclosure determines a target mask map according to a difference map and an initial mask map, and thus segments the target object in the image to be segmented based on the target mask map, which can implement the segmentation of an object in an image, prevent the missed segmentation of the object, and improve the accuracy of object segmentation. Description of the Drawings

[0024] Figure 1 is a flowchart of an object segmentation method in an embodiment of the present disclosure;

[0025] Figure 2aIt is an example diagram of the image to be segmented in the embodiments of the present disclosure;

[0026] Figure 2b It is an example diagram of the initial mask in the embodiments of the present disclosure;

[0027] Figure 2c It is an example diagram of the difference map in the embodiments of the present disclosure;

[0028] Figure 2d It is an example diagram of the target mask in the embodiments of the present disclosure;

[0029] Figure 2e It is a visualization diagram generated based on the initial mask in the embodiments of the present disclosure;

[0030] Figure 2f It is a visualization diagram generated based on the target mask in the embodiments of the present disclosure;

[0031] Figure 3 It is a schematic structural diagram of an object segmentation device in the embodiments of the present disclosure.

[0032] Figure 4 It is a schematic structural diagram of an electronic device in the embodiments of the present disclosure. Detailed implementation manners

[0033] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the accompanying drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0034] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0035] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0036] It should be noted that the concepts such as "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or mutual dependence relationship of the functions performed by these devices, modules or units.

[0037] It should be noted that the modification of "one" and "multiple" mentioned in this disclosure is illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0038] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes, and are not used to limit the scope of these messages or information.

[0039] Embodiment 1

[0040] Figure 1 As shown in the flowchart of an object segmentation method provided for Embodiment 1 of this disclosure, this embodiment is applicable to the situation of segmenting a target object in an image. This method can be executed by an object segmentation device, which can be composed of hardware and / or software and is generally integrated in a device with object segmentation function. The device can be an electronic device such as a server, a mobile terminal or a server cluster. As Figure 1 shown, the method specifically includes the following steps:

[0041] Step 110, perform semantic recognition on the target object in the image to be segmented to obtain an initial mask image.

[0042] Among them, the target object can be any object that needs to be segmented from the image, such as: vehicle, tree, building, sky, etc. In this embodiment, it is mainly for the segmentation of "sky". The size of the initial mask image is the same as that of the image to be segmented, and the gray value of each pixel represents the confidence that the pixel belongs to the target object. Specifically, perform semantic recognition on each pixel of the image to be segmented, determine the confidence that each pixel belongs to the target object, and determine the gray value of each pixel according to the confidence, so as to obtain the initial mask image. Exemplarily, assume that the confidence that a certain pixel belongs to the target object is 200 / 255, then set the gray value of this pixel to 200.

[0043] Optionally, the process of performing semantic recognition on the target object in the image to be segmented to obtain an initial mask image can be: input the image to be segmented into a target object recognition model and output the initial mask image.

[0044] Among them, the target object recognition model can be obtained by training a neural network model with image segmentation data. Input the image to be segmented into the target object recognition model, and output the confidence that each pixel belongs to the target object, so as to obtain the initial mask image. Exemplarily,Figure 2a is the image to be segmented (the original image is a color image), Figure 2b is the initial mask image. Figure 2b is Figure 2a the mask image obtained after performing "sky" recognition on , the closer the gray level is to white, the greater the probability that the pixel point is "sky". In this embodiment, by using the target object recognition model to recognize the target object, the recognition accuracy and efficiency of the target object can be improved.

[0045] Step 120: Determine the initial target object region in the image to be segmented based on the initial mask image.

[0046] Among them, the initial target object region can be understood as the region composed of the target objects determined according to the initial mask image.

[0047] Specifically, the method for determining the initial target object region in the image to be segmented based on the initial mask image can be: obtain the pixel points in the initial mask image with a confidence level greater than the first set value, and determine them as the first target points; determine the region formed by the pixel points corresponding to the first target points in the image to be segmented as the initial target object region.

[0048] Among them, the first set value can be any value between 180 / 255 - 220 / 255. Specifically, determining the pixel points in the initial mask image with a confidence level greater than the first set value as the first target points indicates that the probability that the pixel points corresponding to the first target points in the image to be segmented belong to the target object is greater than the first set value. Therefore, the region formed by the pixel points corresponding to the first target points in the image to be segmented is determined as the initial target object region. In this embodiment, determining the region surrounded by the pixel points with a confidence level greater than the first set value as the initial target object region can roughly segment out the target object.

[0049] Step 130: Perform clustering processing on the pixel points in the initial target object region according to the color values to obtain N color classifications of the target object.

[0050] Among them, N is a positive integer greater than or equal to 1. For example, if N is taken as 3, then clustering of the pixel points in the initial target object region can be performed according to the color values into three categories. Specifically, after obtaining the initial target object region, obtain the color values (Red Green Blue, RGB) of each pixel point in the initial target object region, and then perform N-class clustering on the initial target object region according to the color values, so as to obtain the pixel points of N color classifications of the target object. In this embodiment, any existing clustering algorithm can be used to perform clustering processing on the pixel points in the initial target object region, and no limitation is made here.

[0051] Step 140: Obtain N difference images according to the N color classifications and the image to be segmented.

[0052] Among them, the difference map can be a map obtained by subtracting a certain color value from the image to be segmented. Specifically, the color value of each pixel point in the segmentation map is obtained, and then the color value of each pixel point is subtracted from a certain color value to obtain the color value after subtraction of each pixel point, thereby obtaining the difference map. Among them, the subtraction of color values can be understood as the subtraction of the color values of the three RGB channels respectively.

[0053] Optionally, the process of obtaining N difference maps according to N color classifications and the image to be segmented can be: calculating the average value for each of the N color classifications respectively to obtain N color means; calculating the differences between the image to be segmented and the N color means respectively to obtain N difference maps.

[0054] Among them, calculating the average value for each color classification can be understood as calculating the average value for the three RGB channels in each color classification respectively. In this embodiment, after classifying the pixel points in the initial target object area into N categories, the color values of the pixel points included in each category are extracted, and then the average value of the color values is calculated to obtain N color means, and then the image to be segmented is subtracted from the N color means respectively to obtain N difference maps. Exemplarily, Figure 2c is an example diagram of the difference map in this embodiment, as Figure 2c shown. The color of each pixel point in the figure is the value obtained by subtracting the color mean from the color of the pixel point in the original image. In this embodiment, subtracting the image to be segmented from the N color means respectively to obtain N difference maps can improve the speed of obtaining the difference maps.

[0055] Step 150, determining the target mask map according to the N difference maps and the initial mask map.

[0056] Among them, the target mask map can be a mask map optimized from the initial mask map. Specifically, the confidence of each pixel point in the initial mask map can be adjusted according to the N difference maps, thereby obtaining the target mask map.

[0057] Optionally, the process of determining the target mask map according to the N difference maps and the initial mask map can be: adjusting the confidence of the pixel points in the initial mask map whose confidence falls into the first interval to the first set confidence value; for the pixel points in the initial mask map whose confidence falls into the second interval, if the color value of the pixel point in the N difference maps meets the set conditions, then increasing the confidence of the pixel point by a set ratio, otherwise, reducing the confidence of the pixel point by a set ratio; adjusting the confidence of the pixel points in the initial mask map whose confidence falls into the third interval to the second set confidence value.

[0058] Among them, the first interval is greater than the first set value and less than the first set confidence value; the second interval is greater than the second set value and less than the first set value; the second set value is less than the first set value; the third interval is greater than the second set confidence value and less than the second set value. Exemplarily, assume that the first set value is set to 200 / 255, the first set confidence value is 1, the second set value is set to 40 / 255, and the second set confidence value. Then the first interval is [200 / 255, 255 / 255), the second interval is [40 / 255, 200 / 255), and the third interval is [0, 40 / 255). The set condition can be: the average value of the color values of the pixel points in the N difference maps is less than the set threshold; or the minimum value of the color values of the pixel points in the N difference maps is less than the set threshold.

[0059] In this embodiment, the pixel points in the mask image correspond one-to-one with the pixel points in the difference map. The color value of the pixel points in the N difference maps can be understood as the color values of the corresponding pixel points in the N difference maps. The average value of the color values being less than the set threshold can be understood as the average color values of the RGB three channels are all less than the set threshold. Among them, the set threshold can be any value set to 30 - 50 so far, such as 40. Exemplarily, for a certain pixel point, if the color values of the corresponding pixel points of this pixel point in the N difference maps are (R1, G1, B1), (R2, G2, B2), …… (RN, GN, BN) respectively, then the average value of the color values of this pixel point in the N difference maps is ((R1 + R2 + …… + RN) / N, (G1 + G2 + …… + GN) / N, (B1 + B2 + …… + BN) / N). Similarly, the minimum value of the color values of the pixel points in the N difference maps being less than the set threshold can be understood as the minimum value of the color values of the FBG three channels being less than the set threshold.

[0060] Among them, increasing the set ratio can be understood as expanding the confidence by the multiple corresponding to the set ratio, and the required set ratio can be understood as shrinking the confidence by the multiple corresponding to the set ratio. Exemplarily, assume that the set ratio is m and the confidence is A. Then increasing the confidence by the set ratio is expressed as A * m, and shrinking the confidence by the set ratio is expressed as A / m.

[0061] Specifically, for the pixel points whose confidence levels fall within [200 / 255, 255 / 255), directly adjust the confidence level of the pixel points to 255 / 255. For the pixel points in the initial mask image whose confidence levels fall within [40 / 255, 200 / 255), if the average value of the color values of the pixel points in the N difference images is less than the set threshold or the minimum value of the color values of the pixel points in the N difference images is less than the set threshold, then increase the confidence level of the pixel points by a set ratio; otherwise, reduce the confidence level of the pixel points by the set ratio. For the pixel points whose confidence levels fall within [0, 40 / 255), directly adjust the confidence level of the pixel points to 0. Exemplarily, Figure 2d is an example diagram of the target mask image in this embodiment, such as Figure 2d shown, the boundary between the target object and other regions is more obvious. In this embodiment, the confidence levels of the pixel points in the initial mask image are adjusted to 0 or 255 / 255 according to the initial confidence level and the N difference images, making the boundary between the target object and other regions in the mask image more obvious, thereby improving the segmentation accuracy of the target object.

[0062] Optionally, after increasing the confidence level of the pixel points by a set ratio, the following steps are further included: if the increased confidence level exceeds the first set confidence level value, then set the pixel points to the first set confidence level value. The advantage of doing this is to ensure that the pixel points in the mask image are within [0, 255 / 255].

[0063] Step 160, segment the target object in the image to be segmented based on the target mask image.

[0064] Among them, the target mask image characterizes the confidence levels of each pixel point belonging to the target, and the target object can be segmented according to the confidence levels.

[0065] Specifically, the process of segmenting the image to be segmented based on the target mask image can be: determine the pixel points with the confidence level of the first set confidence level value in the target mask image as the second target points; determine the region formed by the pixel points corresponding to the second target points in the image to be segmented as the final target object region.

[0066] Among them, the first set confidence level value is 255 / 255. Specifically, determining the pixel points with the confidence level of the first set confidence level value in the target mask image as the second target points indicates that the probability that the pixel points corresponding to the second target points in the image to be segmented belong to the target object is 255 / 255. Therefore, determine the region formed by the pixel points corresponding to the second target points in the image to be segmented as the final target object region. Exemplarily, Figure 2e is a visualization diagram generated based on the initial mask image (the original image is a color image), Figure 2f is a visualization diagram generated based on the target mask image (the original image is a color image). It can be seen from the figure that Figure 2f compared withFigure 2e Compared with other regions, the boundary between the "sky" and other regions is more obvious. In this embodiment, the region surrounded by the pixel points with a confidence level greater than or equal to the first set confidence value is determined as the final target object region, and the target object can be accurately segmented.

[0067] According to the technical solution of the present disclosure, semantic recognition is performed on the target object in the image to be segmented to obtain an initial mask image; an initial target object region in the image to be segmented is determined based on the initial mask image; pixel points in the initial target object region are clustered according to color values to obtain N color classifications of the target object; N difference images are obtained according to the N color classifications and the image to be segmented; a target mask image is determined according to the N difference images and the initial mask image; and the image to be segmented is segmented based on the target mask image. The object segmentation method provided by the embodiments of the present disclosure determines the target mask image according to the difference image and the initial mask image, and thus segments the target object in the image to be segmented based on the target mask image, which can realize the segmentation of the object in the image, prevent the missed segmentation of the object, and improve the accuracy of object segmentation.

[0068] Figure 3 is a schematic structural diagram of an object segmentation device provided by an embodiment of the present disclosure, as Figure 3 shown, the device includes:

[0069] An initial mask image acquisition module 210, configured to perform semantic recognition on the target object in the image to be segmented to obtain an initial mask image;

[0070] An initial target object region determination module 220, configured to determine an initial target object region in the image to be segmented based on the initial mask image;

[0071] A clustering module 230, configured to cluster pixel points in the initial target object region according to color values to obtain N color classifications of the target object; where N is a positive integer greater than or equal to 1;

[0072] A difference image acquisition module 240, configured to obtain N difference images according to the N color classifications and the image to be segmented;

[0073] A target mask image acquisition module 250, configured to determine a target mask image according to the N difference images and the initial mask image;

[0074] An image segmentation module 260, configured to segment the target object in the image to be segmented based on the target mask image.

[0075] Optionally, the initial mask image acquisition module 210 is further configured to:

[0076] Input the image to be segmented into a target object recognition model and output an initial mask image.

[0077] Optionally, the initial target object area determination module 220 is further configured to:

[0078] Obtain the pixel points in the initial mask image with confidence greater than the first set value, and determine them as the first target points;

[0079] Determine the area formed by the pixel points corresponding to the first target points in the image to be segmented as the initial target object area.

[0080] Optionally, the difference map acquisition module 240 is further configured to:

[0081] Calculate the average value for each of the N color classifications to obtain the N color means;

[0082] Calculate the differences between the image to be segmented and the N color means respectively to obtain N difference maps.

[0083] Optionally, the target mask image acquisition module 250 is further configured to:

[0084] Adjust the confidence of the pixel points in the initial mask image whose confidence falls within the first interval to the first set confidence value; wherein, the first interval is greater than the first set value and less than the first set confidence value;

[0085] For the pixel points in the initial mask image whose confidence falls within the second interval, if the color value of the pixel points in the N difference maps meets the set conditions, increase the confidence of the pixel points by a set ratio, otherwise, reduce the confidence of the pixel points by a set ratio; wherein, the second interval is greater than the second set value and less than the first set value; the second set value is less than the first set value;

[0086] Adjust the confidence of the pixel points in the initial mask image whose confidence falls within the third interval to the second set confidence value; wherein, the third interval is greater than the second set confidence value and less than the second set value.

[0087] Optionally, the target mask image acquisition module 250 is further configured to:

[0088] If the increased confidence exceeds the first set confidence value, set the pixel point to the first set confidence value.

[0089] Optionally, the image segmentation module 260 is further configured to:

[0090] Determine the pixel points in the target mask image with confidence as the first set confidence value as the second target points; determine the area formed by the pixel points corresponding to the second target points in the image to be segmented as the final target object area.

[0091] The above-mentioned device can execute the methods provided in all the foregoing embodiments of the present disclosure, and has corresponding functional modules and beneficial effects for executing the above methods. For technical details not described in detail in this embodiment, reference may be made to the methods provided in all the foregoing embodiments of the present disclosure.

[0092] Reference is now made to Figure 4 , which shows a schematic structural diagram of an electronic device 300 suitable for implementing the embodiments of the present disclosure. The electronic device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc., or various forms of servers, such as independent servers or server clusters. Figure 4 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0093] As Figure 4 shown, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only storage device (ROM) 302 or a program loaded from a storage device 308 into a random access storage device (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0094] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 can allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included.

[0095] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the word recommendation method. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-described functions defined in the method of the embodiments of the present disclosure are performed.

[0096] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0097] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed network.

[0098] The above computer-readable medium can be included in the above electronic device; or can exist separately without being assembled into the electronic device.

[0099] The above computer-readable medium carries one or more programs, which when executed by the electronic device, cause the electronic device to: perform semantic recognition on the target object in the image to be segmented to obtain an initial mask image; determine an initial target object region in the image to be segmented based on the initial mask image; perform clustering processing on the pixel points in the initial target object region according to the color values to obtain N color classifications of the target object; where N is a positive integer greater than or equal to 1; obtain N difference images according to the N color classifications and the image to be segmented; determine a target mask image according to the N difference images and the initial mask image; segment the target object in the image to be segmented based on the target mask image.

[0100] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0101] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0102] The units involved in the embodiments of the present disclosure can be implemented in software or in hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.

[0103] The functions described above herein can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), Systems on Chip (SOC), Complex Programmable Logic Devices (CPLD), and the like.

[0104] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or Flash memory), an optical fiber, a portable Compact Disc Read-Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0105] According to one or more embodiments of the embodiments of the present disclosure, the embodiments of the present disclosure disclose an object segmentation method, including:

[0106] Perform semantic recognition on the target object in the image to be segmented to obtain an initial mask image;

[0107] Based on the initial mask image, determine the initial target object region in the image to be segmented;

[0108] Cluster the pixel points in the initial target object region according to the color values to obtain N color classifications of the target object; where N is a positive integer greater than or equal to 1;

[0109] Obtain N difference maps according to the N color classifications and the image to be segmented;

[0110] Determine the target mask image according to the N difference maps and the initial mask image;

[0111] Segment the target object in the image to be segmented based on the target mask image.

[0112] Further, performing semantic recognition on the target object in the image to be segmented to obtain an initial mask image includes:

[0113] Input the image to be segmented into the target object recognition model and output the initial mask image.

[0114] Further, based on the initial mask image, determining the initial target object region in the image to be segmented includes:

[0115] Obtain the pixel points in the initial mask image with a confidence greater than the first set value and determine them as the first target points;

[0116] Determine the region formed by the pixel points corresponding to the first target points in the image to be segmented as the initial target object region.

[0117] Further, obtaining N difference maps according to the N color classifications and the image to be segmented includes:

[0118] Calculate the average value for each of the N color classifications to obtain N color means;

[0119] Calculate the differences between the image to be segmented and the N color means respectively to obtain N difference maps.

[0120] Further, determining the target mask image according to the N difference maps and the initial mask image includes:

[0121] Adjust the confidence of the pixel points in the initial mask image whose confidence falls within the first interval to the first set confidence value; where the first interval is greater than the first set value and less than the first set confidence value;

[0122] For the pixel points in the initial mask image whose confidence levels fall within the second interval, if the color values of the pixel points in the N difference images meet the set conditions, increase the confidence levels of the pixel points by a set proportion; otherwise, reduce the confidence levels of the pixel points by the set proportion; wherein, the second interval is greater than a second set value and less than the first set value; the second set value is less than the first set value.

[0123] Adjust the confidence levels of the pixel points in the initial mask image whose confidence levels fall within the third interval to the second set confidence level value; wherein, the third interval is greater than the second set confidence level value and less than the second set value.

[0124] Further, after increasing the confidence levels of the pixel points by the set proportion, it further includes:

[0125] If the increased confidence level exceeds the first set confidence level value, set the pixel point to the first set confidence level value.

[0126] Further, segmenting the image to be segmented based on the target mask image includes:

[0127] Determine the pixel points with the first set confidence level value in the target mask image as the second target points;

[0128] Determine the region formed by the pixel points corresponding to the second target points in the image to be segmented as the final target object region.

[0129] Note that the above is only the preferred embodiment of the present disclosure and the applied technical principles. Those skilled in the art will understand that the present disclosure is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present disclosure. Therefore, although the present disclosure has been described in detail through the above embodiments, the present disclosure is not limited to the above embodiments. Without departing from the concept of the present disclosure, more other equivalent embodiments can be included, and the scope of the present disclosure is determined by the scope of the appended claims.

Claims

1. An object segmentation method, characterized in that Including: Performing semantic recognition on the target object in the image to be segmented to obtain an initial mask image; Determining an initial target object region in the image to be segmented based on the initial mask image; Performing clustering processing on the pixel points in the initial target object region according to color values to obtain N color classifications of the target object; where N is a positive integer greater than or equal to 1; Obtaining N difference maps according to the N color classifications and the image to be segmented; Determining a target mask image according to the N difference maps and the initial mask image; Segmenting the target object in the image to be segmented based on the target mask image; Obtaining N difference maps according to the N color classifications and the image to be segmented, including: Calculating the average value for each of the N color classifications respectively to obtain N color means; Calculating the differences between the image to be segmented and the N color means respectively to obtain N difference maps; Determining a target mask image according to the N difference maps and the initial mask image, including: Adjusting the confidence of the pixel points in the initial mask image whose confidence falls within the first interval to a first set confidence value; where the first interval is greater than a first set value and less than the first set confidence value; For the pixel points in the initial mask image whose confidence falls within the second interval, if the color value of the pixel point in the N difference maps meets the set condition, then increasing the confidence of the pixel point by a set proportion, otherwise, reducing the confidence of the pixel point by the set proportion; where the second interval is greater than a second set value and less than the first set value; the second set value is less than the first set value; Adjusting the confidence of the pixel points in the initial mask image whose confidence falls within the third interval to a second set confidence value; where the third interval is greater than the second set confidence value and less than the second set value.

2. The method according to claim 1, wherein Performing semantic recognition on the target object in the image to be segmented to obtain an initial mask image, including: Inputting the image to be segmented into a target object recognition model and outputting an initial mask image.

3. The method according to claim 1, wherein Determining an initial target object region in the image to be segmented based on the initial mask image, including: Obtaining the pixel points in the initial mask image whose confidence is greater than a first set value and determining them as first target points; Determining the region formed by the pixel points corresponding to the first target points in the image to be segmented as the initial target object region.

4. The method according to claim 1, characterized in that, After increasing the confidence of the pixel point by the set proportion, further including: If the increased confidence exceeds the first set confidence value, then setting the pixel point to the first set confidence value.

5. The method according to claim 1, wherein Segmenting the image to be segmented based on the target mask image, including: Determining the pixel points in the target mask image whose confidence is the first set confidence value as second target points; Determining the region formed by the pixel points corresponding to the second target points in the image to be segmented as the final target object region.

6. An object segmentation device, characterized in that, Including: An initial mask image acquisition module, configured to perform semantic recognition on the target object in the image to be segmented to obtain an initial mask image; An initial target object region determination module, configured to determine an initial target object region in the image to be segmented based on the initial mask image; A clustering module, configured to perform clustering processing on pixel points in the initial target object region according to color values to obtain N color classifications of the target object; where N is a positive integer greater than or equal to 1; A difference map acquisition module, configured to obtain N difference maps according to the N color classifications and the image to be segmented; A target mask image acquisition module, configured to determine a target mask image according to the N difference maps and the initial mask image; An image segmentation module, configured to segment the target object in the image to be segmented based on the target mask image; The difference map acquisition module is further configured to: Calculate the average value for each of the N color classifications to obtain N color means; Calculate the differences between the image to be segmented and the N color means respectively to obtain N difference maps; The target mask image acquisition module is further configured to: Adjust the confidence level of pixel points in the initial mask image whose confidence level falls within a first interval to a first set confidence level value; where the first interval is greater than a first set value and less than the first set confidence level value; For pixel points in the initial mask image whose confidence level falls within a second interval, if the color value of the pixel point in the N difference maps meets a set condition, increase the confidence level of the pixel point by a set proportion, otherwise, reduce the confidence level of the pixel point by the set proportion; where the second interval is greater than a second set value and less than the first set value; the second set value is less than the first set value; Adjust the confidence level of pixel points in the initial mask image whose confidence level falls within a third interval to a second set confidence level value; where the third interval is greater than the second set confidence level value and less than the second set value.

7. An electronic device, characterized in that, The electronic device includes: One or more processing devices; A storage device, configured to store one or more programs; When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement the object segmentation method according to any one of claims 1-5.

8. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processing device, it implements the object segmentation method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Method and device for detecting target

    CN108229575A

  • Image instance segmentation method, device, apparatus, and storage medium

    CN109242869A