Image segmentation method, device, equipment, computer readable medium and program product
By obtaining the segmentation image feature information set and the overall image feature information, and using the target object segmentation model to generate overall and local mask images, the problem of inaccurate segmentation caused by resolution differences in the image segmentation model is solved, and accurate image segmentation is achieved.
Patent Information
- Application Number
- CN202510765883.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-12
AI Technical Summary
The existing image segmentation model has a large difference between the resolution of the output mask image and the resolution of the target image due to the large downsampling operation, resulting in inaccurate image segmentation results.
By obtaining the segmentation image feature information set and the overall image feature information corresponding to the target image, where the resolution of the segmentation image feature information is the same as the resolution of the mask image output by the target object segmentation model, the target object segmentation model is used to generate overall and local mask images, and a multi-level segmentation method is adopted to ensure that the resolution of the segmentation image is consistent with the resolution of the output mask image.
It achieves the goal of accurately generating the segmentation results of the target image through a multi-level segmentation method while ensuring the consistency of the resolution of the mask image and the input image, thus avoiding the reduction of segmentation accuracy due to downsampling.
Smart Images

Figure CN120635103A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of image segmentation, and in particular to an image segmentation method, apparatus, device, computer-readable medium, and program product. Background Art
[0002] With the continuous development of artificial intelligence, image segmentation technology is now widely used in our daily lives. The typical approach for segmenting a target image is to first directly input the target image into an image segmentation model to generate a mask image. Then, based on the mask image, the target image is segmented.
[0003] However, the inventors have discovered that when using the above method to segment an image, the following technical problems often occur:
[0004] Due to the large-scale downsampling operation in the image segmentation model, the resolution of the output mask image is significantly different from that of the target image, so that one segmentation pixel in the mask image corresponds to multiple pixels in the target image, making the obtained image segmentation result inaccurate.
[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the Invention
[0006] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0007] Some embodiments of the present disclosure provide image segmentation methods, devices, equipment, computer-readable media, and program products to solve the technical problems mentioned in the above background technology section.
[0008] In a first aspect, some embodiments of the present disclosure provide an image segmentation method, including: obtaining a divided image feature information set and overall image feature information corresponding to a target image, wherein a first resolution of the divided image corresponding to the divided image feature information is the same as a second resolution of a mask image output by a target object segmentation model; generating an overall mask image based on the overall image feature information using the target object segmentation model; performing image segmentation on the overall mask image according to an image segmentation method corresponding to the divided image set to obtain a divided mask image set; in response to determining that the target object is to be segmented based on the overall image content, for each divided image feature information in the divided image feature information set, generating a local mask image corresponding to the divided image feature information based on the divided image feature information and the divided mask image having a position correspondence, using the target object segmentation model; generating object segmentation information corresponding to the target image based on the obtained local mask image set.
[0009] Optionally, the target object segmentation model is an interactive segmentation model; and the above-mentioned target object segmentation model is used to generate an overall mask image based on the above-mentioned overall image feature information, including: generating first image prompt information corresponding to the above-mentioned target image; inputting the above-mentioned overall image feature information and the above-mentioned first image prompt information into the above-mentioned interactive segmentation model to generate the above-mentioned overall mask image.
[0010] Optionally, the above-mentioned target object segmentation model is used to generate a local mask image corresponding to the above-mentioned segmentation image feature information based on the above-mentioned segmentation image feature information and the segmentation mask image with a positional correspondence, including: generating second image prompt information corresponding to the above-mentioned segmentation mask image; inputting the above-mentioned segmentation mask image into a vector conversion model to generate mask image feature information; inputting the above-mentioned mask image feature information, the above-mentioned segmentation image feature information and the above-mentioned second image prompt information into the above-mentioned target object segmentation model to generate the above-mentioned local mask image.
[0011] Optionally, the above method also includes: in response to determining to segment the target object based on local image content, obtaining the bounding box size ratio and the bounding box circled on the target image; selecting a divided image subset from the divided image set based on the bounding box size ratio and the bounding box, wherein the divided image subset is each adjacent image that meets the bounding box size ratio between the total resolution of the corresponding image and the bounding box size; and using the target object segmentation model to generate object segmentation information corresponding to the divided image subset based on the divided image feature information subset corresponding to the divided image subset.
[0012] Optionally, the above-mentioned divided image feature information set and overall image feature information are generated by the following steps: determining the resolution of the mask image corresponding to the output of the above-mentioned target object segmentation model as the second resolution; dividing the above-mentioned target image according to the first resolution corresponding to the above-mentioned target image and the above-mentioned second resolution to obtain a divided image set, wherein the resolution corresponding to each divided image is the same as the above-mentioned second resolution; inputting each divided image in the above-mentioned divided image set into the vector transformation model to generate divided image feature information to obtain the divided image feature information set; inputting the above-mentioned target image into the above-mentioned vector transformation model to generate the above-mentioned overall image feature information.
[0013] Optionally, the above-mentioned mask image feature information, the above-mentioned segmentation image feature information and the above-mentioned second image prompt information are input into the above-mentioned target object segmentation model to generate the above-mentioned local mask image, including: determining the segmentation image corresponding to the above-mentioned segmentation image feature information as the target segmentation image; screening out at least one segmentation image that has an image proximity relationship with the above-mentioned target segmentation image from the above-mentioned segmentation image set as at least one adjacent image; combining the above-mentioned at least one adjacent image and the above-mentioned target segmentation image to generate a combined image; generating image prompt information and a subset of segmentation image feature information corresponding to the above-mentioned combined image as third image prompt information and a subset of target segmentation image feature information, respectively; inputting the above-mentioned target segmentation image feature information subset and the above-mentioned third image prompt information into the above-mentioned target object segmentation model to generate a combined mask image; screening out a sub-mask image corresponding to the above-mentioned target segmentation image from the above-mentioned combined mask image; inputting the image feature information corresponding to the above-mentioned sub-mask image, the above-mentioned mask image feature information, the above-mentioned segmentation image feature information and the above-mentioned second image prompt information into the above-mentioned target object segmentation model to generate the above-mentioned local mask image.
[0014] Optionally, the above-mentioned target object segmentation model is used to generate the local object segmentation information corresponding to the above-mentioned divided image subset based on the divided image feature information subset corresponding to the above-mentioned divided image subset, including: determining the division mask image corresponding to each division image in the above-mentioned divided image subset to obtain the division mask image subset; based on the above-mentioned division mask image subset and the above-mentioned division image feature information subset, the above-mentioned target object segmentation model is used to generate the above-mentioned local object segmentation information.
[0015] Optionally, after performing image division on the overall mask image according to the image division method corresponding to the division image set to obtain the division mask image set, the method further includes: screening out target division mask images whose corresponding mask area ratio is higher than the target ratio from the division mask image set to obtain at least one target division mask image; determining the at least one target division mask image as the division mask image set, and determining at least one division image feature information corresponding to the at least one target division mask image as the division image feature information set.
[0016] In the second aspect, some embodiments of the present disclosure provide an image segmentation device, comprising: an acquisition unit, configured to acquire a divided image feature information set and overall image feature information corresponding to a target image, wherein the first resolution of the divided image corresponding to the divided image feature information is the same as the second resolution of the mask image output by the target object segmentation model; a first generation unit, configured to generate an overall mask image based on the overall image feature information and using the target object segmentation model; an image division unit, configured to perform image division on the overall mask image according to the image division method corresponding to the divided image set, to obtain a divided mask image set; a second generation unit, configured to, in response to determining that the target object is to be segmented based on the overall image content, generate, for each divided image feature information in the divided image feature information set, a local mask image corresponding to the divided image feature information based on the divided image feature information and the divided mask image having a position correspondence, using the target object segmentation model; a third generation unit, configured to generate object segmentation information corresponding to the target image based on the obtained local mask image set.
[0017] Optionally, the target object segmentation model is an interactive segmentation model; and the first generation unit can be configured to: generate first image prompt information corresponding to the target image; input the overall image feature information and the first image prompt information into the interactive segmentation model to generate the overall mask image.
[0018] Optionally, the second generation unit can be configured to: generate second image prompt information corresponding to the above-mentioned divided mask image; input the above-mentioned divided mask image into a vector transformation model to generate mask image feature information; input the above-mentioned mask image feature information, the above-mentioned divided image feature information and the above-mentioned second image prompt information into the above-mentioned target object segmentation model to generate the above-mentioned local mask image.
[0019] Optionally, the device also includes: in response to determining to segment the target object based on local image content, obtaining the bounding box size ratio and the above-mentioned bounding box size; selecting a divided image subset group from the above-mentioned divided image set based on the above-mentioned bounding box size ratio and the above-mentioned bounding box size, wherein the divided image subsets are each adjacent image that meets the above-mentioned bounding box size ratio between the total resolution of the corresponding image and the above-mentioned bounding box size; according to predetermined divided image sampling constraint information, screening out at least one divided image subset from the above-mentioned divided image subset group; for each divided image subset in the above-mentioned at least one divided image subset, based on the divided image feature information subset corresponding to the above-mentioned divided image subset, using the above-mentioned target object segmentation model, generate local object segmentation information corresponding to the above-mentioned divided image subset.
[0020] Optionally, the second generation unit can be configured to: determine the segmented image corresponding to the above-mentioned segmented image feature information as the target segmented image; filter out at least one segmented image that has an image proximity relationship with the above-mentioned target segmented image from the above-mentioned segmented image set as at least one adjacent image; combine the above-mentioned at least one adjacent image and the above-mentioned target segmented image to generate a combined image; generate image prompt information and a subset of segmented image feature information corresponding to the above-mentioned combined image as the third image prompt information and the target segmented image feature information subset, respectively; input the above-mentioned target segmented image feature information subset and the above-mentioned third image prompt information into the above-mentioned target object segmentation model to generate a combined mask image; filter out the sub-mask image corresponding to the above-mentioned target segmented image from the above-mentioned combined mask image; input the image feature information corresponding to the above-mentioned sub-mask image, the above-mentioned mask image feature information, the above-mentioned segmented image feature information and the above-mentioned second image prompt information into the above-mentioned target object segmentation model to generate the above-mentioned local mask image.
[0021] Optionally, the device also includes: determining the segmentation mask image corresponding to each segmentation image in the above-mentioned segmentation image subset to obtain the segmentation mask image subset; based on the above-mentioned segmentation mask image subset and the above-mentioned segmentation image feature information subset, using the above-mentioned target object segmentation model, generating the above-mentioned local object segmentation information.
[0022] Optionally, the device also includes: filtering out target partitioning mask images whose corresponding mask area ratio is higher than the target ratio from the partitioning mask image set to obtain at least one target partitioning mask image; determining the at least one target partitioning mask image as a partitioning mask image set, and determining at least one partitioning image feature information corresponding to the at least one target partitioning mask image as a partitioning image feature information set.
[0023] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.
[0024] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner in the first aspect is implemented.
[0025] In a fifth aspect, some embodiments of the present disclosure provide a computer program product, including a computer program, which implements the method described in any implementation manner in the first aspect when executed by a processor.
[0026] The above-described embodiments of the present disclosure have the following advantageous effects: Through the image segmentation methods of some embodiments of the present disclosure, accurate image segmentation of a target image is achieved through a multi-level segmentation approach, while ensuring the same resolution between the aggregated mask image and the input image. Specifically, the reason for the inaccurate segmentation result of the target image is that due to the significant downsampling operation of the image segmentation model, the resolution of the output mask image is significantly different from that of the target image, resulting in a single segmentation pixel in the mask image corresponding to multiple pixels in the target image, resulting in an inaccurate image segmentation result. Based on this, the image segmentation methods of some embodiments of the present disclosure first obtain a set of segmented image feature information and overall image feature information corresponding to the target image, wherein the first resolution of the segmented image corresponding to the segmented image is the same as the second resolution of the mask image output by the target object segmentation model. Here, by pre-segmenting the target image, the resolution of the segmented image is ensured to be the same as the output resolution of the target object segmentation model. Therefore, when the target object segmentation model receives the segmented image as input, the output result does not suffer from the significant resolution difference. By determining the segmentation result corresponding to each segmented image from the perspective of the segmented image, and then comprehensively considering each segmentation result, the precise segmentation result corresponding to the target image can be accurately obtained. By obtaining the segmented image feature information set, the image feature semantic content corresponding to each segmented image is determined, so that the subsequent target object segmentation model can obtain the image semantic content corresponding to the segmented image to accurately generate a local mask image. By obtaining the overall image feature information, the subsequent target object segmentation model can obtain the image semantic content corresponding to the target image to accurately generate an overall mask image. Then, based on the above-mentioned overall image feature information, using the above-mentioned target object segmentation model, an overall mask image can be accurately generated to provide preliminary image segmentation results for each subsequent segmented image, thereby improving the accuracy of the generation of local mask images. Next, according to the image segmentation method corresponding to the segmented image set, the above-mentioned overall mask image is segmented to obtain a segmented mask image set to provide corresponding preliminary image segmentation results for each subsequent mask image. Furthermore, in response to determining to segment the target object based on the overall image content, for each segmented image feature information in the segmented image feature information set, based on the segmented image feature information and the segmented mask image that has a positional correspondence, the target object segmentation model can be used to accurately generate a local mask image corresponding to the segmented image feature information. Here, the resolution corresponding to the local mask image is the same as the resolution corresponding to the segmented image, so the pixels in the subsequently generated local mask image have a one-to-one correspondence with the pixels in the segmented image. Therefore, in the process of generating the local mask image, there will be no reduction in segmentation accuracy due to downsampling. In this way, the local mask image can be accurately generated.Finally, based on the obtained local mask image set, the object segmentation information corresponding to the target image can be accurately generated. In summary, by dividing the target image, the resolution of the obtained divided images is equal to the resolution of the output mask image, so that there is no pixel segmentation deviation in determining the event of each divided image corresponding to the local mask image. Therefore, while ensuring that the resolution of the aggregated mask image and the input image is the same, accurate image segmentation of the target image can be achieved through a multi-level segmentation method. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0028] Figure 1 is a schematic diagram of an application scenario of the image segmentation method according to some embodiments of the present disclosure;
[0029] Figure 2 is a flowchart of some embodiments of the image segmentation method according to the present disclosure;
[0030] Figure 3-4 is a schematic diagram of dividing an image set in some embodiments of the image segmentation method according to the present disclosure;
[0031] Figure 5 is an algorithm flow chart of an image segmentation method according to some embodiments of the image segmentation method of the present disclosure;
[0032] Figure 6 is a schematic diagram of generating overall image feature information in some embodiments of the image segmentation method disclosed herein;
[0033] Figure 7 is a schematic diagram of overall image feature information and segmented image feature information in some embodiments of the image segmentation method disclosed herein;
[0034] Figure 8 is a schematic diagram of generating an overall mask image in some embodiments of the image segmentation method according to the present disclosure;
[0035] Figure 9 is a schematic diagram of a target segmentation image and at least one adjacent image in some embodiments of the image segmentation method according to the present disclosure;
[0036] Figure 10 is a flowchart of other embodiments of the image segmentation method according to the present disclosure;
[0037] Figure 11 is an overall flow chart of segmenting a target object based on overall image content in some embodiments of the image segmentation method disclosed herein;
[0038] Figure 12 is a schematic structural diagram of some embodiments of the image segmentation device according to the present disclosure;
[0039] Figure 13 It is a structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0040] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0041] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.
[0042] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0043] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0044] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0045] Before performing operations such as the collection, storage, and use of the information involved in this disclosure (such as the division of image feature information sets and overall image feature information), relevant organizations or individuals must fulfill their obligations, including conducting information security impact assessments, fulfilling their obligation to inform information subjects, and obtaining prior authorization and consent from information subjects.
[0046] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0047] Figure 1It is a schematic diagram of an application scenario of the image segmentation method according to some embodiments of the present disclosure.
[0048] exist Figure 1 In the application scenario, first, the electronic device 101 can obtain the segmented image feature information set 104 and the overall image feature information 105 corresponding to the target image 101. The first resolution of the segmented image corresponding to the segmented image feature information is the same as the second resolution of the mask image output by the target object segmentation model 106. In this application scenario, the segmented image feature information set 104 includes: face segmented image feature information 1041 corresponding to the segmented image 1031, face segmented image feature information 1042 corresponding to the segmented image 1032, face segmented image feature information 1043 corresponding to the segmented image 1033, and face segmented image feature information 1044 corresponding to the segmented image 1034. Then, the electronic device 101 can generate an overall mask image 107 based on the above-mentioned overall image feature information 105 using the above-mentioned target object segmentation model 106. Next, the electronic device 101 can perform image segmentation on the above-mentioned overall mask image 107 according to the image segmentation method corresponding to the segmented image set 103 to obtain a segmented mask image set 108. Furthermore, in response to determining to segment the target object based on the overall image content, for each segmented image feature information in the segmented image feature information set 104, the electronic device 101 may generate a local mask image corresponding to the segmented image feature information based on the segmented image feature information and the segmented mask image with a corresponding position relationship, using the target object segmentation model 106. In this application scenario, based on the facial segmented image feature information 1041 and the segmented mask image 1031 with a corresponding position relationship, the target object segmentation model 106 is used to generate a local mask image corresponding to the facial segmented image feature information 1041. Based on the facial segmented image feature information 1042 and the segmented mask image 1032 with a corresponding position relationship, the target object segmentation model 106 is used to generate a local mask image corresponding to the facial segmented image feature information 1042. Based on the facial segmented image feature information 1043 and the segmented mask image 1033 with a corresponding position relationship, the target object segmentation model 106 is used to generate a local mask image corresponding to the facial segmented image feature information 1043. Based on the facial segmentation image feature information 1044 and the positionally corresponding image 1034, the target object segmentation model 106 is used to generate a local mask image corresponding to the facial segmentation image feature information 1044. Finally, the electronic device 101 can generate object segmentation information 110 corresponding to the target image 102 based on the obtained local mask image set 109.
[0049] It should be noted that the electronic device 101 can be hardware or software. When the electronic device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or it can be implemented as a single server or a single terminal device. When the electronic device is embodied as software, it can be installed in the hardware devices listed above. It can be implemented as multiple software or software modules for providing distributed services, for example, or it can be implemented as a single software or software module. No specific limitation is made here.
[0050] It should be understood that Figure 1 The number of electronic devices in the embodiment is merely illustrative. Any number of electronic devices may be provided according to implementation requirements.
[0051] Continue to refer Figure 2 , shows a process 200 of some embodiments of the image segmentation method according to the present disclosure. The image segmentation method comprises the following steps:
[0052] Step 201: Obtain the segmented image feature information set and the overall image feature information corresponding to the target image.
[0053] In some embodiments, the execution subject of the above image segmentation method (for example Figure 1 The electronic device 101 shown) can obtain the segmented image feature information set and the overall image feature information corresponding to the target image through a wired connection or a wireless connection. The target image can be an image to be segmented. For example, the target image can be a face image, or a main view image corresponding to the target object. The segmented image feature information can represent the image semantic content corresponding to the segmented image. That is, the segmented image feature information can be the semantic content corresponding to a partial image area in the target image. The segmented image feature information can be feature information in the form of a vector. There is a one-to-one correspondence between the segmented image feature information in the segmented image feature information set and the segmented images in the segmented image set. The overall image feature information can be a vector representing the overall image feature semantic content corresponding to the target image. The segmented image feature information set and the overall image feature information are pre-generated offline. The segmented image feature information set, the overall image feature information and the target image are stored in the image feature information database for subsequent online segmentation.
[0054] In some optional implementations of some embodiments, the above-mentioned divided image feature information sets and overall image feature information are generated by the following steps:
[0055] The first step is to determine the resolution of the mask image output by the target object segmentation model as the second resolution. If the input and output of the target object segmentation model and the model structure are predetermined, the corresponding second resolution is also predetermined.
[0056] In the second step, the target image is divided according to the first resolution and the second resolution corresponding to the target image to obtain a set of divided images, wherein the resolution corresponding to each divided image is the same as the second resolution.
[0057] As an example, the execution entity may determine whether the number of pixels in the horizontal and vertical directions of the screen at the first resolution is an integer multiple of the number of pixels in the horizontal and vertical directions of the screen at the second resolution. Then, in response to the determination that the number of pixels is equal, the image corresponding to the first resolution is evenly divided using the area corresponding to the second resolution as a division unit to obtain a set of divided images. Next, in response to the determination that the number of pixels is equal, the target image at the first resolution is filled in the corresponding direction so that the resolution of the filled image is an integer multiple of the resolution of the second resolution in each direction. Finally, the filled image is evenly divided using the area corresponding to the second resolution as a division unit to obtain a set of divided images.
[0058] like Figure 3-Figure 4 As shown, a schematic diagram of dividing an image set is shown.
[0059] like Figure 3 As shown, for a target image, the first resolution is 512*512 and the second resolution is 256*256. The target image can be divided into four blocks by gridding, and the resolution of each block corresponding to the grid is 256*256.
[0060] like Figure 4 As shown in the figure, for a target image with a first resolution of 1920*1080 and a second resolution of 256*256, 1920 and 1080 cannot divide 256 evenly. Therefore, padding (pixel addition) is performed on the target image to obtain a padded image of 2048*1280. The padded image of 2048*1280 is then evenly divided to obtain a set of divided images, each of which is 256*256.
[0061] In the third step, each segmented image in the segmented image set is input into a vector conversion model to generate segmented image feature information, thereby obtaining a segmented image feature information set.
[0062] The fourth step is to input the target image into the vector conversion model to generate the overall image feature information.
[0063] It should be noted that the segmentation image feature information set and the overall image feature information are generated offline. During the online target object segmentation process of the target image, the segmentation image feature information set and the overall image feature information generated offline can be directly obtained, thereby greatly improving the image segmentation efficiency.
[0064] like Figure 5 , which shows an algorithm flow chart of the image segmentation method.
[0065] In the offline feature extraction process, the target image is first segmented to obtain a set of segmented images. Feature extraction (i.e., top-level image feature extraction) is then performed on the target image to obtain overall image feature information. Next, feature extraction (i.e., grid image feature extraction) is performed on each segmented image to obtain segmented image feature information.
[0066] In the online segmentation reasoning phase, annotation points (i.e., positive and negative points) are input to perform segmentation reasoning on the top-level image, generating a global mask image corresponding to the target image. Based on this global mask image, segmentation reasoning is performed on the grid image (i.e., the partitioned image) to obtain the final segmentation result corresponding to the target image.
[0067] like Figure 6 As shown, a schematic diagram of generating overall image feature information is shown.
[0068] like Figure 6 As shown in FIG, the input image (ie, the target image) is input into the Embedding extraction model to obtain an n*n*d Embedding vector (ie, the overall image feature information).
[0069] like Figure 7 As shown, a schematic diagram of overall image feature information and divided image feature information is shown.
[0070] The overall image feature information is an n*n*d Embedding vector. Figure 3 There are 4 partitioned image sets in . The partitioned image feature information corresponding to Grid(1,1) is an n*n*d vector. The partitioned image feature information corresponding to Grid(1,2) is an n*n*d vector. The partitioned image feature information corresponding to Grid(2,1) is an n*n*d vector. The partitioned image feature information corresponding to Grid(2,2) is an n*n*d vector.
[0071] Step 202 : Generate an overall mask image based on the overall image feature information and using the target object segmentation model.
[0072] In some embodiments, the execution entity may generate an overall mask image using the target object segmentation model based on the overall image feature information. The overall mask image may be an image that represents the mask position information of the target object in the target image. The image region where the target object is located in the overall mask image is set to a first color, and the image region where the non-target object is located is set to a second color. For example, the first color may be black, and the second color may be white. The target object segmentation model may be a neural network model that segments the target object in the image. For example, the target object segmentation model may be an encoding and decoding model. For example, the target object segmentation model may be a U-net model. The resolution corresponding to the overall mask image is the same as the resolution corresponding to the segmented image. The target object may be a segmented object.
[0073] As an example, in response to determining that the target object segmentation model is a non-interactive segmentation model, the execution entity may directly input the overall image feature information into the target object segmentation model to obtain an overall mask image.
[0074] In some optional implementations of some embodiments, the target object segmentation model is an interactive segmentation model. The first interactive segmentation model may be a neural network model that performs object segmentation based on interactive information. For example, the first interactive segmentation model may be a SAM (Segment Anything Model) model.
[0075] Optionally, the execution entity may generate an overall mask image based on the overall image feature information and using the target object segmentation model, including the following steps:
[0076] The first step is to generate the first image prompt information corresponding to the target image. The first image prompt information can be a positive example set and a negative example set of the corresponding image content in the target image. The positive example set and the negative example set can be pre-set on the target image, and the corresponding numbers can be calibrated in sequence by observing the segmentation effect. In practice, the positive example can be that the corresponding point (or mask area) is on the target object in the target image. The negative example can be that the corresponding point (or mask area) is not on the target object in the target image. For example, the positive example can be a green point. The negative example can be a red point.
[0077] In the second step, the overall image feature information and the first image prompt information are input into the interactive segmentation model to generate the overall mask image.
[0078] like Figure 8 As shown, a schematic diagram of generating the overall mask image is shown.
[0079] The overall image feature information and the first image prompt information are input into the interactive segmentation model to generate the overall mask image. The interactive segmentation model includes an encoder model and a decoder model.
[0080] Step 203 : performing image division on the overall mask image according to the image division method corresponding to the division image set to obtain a division mask image set.
[0081] In some embodiments, the execution entity may perform image segmentation on the overall mask image according to an image segmentation method corresponding to the segmentation image set to obtain a segmentation mask image set. The image segmentation method may be a method for segmenting the target image into the segmentation image set. The segmentation mask images in the segmentation mask image set have a one-to-one positional correspondence with the segmentation images in the segmentation image set. The segmentation mask images may be images that represent mask position information of the target object in the segmentation image.
[0082] In some optional implementations of some embodiments, after step 203, the execution entity may further include the following steps:
[0083] The first step is to select, from the set of segmentation mask images, target segmentation mask images whose corresponding mask area ratio is higher than the target ratio, thereby obtaining at least one target segmentation mask image. The mask area ratio may be the resolution ratio between the area corresponding to the mask content and the area corresponding to the segmentation mask image. In practice, the target ratio may be a preset ratio value. For example, the target ratio may be 5%.
[0084] In the second step, the at least one target segmentation mask image is determined as a segmentation mask image set, and at least one segmentation image feature information corresponding to the at least one target segmentation mask image is determined as a segmentation image feature information set.
[0085] Step 204, in response to determining to segment the target object based on the overall image content, for each segmentation image feature information in the above-mentioned segmentation image feature information set, based on the above-mentioned segmentation image feature information and the segmentation mask image with a positional correspondence, the above-mentioned target object segmentation model is used to generate a local mask image corresponding to the above-mentioned segmentation image feature information.
[0086] In some embodiments, in response to determining to segment the target object based on the overall image content, for each segmented image feature information in the segmented image feature information set, the execution entity may generate a local mask image corresponding to the segmented image feature information based on the segmented image feature information and the segmented mask image with a positional correspondence, using the target object segmentation model. Wherein, if there is a segmented image with a corresponding positional relationship for the segmented image feature information, there is also a segmented mask image with a positional correspondence. Wherein, the local mask image may be an image of the mask position information of the target object in the segmented image corresponding to the segmented image feature information. The resolution corresponding to the local mask image is equivalent to that of the corresponding segmented image. Segmenting the target object based on the overall image content may be a global segmentation of the image content corresponding to the target image. That is, all object content in the target image that is related to the target object is segmented out.
[0087] As an example, in response to determining that the target object segmentation model is a non-interactive segmentation model, the above-mentioned execution entity can directly input the above-mentioned segmentation image feature information and the segmentation mask image with a position correspondence into the target object segmentation model to generate a local mask image.
[0088] In some optional implementations of some embodiments, the execution entity may generate a local mask image corresponding to the segmented image feature information using the target object segmentation model based on the segmented image feature information and the segmented mask image having a positional correspondence, including the following steps:
[0089] The first step is to generate second image prompt information corresponding to the segmentation mask image. The explanation of the second image prompt information can refer to the explanation of the first image prompt information. The difference between the second image prompt information and the first image prompt information is that they target different image objects.
[0090] The second step is to input the segmented mask image into a vector transformation model to generate mask image feature information. The vector transformation model can be a neural network model that transforms the image into a vector form and represents the semantic content of the image features. For example, the vector transformation model can be an embedding model.
[0091] In the third step, the mask image feature information, the segmented image feature information and the second image prompt information are input into the target object segmentation model to generate the local mask image.
[0092] In some optional implementations of some embodiments, the execution entity may input the mask image feature information, the segmented image feature information, and the second image prompt information into the target object segmentation model to generate the local mask image, including the following steps:
[0093] The first step is to determine the segmented image corresponding to the above segmented image feature information as the target segmented image.
[0094] As an example, the execution subject determines the segmented image corresponding to the segmented image feature information as the target segmented image through the position correspondence relationship.
[0095] In the second step, at least one segmented image that is in proximity to the target segmented image is selected from the segmented image set as at least one adjacent image. The adjacent image may be image content in the target image that is in proximity to the target segmented image.
[0096] like Figure 9 As shown, a schematic diagram of a target segmented image and at least one adjacent image is shown.
[0097] The at least one adjacent image is each segmented image corresponding to an outer encirclement of the target segmented image.
[0098] The third step is to combine the at least one adjacent image and the target segmented image to generate a combined image.
[0099] The fourth step is to generate image prompt information and a subset of divided image feature information corresponding to the above-mentioned combined image, which are used as the third image prompt information and the target divided image feature information subset respectively.
[0100] In the fifth step, the target segmentation image feature information subset and the third image prompt information are input into the target object segmentation model to generate a combined mask image, wherein the combined mask image can represent the position information of the target object in the image area corresponding to the combined image.
[0101] In the sixth step, the sub-mask image corresponding to the target segmentation image is selected from the combined mask image.
[0102] As an example, the execution entity first determines a corresponding image segmentation method based on an image combination method corresponding to at least one adjacent image and the target segmented image. Then, based on the image segmentation method, the execution entity selects a sub-mask image corresponding to the target segmented image from the combined mask image.
[0103] In the seventh step, the image feature information corresponding to the sub-mask image, the mask image feature information, the segmented image feature information and the second image prompt information are input into the target object segmentation model to generate the local mask image.
[0104] Here, the mask image feature information represents the mask content corresponding to the segmentation image feature information after the target image is segmented into the target object. However, the first resolution and the second resolution may be quite different, resulting in the generated mask image feature information corresponding to one pixel of the mask image may correspond to multiple pixels on the segmentation image. In the case where the pixel correspondence ratio is large, the mask image feature information corresponds to the mask image more seriously, and cannot provide more and higher-quality mask feature information for the target object segmentation model. Therefore, the present disclosure determines the target object segmentation result in the corresponding area of the combined image by taking the target segmentation image as the center and at least one adjacent image as the surrounding image information corresponding to the target segmentation image. The ratio of the corresponding resolution of the combined image to the corresponding resolution of the segmentation image is less than the ratio of the corresponding resolution of the target image to the corresponding resolution of the segmentation image. Thus, by inputting the regional mask feature centered on the target segmentation image into the target object segmentation model, the model learns mask content of higher quality relative to the mask image feature information, so that the subsequent local mask image is more accurate.
[0105] Step 205 : generating object segmentation information corresponding to the target image according to the obtained local mask image set.
[0106] In some embodiments, the execution entity may generate object segmentation information corresponding to the target image based on the obtained local mask image set. The object segmentation information may be image position information of the target object in the target image, or an image after semantic segmentation of the target object.
[0107] As an example, first, the execution subject may combine the local mask images in the local mask image set according to the image segmentation method to generate a combined mask image. Then, object segmentation information corresponding to the target object is generated based on the combined mask image.
[0108] See also Figure 10 , shows the overall flow chart of target object segmentation based on overall image content.
[0109] For the top-level segmentation flow (i.e., performing overall object segmentation on the target image), the first image prompt information, including the positive and negative point sets, and the corresponding overall image feature information (i.e., the Top Embedding in the figure) are input into the encoding and decoding model to obtain the overall mask image. The Top Embedding is an n*n*d vector. The encoding and decoding model includes an encoding layer (the Encoder in the figure) and a decoding layer (the Decoder in the figure). The overall mask image is then split into four partitioned mask images (i.e., the partitioned mask image set) according to the partitioning method of the partitioned image set.
[0110] For the grid-layer segmentation flow (i.e., the process of generating a local mask image), the partition mask image at the upper left corner, the corresponding partition image feature information (i.e., Grid (1, 1) Embedding), and the second image hint information at the upper left corner are input into the encoding and decoding model to obtain a local mask image below the upper left corner. The partition mask image at the upper right corner, the corresponding partition image feature information (i.e., Grid (1, 2) Embedding), and the second image hint information at the upper right corner are input into the encoding and decoding model to obtain a local mask image below the upper right corner. The partition mask image at the lower left corner, the corresponding partition image feature information (i.e., Grid (2, 1) Embedding), and the second image hint information at the lower left corner are input into the encoding and decoding model to obtain a local mask image below the lower left corner. The partition mask image at the lower right corner, the corresponding partition image feature information (i.e., Grid (2, 2) Embedding), and the second image hint information at the lower right corner are input into the encoding and decoding model to obtain a local mask image below the lower right corner. The local mask image under the upper left corner, the local mask image under the lower left corner, the local mask image under the upper right corner, and the local mask image under the lower right corner are spliced to obtain object segmentation information.
[0111] The above-described embodiments of the present disclosure have the following advantageous effects: Through the image segmentation methods of some embodiments of the present disclosure, accurate image segmentation of a target image is achieved through a multi-level segmentation approach, while ensuring the same resolution between the aggregated mask image and the input image. Specifically, the reason for the inaccurate segmentation result of the target image is that due to the significant downsampling operation of the image segmentation model, the resolution of the output mask image is significantly different from that of the target image, resulting in a single segmentation pixel in the mask image corresponding to multiple pixels in the target image, resulting in an inaccurate image segmentation result. Based on this, the image segmentation methods of some embodiments of the present disclosure first obtain a set of segmented image feature information and overall image feature information corresponding to the target image, wherein the first resolution of the segmented image corresponding to the segmented image is the same as the second resolution of the mask image output by the target object segmentation model. Here, by pre-segmenting the target image, the resolution of the segmented image is ensured to be the same as the output resolution of the target object segmentation model. Therefore, when the target object segmentation model receives the segmented image as input, the output result does not suffer from the significant resolution difference. By determining the segmentation result corresponding to each segmented image from the perspective of the segmented image, and then comprehensively considering each segmentation result, the precise segmentation result corresponding to the target image can be accurately obtained. By obtaining the segmented image feature information set, the image feature semantic content corresponding to each segmented image is determined, so that the subsequent target object segmentation model can obtain the image semantic content corresponding to the segmented image to accurately generate a local mask image. By obtaining the overall image feature information, the subsequent target object segmentation model can obtain the image semantic content corresponding to the target image to accurately generate an overall mask image. Then, based on the above-mentioned overall image feature information, using the above-mentioned target object segmentation model, an overall mask image can be accurately generated to provide preliminary image segmentation results for each subsequent segmented image, thereby improving the accuracy of the generation of local mask images. Next, according to the image segmentation method corresponding to the segmented image set, the above-mentioned overall mask image is segmented to obtain a segmented mask image set to provide corresponding preliminary image segmentation results for each subsequent mask image. Furthermore, in response to determining to segment the target object based on the overall image content, for each segmented image feature information in the segmented image feature information set, based on the segmented image feature information and the segmented mask image that has a positional correspondence, the target object segmentation model can be used to accurately generate a local mask image corresponding to the segmented image feature information. Here, the resolution corresponding to the local mask image is the same as the resolution corresponding to the segmented image, so the pixels in the subsequently generated local mask image have a one-to-one correspondence with the pixels in the segmented image. Therefore, in the process of generating the local mask image, there will be no reduction in segmentation accuracy due to downsampling. In this way, the local mask image can be accurately generated.Finally, based on the obtained local mask image set, the object segmentation information corresponding to the target image can be accurately generated. In summary, by dividing the target image, the resolution of the obtained divided images is equal to the resolution of the output mask image, so that there is no pixel segmentation deviation in determining the event of each divided image corresponding to the local mask image. Therefore, while ensuring that the resolution of the aggregated mask image and the input image is the same, accurate image segmentation of the target image can be achieved through a multi-level segmentation method.
[0112] Further references Figure 11 , shows a process 1100 of another embodiment of the image segmentation method according to the present disclosure. The image segmentation method includes the following steps:
[0113] Step 1101: Obtain the segmented image feature information set and the overall image feature information corresponding to the target image.
[0114] Step 1102 : Generate an overall mask image based on the overall image feature information and using the target object segmentation model.
[0115] Step 1103 : performing image division on the overall mask image according to the image division method corresponding to the division image set to obtain a division mask image set.
[0116] Step 1104, in response to determining to segment the target object based on the overall image content, for each segmentation image feature information in the above-mentioned segmentation image feature information set, based on the above-mentioned segmentation image feature information and the segmentation mask image with a position correspondence, the above-mentioned target object segmentation model is used to generate a local mask image corresponding to the above-mentioned segmentation image feature information.
[0117] Step 1105 : generating object segmentation information corresponding to the target image according to the obtained local mask image set.
[0118] In some embodiments, the specific implementation of steps 1101-1105 and the technical effects thereof can be referred to in Figure 2 Steps 201-205 in the corresponding embodiment are not repeated here.
[0119] Step 1106 : In response to determining to segment the target object based on the local image content, obtain a bounding box size ratio and a bounding box defined on the target image.
[0120] In some embodiments, in response to determining to segment the target object based on the local image content, the subject (eg Figure 1The electronic device 101 shown) can obtain the bounding box size ratio and the bounding box circled on the target image. The local image content can be the local (partial) image content in the target image. Segmenting the target object based on the local image content can be segmenting the target object on the local content in the target image. The bounding box size ratio can be the ratio between the bounding box size and the resolution size corresponding to the local image content. The bounding box size can be the pixel frame size set by the bounding box. The bounding box size ratio can be pre-set. For example, the bounding box size ratio can be 0.9:1. The bounding box circled on the target image can be a bounding box circled on the target image that focuses on segmenting the content in the bounding box.
[0121] Step 1107 : Select a partitioned image subset from the partitioned image set according to the bounding box size ratio and the bounding box.
[0122] In some embodiments, the execution entity may select a subset of divided images from the set of divided images based on the bounding box size ratio and the bounding box, wherein the subset of divided images is each adjacent image having a corresponding total resolution and the bounding box size that meets the bounding box size ratio.
[0123] As an example, the execution subject may expand the region according to the position of the bounding box on the target image and the size of the bounding box to obtain an image expansion region, and determine each segmented image corresponding to the image expansion region as a segmented image subset.
[0124] Step 1108 : For the divided image feature information subset corresponding to the divided image subset, the target object segmentation model is used to generate local object segmentation information corresponding to the divided image subset.
[0125] In some embodiments, the execution entity may generate the local object segmentation information corresponding to the divided image subset using the target object segmentation model based on the divided image feature information subset corresponding to the divided image subset.
[0126] As an example, the execution entity may directly input the divided image feature information subset into the target object segmentation model to generate local object segmentation information.
[0127] In some optional implementations of some embodiments, the execution entity may generate local object segmentation information corresponding to the divided image subset using the target object segmentation model based on the divided image feature information subset corresponding to the divided image subset, including the following steps:
[0128] The first step is to determine the segmentation mask image corresponding to each segmentation image in the segmentation image subset to obtain the segmentation mask image subset, wherein the segmentation mask image subset is a subset in the segmentation mask image set.
[0129] In the second step, the above-mentioned local object segmentation information is generated according to the above-mentioned segmentation mask image subset and the above-mentioned segmentation image feature information subset using the above-mentioned target object segmentation model.
[0130] As an example, in response to determining that the target object segmentation model is an interactive segmentation model, fourth image hint information corresponding to the aforementioned partitioned image feature information subset is determined. Then, the aforementioned partitioned mask image subset is input into a vector transformation model to generate a partitioned mask image feature information subset. Finally, the fourth image hint information, the aforementioned partitioned mask image feature information subset, and the partitioned image feature information subset are input into the target object segmentation model to generate local object segmentation information.
[0131] As another example, in response to determining that the target object segmentation model is a non-interactive segmentation model, the execution entity may directly input the partitioned mask image subset and the partitioned image feature information subset into the target object segmentation model to obtain local object segmentation information.
[0132] from Figure 11 It can be seen that Figure 2 Compared with the description of some corresponding embodiments, Figure 11 In the process 1100 of the image segmentation method in some corresponding embodiments, by setting the size ratio of the bounding box, the target object segmentation of the local image content on the target image can be accurately and flexibly achieved.
[0133] Further references Figure 12 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an image segmentation device. These device embodiments are similar to Figure 2 Corresponding to the method embodiments shown, the image segmentation device can be specifically applied to various electronic devices.
[0134] like Figure 12As shown, an image segmentation device 1200 includes: an acquisition unit 501 , a first generation unit 1202 , an image division unit 1203 , a second generation unit 1204 and a third generation unit 1205 . Among them, the acquisition unit 1201 is configured to acquire the segmented image feature information set and the overall image feature information corresponding to the target image, wherein the first resolution of the segmented image corresponding to the segmented image feature information is the same as the second resolution of the mask image output by the target object segmentation model; the first generation unit 1202 is configured to generate an overall mask image based on the above-mentioned overall image feature information using the above-mentioned target object segmentation model; the image division unit 1203 is configured to perform image division on the above-mentioned overall mask image according to the image division method corresponding to the segmented image set to obtain a segmented mask image set; the second generation unit 1204 is configured to, in response to determining that the target object is segmented based on the overall image content, generate a local mask image corresponding to the above-mentioned segmented image feature information for each segmented image feature information in the above-mentioned segmented image feature information set according to the above-mentioned segmented image feature information and the segmented mask image with a position correspondence, using the above-mentioned target object segmentation model; the third generation unit 1205 is configured to generate the object segmentation information corresponding to the above-mentioned target image based on the obtained local mask image set.
[0135] In some optional implementations of some embodiments, the target object segmentation model is an interactive segmentation model; and the first generation unit 1202 can be further configured to: generate first image prompt information corresponding to the target image; input the overall image feature information and the first image prompt information into the interactive segmentation model to generate the overall mask image.
[0136] In some optional implementations of some embodiments, the above-mentioned second generation unit 1204 can be further configured to: generate second image prompt information corresponding to the above-mentioned divided mask image; input the above-mentioned divided mask image into a vector transformation model to generate mask image feature information; input the above-mentioned mask image feature information, the above-mentioned divided image feature information and the above-mentioned second image prompt information into the above-mentioned target object segmentation model to generate the above-mentioned local mask image.
[0137] In some optional implementations of some embodiments, the above-mentioned device 1200 further includes: an information acquisition unit, a selection unit, a screening unit and a fourth generation unit (not shown in the figure). The above-mentioned information acquisition unit can be configured to: in response to determining to segment the target object based on the local image content, obtain the bounding box size ratio and the bounding box circled on the target image. The selection unit can be configured to: select a divided image subset from the divided image set based on the bounding box size ratio and the bounding box, wherein the divided image subset is each adjacent image that meets the bounding box size ratio between the total resolution of the corresponding image and the bounding box size. The fourth generation unit can be configured to: generate object segmentation information corresponding to the divided image subset using the target object segmentation model based on the divided image feature information subset corresponding to the divided image subset.
[0138] In some optional implementations of some embodiments, the above-mentioned second generation unit 1204 can be further configured to: determine the segmented image corresponding to the above-mentioned segmented image feature information as the target segmented image; filter out at least one segmented image that has an image proximity relationship with the above-mentioned target segmented image from the above-mentioned segmented image set as at least one adjacent image; combine the above-mentioned at least one adjacent image and the above-mentioned target segmented image to generate a combined image; generate image prompt information and a subset of segmented image feature information corresponding to the above-mentioned combined image as the third image prompt information and the target segmented image feature information subset, respectively; input the above-mentioned target segmented image feature information subset and the above-mentioned third image prompt information into the above-mentioned target object segmentation model to generate a combined mask image; filter out the sub-mask image corresponding to the above-mentioned target segmented image from the above-mentioned combined mask image; input the image feature information corresponding to the above-mentioned sub-mask image, the above-mentioned mask image feature information, the above-mentioned segmented image feature information and the above-mentioned second image prompt information into the above-mentioned target object segmentation model to generate the above-mentioned local mask image.
[0139] In some optional implementations of some embodiments, the above-mentioned fourth generation unit can be further configured to: determine the segmentation mask image corresponding to each segmentation image in the above-mentioned segmentation image subset to obtain the segmentation mask image subset; based on the above-mentioned segmentation mask image subset and the above-mentioned segmentation image feature information subset, use the above-mentioned target object segmentation model to generate the above-mentioned local object segmentation information.
[0140] In some optional implementations of some embodiments, the image segmentation device 1200 also includes: filtering out target partitioning mask images whose corresponding mask area ratio is higher than the target ratio from the partitioning mask image set to obtain at least one target partitioning mask image; determining the at least one target partitioning mask image as a partitioning mask image set, and determining at least one partitioning image feature information corresponding to the at least one target partitioning mask image as a partitioning image feature information set.
[0141] It is understandable that the units described in the image segmentation device 1200 are similar to those described in the reference Figure 2 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the image segmentation device 1200 and the units included therein, and will not be described in detail here.
[0142] Reference below Figure 13 , which shows an electronic device (eg, Figure 1 Schematic diagram of the structure of the electronic device 101)1300. Figure 13 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0143] like Figure 13 As shown, the electronic device 1300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1301, which can perform various appropriate actions and processes according to a program stored in a read-only memory 1302 or a program loaded from a storage device 1308 into a random access memory 1303. Various programs and data required for the operation of the electronic device 1300 are also stored in the random access memory 1303. The processing device 1301, the read-only memory 1302, and the random access memory 1303 are connected to each other via a bus 1304. An input / output interface 1305 is also connected to the bus 1304.
[0144] Typically, the following devices may be connected to the input / output interface 1305: an input device 1306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1309. The communication device 909 may allow the electronic device 1300 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 13 The electronic device 1300 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 13Each block shown in the figure may represent one device, or may represent multiple devices as needed.
[0145] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 1309, or installed from the storage device 1308, or installed from the read-only memory 1302. When the computer program is executed by the processing device 1301, the above-mentioned functions defined in the method of some embodiments of the present disclosure are performed.
[0146] It should be noted that in some embodiments of the present disclosure, the computer-readable medium mentioned above may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0147] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0148] The computer-readable medium may be included in the electronic device, or may exist separately and not be incorporated into the electronic device. The computer-readable medium carries one or more programs, and when executed by the electronic device, causes the electronic device to: obtain a segmented image feature information set and overall image feature information corresponding to a target image, wherein the first resolution of the segmented image corresponding to the segmented image feature information is the same as the second resolution of the mask image output by the target object segmentation model; generate an overall mask image based on the overall image feature information using the target object segmentation model; perform image segmentation on the overall mask image according to the image segmentation method corresponding to the segmented image set to obtain a segmented mask image set; in response to determining that the target object is to be segmented based on the overall image content, generate a local mask image corresponding to each segmented image feature information in the segmented image feature information set using the target object segmentation model based on the segmented image feature information and the segmented mask image with a corresponding position relationship; and generate object segmentation information corresponding to the target image based on the obtained local mask image set.
[0149] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0150] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0151] The units described in some embodiments of the present disclosure may be implemented by software or by hardware. The described units may also be provided in a processor. For example, they may be described as follows: a processor includes an acquisition unit, a first generation unit, an image division unit, a second generation unit, and a third generation unit. The names of these units do not, in some cases, constitute limitations on the units themselves. For example, the acquisition unit may also be described as a "unit for acquiring a set of divided image feature information and overall image feature information corresponding to a target image."
[0152] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0153] Some embodiments of the present disclosure further provide a computer program product, including a computer program, which implements any of the above-mentioned image segmentation methods when executed by a processor.
[0154] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. An image segmentation method, comprising: Obtaining a segmented image feature information set and overall image feature information corresponding to the target image, wherein a first resolution of the segmented image corresponding to the segmented image feature information is the same as a second resolution of the mask image output by the target object segmentation model; generating an overall mask image using the target object segmentation model according to the overall image feature information; Performing image division on the overall mask image according to an image division method corresponding to the division image set to obtain a division mask image set; In response to determining to segment the target object based on the overall image content, for each segmented image feature information in the segmented image feature information set, generating a local mask image corresponding to the segmented image feature information using the target object segmentation model based on the segmented image feature information and the segmented mask image having a positional correspondence relationship; According to the obtained local mask image set, object segmentation information corresponding to the target image is generated.
2. The method according to claim 1, wherein The target object segmentation model is an interactive segmentation model; as well as The step of generating an overall mask image based on the overall image feature information and utilizing the target object segmentation model comprises: generating first image prompt information corresponding to the target image; The overall image feature information and the first image prompt information are input into the interactive segmentation model to generate the overall mask image.
3. The method according to claim 1, wherein The method of generating a local mask image corresponding to the segmented image feature information by using the target object segmentation model based on the segmented image feature information and the segmented mask image having a position correspondence relationship comprises: generating second image prompt information corresponding to the segmentation mask image; Inputting the segmented mask image into a vector transformation model to generate mask image feature information; The mask image feature information, the segmented image feature information and the second image prompt information are input into the target object segmentation model to generate the local mask image.
4. The method according to claim 1, wherein The method further comprises: In response to determining to segment the target object based on the local image content, obtaining a bounding box size ratio and a bounding box circled on the target image; selecting, from the divided image set, a divided image subset according to the bounding box size ratio and the bounding box, wherein the divided image subset is each adjacent image having a total resolution of the corresponding image and a size of the bounding box that meets the bounding box size ratio; According to the divided image feature information subset corresponding to the divided image subset, the object segmentation information corresponding to the divided image subset is generated using the target object segmentation model.
5. The method according to claim 1, wherein The divided image feature information set and the overall image feature information are generated by the following steps: Determining a resolution of the mask image outputted by the target object segmentation model as a second resolution; Dividing the target image according to the first resolution and the second resolution corresponding to the target image to obtain a set of divided images, wherein the resolution corresponding to each divided image is the same as the second resolution; Inputting each segmented image in the segmented image set into a vector conversion model to generate segmented image feature information, thereby obtaining a segmented image feature information set; The target image is input into the vector transformation model to generate the overall image feature information.
6. The method according to claim 3, wherein: The step of inputting the mask image feature information, the segmented image feature information, and the second image prompt information into the target object segmentation model to generate the local mask image includes: Determine a segmented image corresponding to the segmented image feature information as a target segmented image; Selecting at least one segmented image that has an image proximity relationship with the target segmented image from the segmented image set as at least one adjacent image; combining the at least one adjacent image and the target segmented image to generate a combined image; generating image prompt information and a divided image feature information subset corresponding to the combined image as third image prompt information and a target divided image feature information subset respectively; Inputting the target segmentation image feature information subset and the third image prompt information into the target object segmentation model to generate a combined mask image; Filtering out a sub-mask image corresponding to the target segmented image from the combined mask image; The image feature information corresponding to the sub-mask image, the mask image feature information, the segmented image feature information and the second image prompt information are input into the target object segmentation model to generate the local mask image.
7. The method according to claim 4, wherein: The generating of the local object segmentation information corresponding to the divided image subset by using the target object segmentation model according to the divided image feature information subset corresponding to the divided image subset includes: Determining a partition mask image corresponding to each partition image in the partition image subset to obtain a partition mask image subset; The local object segmentation information is generated according to the segmentation mask image subset and the segmentation image feature information subset using the target object segmentation model.
8. The method according to claim 1, wherein After performing image division on the overall mask image according to the image division method corresponding to the divided image set to obtain the divided mask image set, the method further includes: Filtering target segmentation mask images whose corresponding mask area ratio is higher than the target ratio from the segmentation mask image set to obtain at least one target segmentation mask image; The at least one target segmentation mask image is determined as a segmentation mask image set, and at least one segmentation image feature information corresponding to the at least one target segmentation mask image is determined as a segmentation image feature information set.
9. An image segmentation device, comprising: an acquisition unit configured to acquire a segmented image feature information set and overall image feature information corresponding to the target image, wherein a first resolution of the segmented image corresponding to the segmented image feature information is the same as a second resolution of the mask image output by the target object segmentation model; a first generating unit configured to generate an overall mask image based on the overall image feature information and using the target object segmentation model; an image division unit configured to perform image division on the overall mask image according to an image division method corresponding to the division image set to obtain a division mask image set; a second generating unit configured to, in response to determining to segment the target object based on the overall image content, generate, for each segmented image feature information in the segmented image feature information set, a local mask image corresponding to the segmented image feature information based on the segmented image feature information and a segmented mask image having a positional correspondence, using the target object segmentation model; The third generating unit is configured to generate object segmentation information corresponding to the target image according to the obtained local mask image set.
10. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 8.
11. A computer-readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
12. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.