Method and apparatus for image processing, device, and storage medium

By employing sub-image segmentation and background segmentation techniques, the background segmentation problem under different image layouts is solved, achieving accuracy and stability in image segmentation and adapting to diverse image layouts.

WO2025222884A1PCT designated stage Publication Date: 2025-10-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/139157
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-25
Filing Date
2024-12-13
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing technologies cannot effectively handle image background segmentation under different image layouts, resulting in inaccurate segmentation results and affecting subsequent image processing or analysis.

Method used

By acquiring the target image and layout information, sub-image segmentation and background segmentation are performed to generate sub-mask images of the sub-images. Based on the sub-image mask images and the size of the target image, the mask image of the target image is determined, thereby achieving accurate identification of the foreground and background regions.

Benefits of technology

It improves the accuracy and stability of image background segmentation, can adapt to images with different layouts, avoids interference, and obtains high-precision segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024139157_30102025_PF_FP_ABST
    Figure CN2024139157_30102025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a method and apparatus for image processing, a device, and a storage medium. The method comprises: acquiring a target image and layout information, wherein the target image comprises at least one sub-image, and the layout information indicates the layout of each sub-image among the at least one sub-image in the target image; on the basis of the target image and the layout information, performing sub-image segmentation on the target image to obtain at least one sub-image; separately performing background segmentation on the at least one obtained sub-image to generate respective sub-mask maps of the at least one sub-image, wherein each sub-mask map is used for identifying a foreground area and a background area in the corresponding sub-image; and determining a mask map of the target image on the basis of the respective sub-mask maps of the at least one sub-image and the size of the target image, wherein the mask map is used for identifying a foreground area and a background area in the target image. In this way, the mask map of the target image having high accuracy can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus, devices and storage media for image processing

[0001] This application claims priority to Chinese Patent Application No. 202410509950.0, filed on April 25, 2024, entitled "Method, Apparatus, Device and Storage Medium for Image Processing", the entire contents of which are incorporated herein by reference. Technical Field

[0002] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatuses, devices, and computer-readable storage media for image processing. Background Technology

[0003] With the rapid development of computer technology, image background segmentation algorithms (also known as background separation) have been widely used in the field of image processing. Background segmentation separates the foreground and background regions in an image. Background segmentation can be used for image analysis. However, related techniques cannot effectively handle background segmentation in different image layouts. Summary of the Invention

[0004] In a first aspect of this disclosure, a method for image processing is provided. The method includes: acquiring a target image and layout information, the target image including at least one sub-image, the layout information indicating the layout of each sub-image in the target image; performing sub-image segmentation on the target image based on the target image and the layout information to segment at least one sub-image; generating a sub-mask image for each of the at least one sub-image by performing background segmentation on each of the segmented sub-images, each sub-mask image being used to identify a foreground region and a background region in the corresponding sub-image; and determining a mask image for the target image based on the sub-mask images of each of the at least one sub-images and the dimensions of the target image, the mask image being used to identify the foreground region and the background region in the target image.

[0005] In a second aspect of this disclosure, an apparatus for image processing is provided. The apparatus includes: an image acquisition module configured to acquire a target image and layout information, the target image including at least one sub-image, the layout information indicating the layout of each sub-image in the target image; an image segmentation module configured to perform sub-image segmentation on the target image based on the target image and the layout information to segment at least one sub-image; a background segmentation module configured to generate a sub-mask image for each of the at least one segmented sub-images by performing background segmentation on each of the segmented sub-images, each sub-mask image identifying a foreground region and a background region in the corresponding sub-image; and an image determination module configured to determine a mask image for the target image based on the sub-mask images of each of the at least one sub-images and the dimensions of the target image, the mask image identifying the foreground region and the background region in the target image.

[0006] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.

[0008] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.

[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0011] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0012] Figure 2 shows a flowchart of a process for image processing according to some embodiments of the present disclosure;

[0013] Figure 3 shows a schematic diagram of an example of image processing according to some embodiments of the present disclosure;

[0014] Figure 4 shows a flowchart of a process for image processing according to some embodiments of the present disclosure;

[0015] Figure 5 shows a schematic structural block diagram of an apparatus for image processing according to certain embodiments of the present disclosure; and

[0016] Figure 6 shows a block diagram of an electronic device capable of implementing one or more embodiments of the present disclosure. Detailed Implementation

[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0018] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0019] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.

[0020] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0021] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.

[0022] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information, thereby enabling the user to choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers or storage media that perform the operation of the technical solution disclosed herein, based on the prompt message.

[0023] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, such as a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0024] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0025] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0026] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Environment 100 includes one or more users 110-1, 110-2, 110-3, ..., 110-N who can send and receive messages through their respective associated terminal devices 120-1, 120-2, 120-3, ..., 120-N. For ease of discussion, users 110-1, 110-2, 110-3, ..., 110-N can be collectively referred to as user 110 or individually, and terminal devices 120-1, 120-2, 120-3, ..., 120-N can be collectively referred to as terminal device 120 or individually. In some scenarios, user 110 can publish and comment on works, live streams, etc., on a target platform through associated terminal devices 120.

[0027] Terminal device 120 may have an application 125 that supports message interaction installed (i.e., application 125-1 is installed on terminal device 120-1, application 125-2 is installed on terminal device 120-2, application 125-3 is installed on terminal device 120-3, ..., application 125-N is installed on terminal device 120-N). It should be noted that the application 125 installed on different terminal devices 120 can be the exact same application or different applications (e.g., different versions). Application 125 can be any suitable application with message sending and receiving functions, such as a dedicated chat application, a social application, a content sharing application, a content (image) editing application, an office support application, etc.

[0028] In environment 100 of Figure 1, if application 125 is active, terminal device 120 can acquire an image of application 125. This image may include various interfaces provided by application 125, such as user interfaces supporting message interaction, user interfaces supporting content browsing, message sending and receiving interfaces, etc. Through different interfaces, application 125 can provide different content to user 110. Through appropriate means, such as clicking or selecting any appropriate element in the user interface, application 125 can also provide user 110 with the option to select and switch the presentation mode of related content.

[0029] In some embodiments, different terminal devices 120 can also communicate with server devices 130 via network 132 to provide services to application 125.

[0030] Terminal device 120 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 120 may also support any type of user-facing interface (such as "wearable" circuitry). Server device 130 can be various types of computing systems / servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0031] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0032] As briefly described above, with the rapid development of computer technology, image background segmentation has a wide range of applications in the field of image processing. Current image background segmentation techniques can perform image segmentation under certain layout conditions. However, this method, based on the acquired image, cannot effectively handle situations where the layout of the acquired image varies greatly. For example, the acquired image may contain not only the relevant valid image in the current scene but also other interfering background elements. When performing image background segmentation on such an image, the results will be affected by these interfering background elements, leading to inaccurate results. If the image background segmentation is inaccurate or does not meet the scene requirements, subsequent image processing or analysis may not achieve the expected results. For example, it may be impossible to guarantee that the image transmitted to the rendering engine is the valid image to be rendered (it might only be a foreground or background area).

[0033] In view of this, embodiments of the present disclosure provide an improved scheme for image processing. In this scheme, a target image and layout information are acquired. The target image includes at least one sub-image, and the layout information indicates the layout of each sub-image within the target image. Based on the target image and the layout information, sub-image segmentation is performed on the target image to segment at least one sub-image. Background segmentation is then performed on each of the segmented sub-images to generate a sub-mask image for each sub-image. Each sub-mask image is used to identify the foreground and background regions in the corresponding sub-image. Finally, based on the sub-masks of each sub-image and the dimensions of the target image, a mask image for the target image is determined. This mask image is used to identify the foreground and background regions in the target image. This approach effectively handles background segmentation of images with different layouts, avoids interference, and achieves accurate and stable background segmentation results.

[0034] The various example implementations of this disclosure will be described in detail below.

[0035] Figure 2 shows a flowchart of a process 200 for image processing according to some embodiments of the present disclosure. The following description will refer to the environment 100 of Figure 1 for ease of discussion.

[0036] In some embodiments, process 200 may be implemented at terminal device 120, for example, terminal device 120 performs processing on an image to be processed or presented locally. In some embodiments, process 200 may also be implemented at server device 130, for example, server device 130 remotely processes an image to be presented at terminal device 120, or server device 130 processes an image for image analysis purposes at the server.

[0037] In the embodiments described below, for ease of description only, process 200 is implemented at terminal device 120 as an example, but the same operation can also be implemented in server device or any other suitable electronic device.

[0038] In frame 210, terminal device 120 acquires target image and layout information.

[0039] The target image includes at least one sub-image (also referred to as a "sub-screen"). In some embodiments, in addition to the areas occupied by the individual sub-images, the target image may also include an image display area. The image display area refers to a drawable or visible area in the image defined by certain height and width attributes. Such an image is sometimes also referred to as a non-full-screen image or non-full-screen screen.

[0040] The layout information is used to indicate the layout of each sub-image in the target image within at least one sub-image. That is, the target image may include one or more sub-images. If the target image includes one sub-image, the layout information indicates the layout of that sub-image within the target image. If the target image includes multiple sub-images, the layout information indicates the layout of each sub-image within the target image. In some embodiments, the layout information of each sub-image in the input image can be set in the terminal device 120.

[0041] In some embodiments, the layout information includes the following information corresponding to each sub-image: the coordinates of a predetermined vertex in the sub-image in the target image; and the size of the sub-image.

[0042] As an example, the layout information can be represented as Block{x,y,w,h}. x represents the x-coordinate of a predetermined vertex of the sub-image relative to the target image, and y represents the y-coordinate of the first predetermined vertex of the sub-image relative to the target image. w represents the width of the sub-image, and h represents the height of the sub-image. The predetermined vertex can be, for example, the top-left corner of the sub-image, or it could be the bottom-left corner, top-right corner, bottom-right corner, center point, etc. The predetermined vertex can be used as a positioning reference for the sub-image within the target image.

[0043] In frame 220, terminal device 120 performs sub-image segmentation on the target image.

[0044] Terminal device 120 performs sub-image segmentation on the target image based on the target image and layout information to segment at least one sub-image. If the target image contains one sub-image, the segmented sub-image is obtained after sub-image segmentation. If the target image contains multiple sub-images, the segmented sub-images are obtained after sub-image segmentation.

[0045] In some embodiments, the target image includes multiple sub-images. The terminal device 120 can initiate multiple first threads corresponding to the multiple sub-images respectively. Each first thread is configured to segment the sub-image from the target image based on the layout information of one of the multiple sub-images, and the multiple first threads execute in parallel. By using multi-threaded parallel execution, the efficiency of sub-image segmentation can be improved.

[0046] In frame 230, terminal device 120 performs background segmentation on each sub-image.

[0047] Terminal device 120 performs background segmentation on at least one segmented sub-image to generate a sub-mask image for each sub-image. Each sub-mask image identifies the foreground and background regions in the corresponding sub-image. Typically, each sub-mask image has the same size as the corresponding sub-image. In some embodiments, the sub-mask image corresponding to each sub-image can be a binary image, where each pixel value is either a first value or a second value. The first value indicates that the corresponding pixel belongs to the foreground region, and the second value indicates that the corresponding pixel belongs to the background region. For example, the first value can be 1, and the second value can be 0, or the opposite values ​​can be used. In this way, multiplying the sub-mask image with the corresponding sub-image can separate the foreground and background regions from the sub-image.

[0048] Since each sub-image can be considered an independent image with a complete picture, the terminal device can perform background segmentation on the sub-image according to any appropriate background segmentation algorithm (currently known or developed in the future) to distinguish whether each pixel belongs to the foreground or the background.

[0049] In some embodiments, the target image includes multiple sub-images. The terminal device 120 can initiate multiple second threads corresponding to the multiple sub-images, each second thread being configured to perform background segmentation on one of the multiple sub-images; and the multiple second threads execute in parallel. By using multi-threaded parallel execution, the background segmentation efficiency of the sub-images can be improved. In some embodiments, each second thread can receive the segmented sub-image from the first thread after the processing of each first thread, and use it to perform image background segmentation.

[0050] In frame 240, terminal device 120 determines the mask image of the target image.

[0051] Terminal device 120 determines a mask image of the target image based on the sub-mask images of at least one sub-image and the size of the target image. The mask image is used to identify foreground and background regions in the target image. The mask image of the target image has the same size as the target image.

[0052] In some embodiments, for each of the at least one sub-images, the coordinates of each pixel value in the sub-mask of the sub-image are transformed into coordinates in the target image based on the coordinates of predetermined vertices of the sub-image in the target image and the scaling value of the sub-image relative to the target image. The mask of the target image is initialized; and the initialized mask is updated based on the pixel values ​​corresponding to the coordinates of the respective sub-mask of the at least one sub-image in the target image, to obtain the mask of the target image.

[0053] In some embodiments, the initial mask of the target image is updated based on the pixel values ​​corresponding to the coordinates of the sub-mask images of at least one sub-image in the target image. Specifically, for each sub-mask image, the following determinations are performed: in response to determining that the pixel value corresponding to the first coordinate in the target image of the sub-mask image is a first value, the pixel value corresponding to the first coordinate in the initial mask image is updated to a value indicating a foreground region; and in response to determining that the pixel value corresponding to the first coordinate in the target image of the sub-mask image is a second value, the pixel value corresponding to the first coordinate in the initial mask image is updated to a value indicating a background region.

[0054] In some embodiments, two mask image modes can be set, including a first mask mode and a second mask mode. The first mask mode updates the pixel values ​​in the initialized mask image that have not been updated to values ​​indicating the background area. The second mask mode updates the pixel values ​​in the initialized mask image that have not been updated to values ​​indicating the image display area. These two mask image modes can be applied to different business contexts.

[0055] Figure 3 illustrates a schematic diagram of an example of image processing according to some embodiments of the present disclosure. It should be understood that the target image shown in Figure 3 is merely an example, and various designs are possible in practice. For example, the various graphic elements and / or controls in the target image may have different arrangements and different visual representations, one or more elements and / or controls may be omitted or replaced, and one or more other elements and / or controls may also be present. Furthermore, the target image may contain any suitable content. The scope of the present disclosure is not limited in this respect. For example, the target image may include one sub-image or multiple sub-images.

[0056] In some embodiments, a target image 310 and layout information are acquired. The target image 310 contains multiple sub-images, including sub-images 320, 325, 330, and 335. The layout information includes the coordinates of predetermined vertices in each sub-image of the target image 310 within the target image 310, and the size of each sub-image of the target image 310. The layout information indicates the layout of each sub-image in the target image 310. Sub-images 320, 325, 330, and 335 each include a corresponding foreground region and a background region.

[0057] In some embodiments, the target image 310 includes multiple sub-images, and the electronic device 120 can initiate multiple first threads, each of which corresponds to one sub-image and performs sub-image segmentation. The multiple first threads are executed in parallel.

[0058] In some embodiments, based on the target image 310 and layout information, sub-image segmentation is performed on the target image 310 to segment it into multiple sub-images, including: sub-image 320, sub-image 325, sub-image 330 and sub-image 335.

[0059] In some embodiments, background segmentation is performed on the segmented sub-images to obtain sub-mask images for each sub-image, including: sub-mask image 340, sub-mask image 345, sub-mask image 350, and sub-mask image 355. The correspondence between the multiple sub-mask images and the segmented sub-images is as follows: sub-mask image 340 corresponds to segmented sub-image 320, sub-mask image 345 corresponds to segmented sub-image 325, sub-mask image 350 corresponds to segmented sub-image 330, and sub-mask image 355 corresponds to segmented sub-image 335. Each sub-mask image has the same size as its corresponding sub-image. Each sub-mask image is used to identify the foreground and background regions in the corresponding sub-image.

[0060] In some embodiments, the target image 310 includes multiple sub-images, and the electronic device 120 can initiate multiple second threads, each of which corresponds to one sub-image to perform background segmentation. The multiple second threads can be executed in parallel.

[0061] In some embodiments, a mask 360 or a mask 370 of the target image 310 is determined based on a plurality of sub-mask images and the size of the target image 310. The mask 360 or the mask 370 is used to identify the foreground region and the background region in the target image 310.

[0062] As an example, assuming the predetermined vertex is the top left corner of the sub-mask, the coordinates of each pixel value in sub-mask 340, sub-mask 345, sub-mask 350 and sub-mask 355 are replaced with the coordinates in the target image 310.

[0063] The transformation matrices for the x and y coordinates of each pixel value in the submask image are shown below:

[0064] Formula for the horizontal axis matrix: global x =sx×block x +block_origin x .

[0065] In the above formula for transforming the horizontal coordinate, global_x represents the horizontal coordinate of the current pixel value in the target image 310. sx represents the scaling of the horizontal coordinate of the sub-image corresponding to the current pixel value relative to the original image. block_x represents the horizontal coordinate of the current pixel value in the corresponding sub-image, and block_origin_x represents the horizontal coordinate of the top-left corner of the sub-image in the target image 310.

[0066] As an example, a horizontal coordinate transformation is performed on sub-mask image 340. The horizontal coordinate of the current pixel value in sub-image 320 is obtained by multiplying the product of the scaling of the horizontal coordinate of the current pixel value relative to the original image and the horizontal coordinate of the current pixel value in the corresponding sub-image 320, and then summing this product with the horizontal coordinate of the top-left corner of sub-image 320 in the target image 310. This horizontal coordinate transformation is then performed on each pixel value in sub-mask image 340.

[0067] Vertical coordinate transformation matrix: global y =sy×block y +block o rigin y .

[0068] In the above y-coordinate transformation matrix, global_y represents the y-coordinate of the current pixel value in the target image 310. sy represents the scaling of the y-coordinate of the sub-image corresponding to the current pixel value relative to the original image. block_y represents the y-coordinate of the current pixel value in the corresponding sub-image, and block_origin_y represents the y-coordinate of the top-left corner of the sub-image in the target image 310.

[0069] As an example, a ordinate transformation is performed on sub-mask image 340. The ordinate of the current pixel value in the target image 310 is obtained by multiplying the product of the scaling of the ordinate of the current pixel value in sub-image 320 relative to the original image and the ordinate of the current pixel value in the corresponding sub-image 320, and then summing this product with the ordinate of the top-left corner of sub-image 320 in the target image 310. This ordinate transformation is then performed on each pixel value in sub-mask image 340.

[0070] In some embodiments, mask image 360 ​​or mask image 370 is initialized, i.e., an image of the same size as the target image 305 is created to store mask image 360 ​​or mask image 370. Based on the pixel values ​​corresponding to the coordinates of the sub-masks of each of the multiple sub-images in the target image 310, the initialized mask image 360 ​​or mask image 370 is updated to obtain the mask image 360 ​​or mask image 370 of the target image 310. Each pixel value in the sub-mask corresponding to each of the multiple sub-images can be a first value or a second value. The first value can indicate that the corresponding pixel belongs to the foreground region, and the second value can indicate that the corresponding pixel belongs to the background region.

[0071] In some embodiments, an initialized mask image is updated based on the pixel values ​​corresponding to the coordinates of the respective sub-mask images of the multiple sub-images in the target image. For each sub-mask image, the following is performed: in response to determining that the pixel value corresponding to the first coordinate in the target image 310 of the sub-mask image is a first value, the pixel value corresponding to the first coordinate in the initialized mask image 360 ​​or mask image 370 is updated to a value indicating a foreground region; and in response to determining that the pixel value corresponding to the first coordinate in the target image 310 of the sub-mask image is a second value, the pixel value corresponding to the first coordinate in the initialized mask image 360 ​​or mask image 370 is updated to a value indicating a background region.

[0072] In some embodiments, two mask image modes can be set, including a first mask mode and a second mask mode, which correspond to the example mask image 360 ​​and mask image 370, respectively. The two mask image modes are distinguished based on updating the unupdated pixel values ​​in the initialized mask image to predetermined values. The first mask mode corresponds to updating the unupdated pixel values ​​in the initialized mask image 360 ​​to values ​​indicating the background region, resulting in mask image 360. The multiple pixel values ​​in mask image 360 ​​include values ​​for the foreground region and values ​​for the background region. The second mask mode corresponds to updating the unupdated pixel values ​​in the initialized mask image 370 to values ​​indicating the image display area, resulting in mask image 370. The multiple pixel values ​​in mask image 370 include values ​​for the foreground region, values ​​for the background region, and values ​​for the image display area. These two mask image modes can be applied to different business contexts.

[0073] Figure 4 shows a flowchart of a process 400 for image processing according to some embodiments of the present disclosure. Process 400 can be implemented at a terminal device 120. Process 400 is described below with reference to Figure 1.

[0074] In box 410, terminal device 120 acquires a target image and layout information, the target image including at least one sub-image, and the layout information indicating the layout of each sub-image in the target image.

[0075] In frame 420, terminal device 120 performs sub-image segmentation on the target image based on the target image and layout information to segment at least one sub-image.

[0076] In box 430, terminal device 120 performs background segmentation on at least one segmented sub-image to generate a sub-mask map for each sub-image, with each sub-mask map used to identify the foreground and background regions in the corresponding sub-image.

[0077] In box 440, terminal device 120 determines a mask map of the target image based on the sub-mask map of at least one sub-image and the size of the target image. The mask map is used to identify the foreground and background regions in the target image.

[0078] In some embodiments, the layout information includes the following information corresponding to each sub-image: the coordinates of a predetermined vertex in the sub-image in the target image; and the size of the sub-image.

[0079] In some embodiments, the target image includes a plurality of sub-images, and the terminal device 120 performs sub-image segmentation on the target image by: initiating a plurality of first threads corresponding to the plurality of sub-images respectively, each first thread being configured to segment the sub-image from the target image based on the layout information of one of the sub-images; and enabling the plurality of first threads to execute in parallel.

[0080] In some embodiments, the target image includes a plurality of sub-images, and the terminal device 120 generates a submask map of at least one sub-image by: initiating a plurality of second threads corresponding to the plurality of sub-images respectively, each second thread being configured to perform background segmentation on one of the plurality of sub-images; and causing the plurality of second threads to execute in parallel.

[0081] In some embodiments, each sub-mask has the same size as the corresponding sub-image, and the mask of the target image has the same size as the target image. The terminal device 120 determines the mask of the target image by: for each of at least one sub-images, transforming the coordinates of each pixel value in the sub-mask of the sub-image to coordinates in the target image based on the coordinates of predetermined vertices of the sub-image in the target image and the scaling value of the sub-image relative to the target image; initializing the mask of the target image; and updating the initialized mask based on the pixel values ​​corresponding to the coordinates of the respective sub-mask of the at least one sub-image in the target image, to obtain the mask of the target image.

[0082] In some embodiments, each pixel value in the sub-mask map corresponding to each sub-image is a first value or a second value, the first value indicating that the corresponding pixel belongs to a foreground region and the second value indicating that the corresponding pixel belongs to a background region, and wherein the terminal device 120 updates the initialized mask map based on the pixel values ​​corresponding to the coordinates of the respective sub-mask maps of at least one sub-image in the target image, including: for each sub-mask map, performing the following: in response to determining that the pixel value corresponding to the first coordinate in the target image of the sub-mask map is a first value, updating the pixel value corresponding to the first coordinate in the initialized mask map to a value indicating a foreground region, and in response to determining that the pixel value corresponding to the first coordinate in the target image of the sub-mask map is a second value, updating the pixel value corresponding to the first coordinate in the initialized mask map to a value indicating a background region.

[0083] In some embodiments, the terminal device 120 further includes updating the initialized mask image by: updating unupdated pixel values ​​in the initialized mask image to values ​​indicating background areas in response to determining that the mask image mode of the target image is a first mask mode; and updating unupdated pixel values ​​in the initialized mask image to values ​​indicating image display areas in response to determining that the mask image mode of the target image is a second mask mode.

[0084] In summary, this disclosure obtains more accurate sub-image segmentation results by performing sub-image segmentation and sub-image background segmentation on the sub-image itself. The target image segmentation result is then obtained based on the sub-image segmentation result and the size of the acquired target image. This method effectively handles image segmentation for images with different layouts, avoids interference, and yields highly accurate image segmentation results.

[0085] Figure 5 shows a schematic structural block diagram of an image processing apparatus 500 according to certain embodiments of the present disclosure. The apparatus 500 may be implemented as or included in the terminal device 120. The various modules / components in the apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.

[0086] As shown in the figure, the device 500 includes an image acquisition module 510 configured to acquire a target image and layout information, the target image including at least one sub-image, and the layout information indicating the layout of each sub-image in the target image.

[0087] The apparatus 500 also includes an image segmentation module 520, configured to perform sub-image segmentation on the target image based on the target image and layout information to segment at least one sub-image.

[0088] The apparatus 500 also includes a background segmentation module 530, configured to perform background segmentation on at least one segmented sub-image to generate a sub-mask map for each of the at least one sub-image, each sub-mask map being used to identify the foreground region and background region in the corresponding sub-image.

[0089] The apparatus 500 also includes an image determination module 540 configured to determine a mask of the target image based on a sub-mask of at least one sub-image and the size of the target image, the mask being used to identify foreground and background regions in the target image.

[0090] In some embodiments, the layout information includes the following information corresponding to each sub-image: the coordinates of a predetermined vertex in the sub-image in the target image; and the size of the sub-image.

[0091] In some embodiments, the target image includes multiple sub-images, and the image segmentation module 520 is further configured to initiate multiple first threads corresponding to the multiple sub-images respectively, each first thread being configured to segment the sub-image from the target image based on the layout information of one of the multiple sub-images; and to enable the multiple first threads to execute in parallel.

[0092] In some embodiments, where the target image includes multiple sub-images, the background segmentation module is further configured to initiate multiple second threads corresponding to the multiple sub-images respectively, each second thread being configured to perform background segmentation on one of the multiple sub-images; and to enable the multiple second threads to execute in parallel.

[0093] In some embodiments, where each sub-mask has the same size as the corresponding sub-image, and the mask of the target image has the same size as the target image, the image determination module 540 is further configured to, for each of the at least one sub-images, transform the coordinates of each pixel value in the sub-mask of the sub-image to coordinates in the target image based on the coordinates of a predetermined vertex of the sub-image in the target image and the scaling value of the sub-image relative to the target image; initialize the mask of the target image; and update the initialized mask based on the pixel values ​​corresponding to the coordinates of the respective sub-mask of the at least one sub-image in the target image, to obtain the mask of the target image.

[0094] In some embodiments, each pixel value in the sub-mask image corresponding to each sub-image is a first value or a second value, the first value indicating that the corresponding pixel belongs to the foreground region and the second value indicating that the corresponding pixel belongs to the background region. The image determination module is further configured to perform the following for each sub-mask image: in response to determining that the pixel value corresponding to the first coordinate in the target image of the sub-mask image is the first value, updating the pixel value corresponding to the first coordinate in the initialized mask image to the value indicating the foreground region; and in response to determining that the pixel value corresponding to the first coordinate in the target image of the sub-mask image is the second value, updating the pixel value corresponding to the first coordinate in the initialized mask image to the value indicating the background region.

[0095] In some embodiments, the image determination module is further configured to update the unupdated pixel values ​​in the initialized mask image to values ​​indicating the background region in response to determining that the mask image mode of the target image is a first mask mode; and to update the unupdated pixel values ​​in the initialized mask image to values ​​indicating the image display area in response to determining that the mask image mode of the target image is a second mask mode.

[0096] Figure 6 shows a block diagram illustrating an electronic device 600 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 600 shown in Figure 6 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 600 shown in Figure 6 can be used to implement the terminal device 120 of Figure 1 or the apparatus 500 of Figure 5.

[0097] As shown in Figure 6, the electronic device 600 is in the form of a general-purpose electronic device. Components of the electronic device 600 may include, but are not limited to, one or more processors or processing units 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 950, and one or more output devices 660. The processing unit 610 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 600.

[0098] Electronic device 600 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 630 can be a removable or non-removable medium and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within electronic device 600.

[0099] Electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 6, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 620 may include computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0100] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 600 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0101] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 600 can also communicate with one or more external devices (not shown) via communication unit 640 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 600, or with any device that enables electronic device 600 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0102] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0103] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0104] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0105] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0106] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0107] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for image processing, comprising: Acquire a target image and layout information, wherein the target image includes at least one sub-image, and the layout information indicates the layout of each sub-image in the target image; Based on the target image and the layout information, sub-image segmentation is performed on the target image to segment at least one sub-image; By performing background segmentation on the at least one segmented sub-image, a sub-mask image is generated for each of the at least one sub-images. Each sub-mask image is used to identify the foreground region and background region in the corresponding sub-image. as well as Based on the sub-mask maps of the at least one sub-image and the size of the target image, a mask map of the target image is determined, the mask map being used to identify the foreground and background regions in the target image.

2. The method according to claim 1, wherein the layout information includes the following information corresponding to each sub-image: The coordinates of the predetermined vertex in the target image in the sub-image; and The size of the sub-image.

3. The method of claim 1, wherein the target image comprises a plurality of sub-images, and wherein performing sub-image segmentation on the target image comprises: Multiple first threads corresponding to the multiple sub-images are started, and each first thread is configured to segment the sub-image from the target image based on the layout information of one of the multiple sub-images; as well as This allows the multiple first threads to execute in parallel.

4. The method of claim 1, wherein the target image comprises a plurality of sub-images, and wherein generating a sub-mask image for each of the at least one sub-image comprises: Start multiple second threads corresponding to the multiple sub-images respectively, and each second thread is configured to perform background segmentation on one of the multiple sub-images; as well as This allows the multiple second threads to execute in parallel.

5. The method of claim 1, wherein each sub-mask image has the same size as the corresponding sub-image, and the mask image of the target image has the same size as the target image, wherein determining the mask image of the target image comprises: For each of the at least one sub-images, based on the coordinates of a predetermined vertex of the sub-image in the target image and the scaling value of the sub-image relative to the target image, the coordinates of each pixel value in the sub-mask image of the sub-image are transformed into coordinates in the target image; The mask image of the target image is initialized; as well as Based on the pixel values ​​corresponding to the coordinates of the sub-mask maps of the at least one sub-image in the target image, the initialized mask map is updated to obtain the mask map of the target image.

6. The method of claim 5, wherein each pixel value in the sub-mask image corresponding to each sub-image is a first value or a second value, the first value indicating that the corresponding pixel belongs to the foreground region, and the second value indicating that the corresponding pixel belongs to the background region, and Updating the initialized mask image based on the pixel values ​​corresponding to the coordinates of the sub-mask images of the at least one sub-image in the target image includes: For each submask image, perform the following: In response to determining that the pixel value corresponding to the first coordinate in the target image of the sub-mask image is the first value, the pixel value corresponding to the first coordinate in the initialized mask image is updated to a value indicating the foreground region, and In response to determining that the pixel value corresponding to the first coordinate in the target image of the sub-mask is the second value, the pixel value corresponding to the first coordinate in the initialized mask is updated to a value indicating the background region.

7. The method of claim 6, wherein updating the initialized mask map further comprises: In response to determining that the mask map mode of the target image is a first mask mode, the pixel values ​​in the initialized mask map that have not been updated are updated with values ​​indicating the background region; as well as In response to determining that the mask map mode of the target image is the second mask mode, the pixel values ​​in the initialized mask map that have not been updated are updated with values ​​indicating the image display area.

8. An apparatus for image processing, comprising: An image acquisition module is configured to acquire a target image and layout information, wherein the target image includes at least one sub-image and the layout information indicates the layout of each sub-image in the target image. An image segmentation module is configured to perform sub-image segmentation on the target image based on the target image and the layout information, so as to segment at least one sub-image; The background segmentation module is configured to perform background segmentation on the segmented at least one sub-image respectively, and generate a sub-mask map for each of the at least one sub-image, wherein each sub-mask map is used to identify the foreground region and background region in the corresponding sub-image; as well as An image determination module is configured to determine a mask map of a target image based on the sub-mask maps of the at least one sub-image and the size of the target image, the mask map being used to identify foreground and background regions in the target image.

9. The apparatus of claim 8, wherein the layout information includes the following information corresponding to each sub-image: The coordinates of the predetermined vertex in the target image in the sub-image; and The size of the sub-image.

10. The apparatus of claim 8, wherein the target image comprises a plurality of sub-images, and wherein the image segmentation module is further configured to: Initiate multiple first threads corresponding to the plurality of sub-images, each first thread being configured to segment the sub-image from the target image based on the layout information of one of the sub-images; and This allows the multiple first threads to execute in parallel.

11. The apparatus of claim 8, wherein the target image comprises a plurality of sub-images, and wherein the background segmentation module is further configured to: Start multiple second threads corresponding to the plurality of sub-images, each second thread being configured to perform background segmentation on one of the plurality of sub-images; and This allows the multiple second threads to execute in parallel.

12. The apparatus of claim 8, wherein each sub-mask image has the same size as the corresponding sub-image, and the mask image of the target image has the same size as the target image, wherein the image determination module is further configured to: For each of the at least one sub-images, based on the coordinates of a predetermined vertex of the sub-image in the target image and the scaling value of the sub-image relative to the target image, the coordinates of each pixel value in the sub-mask image of the sub-image are transformed into coordinates in the target image; The mask image of the target image is initialized; as well as Based on the pixel values ​​corresponding to the coordinates of the sub-mask maps of the at least one sub-image in the target image, the initialized mask map is updated to obtain the mask map of the target image.

13. The apparatus of claim 12, wherein each pixel value in the sub-mask image corresponding to each sub-image is a first value or a second value, the first value indicating that the corresponding pixel belongs to the foreground region, and the second value indicating that the corresponding pixel belongs to the background region, and The image determination module is further configured to: For each submask image, perform the following: In response to determining that the pixel value corresponding to the first coordinate in the target image of the sub-mask image is the first value, the pixel value corresponding to the first coordinate in the initialized mask image is updated to a value indicating the foreground region, and In response to determining that the pixel value corresponding to the first coordinate in the target image of the sub-mask is the second value, the pixel value corresponding to the first coordinate in the initialized mask is updated to a value indicating the background region.

14. The apparatus of claim 13, wherein the image determination module is further configured to: In response to determining that the mask map mode of the target image is a first mask mode, the pixel values ​​in the initialized mask map that have not been updated are updated with values ​​indicating the background region; and In response to determining that the mask map mode of the target image is the second mask mode, the pixel values ​​in the initialized mask map that have not been updated are updated with values ​​indicating the image display area.

15. An electronic device comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 7.

16. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 7.

17. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device for instance segmentation, equipment and storage medium

    CN114998592A

  • Image head instance segmentation method and device, electronic equipment and storage medium

    CN117036692A

  • Image instance segmentation method, device and equipment and computer readable storage medium

    CN117456173A

  • Image processing apparatus, imaging apparatus, and image processing method

    US20150116546A1