Image processing method and device, electronic equipment and storage medium
By identifying and processing element information and background images in the image and generating target images, the problems of low and difficult image processing efficiency in the prior art are solved, and accurate identification and editing of image elements are achieved, which facilitates the improvement of image processing efficiency.
Patent Information
- Application Number
- CN202311602501.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to accurately identify and edit different elements in image materials separately, resulting in low image processing efficiency and difficulty.
By identifying the element information of each element in the image to be processed, and determining the background image in the image, a target image is generated. The method includes acquiring the image to be processed, identifying the element information, determining the background image, and generating a target image based on the element information and the background image.
It realizes accurate identification and combination of different elements, which facilitates the subsequent editing and processing of each element, improves image processing efficiency and reduces image processing difficulty.
Smart Images

Figure CN120070655A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing, and in particular, to an image processing method, apparatus, electronic device, and storage medium. Background Art
[0002] In recent years, with the increasing diversification of layout design and the rapid development of image editing technology, images, as planar materials, have been widely used in fields such as commercial promotion. An image material includes various elements, and generally, it is necessary to perform secondary editing and processing on the elements in the image material to improve the utilization rate and economic benefits of the image material.
[0003] However, when using the image processing method of related technologies to perform secondary editing and processing on an image, it is difficult to accurately identify and obtain different types of elements and perform separate editing and processing, resulting in problems such as low image processing efficiency and high image processing difficulty. Summary of the Invention
[0004] To overcome the problems existing in related technologies, the present disclosure provides an image processing method, apparatus, electronic device, and storage medium.
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an image processing method, the image processing method including:
[0006] Obtain an image to be processed;
[0007] Identify the element information corresponding to each element in the image to be processed;
[0008] Determine the background image in the image to be processed;
[0009] Generate a target image according to the element information and the background image.
[0010] In some embodiments of the present disclosure, the elements in the image to be processed include picture elements and / or text elements;
[0011] The identifying the element information corresponding to each element in the image to be processed includes:
[0012] Sequentially identify the picture element information corresponding to the picture elements and the text element information corresponding to the text elements.
[0013] In some embodiments of the present disclosure, the image processing method further includes:
[0014] When the image to be processed includes picture elements, before identifying the text element information corresponding to the text elements, perform a masking process on the picture elements.
[0015] In some embodiments of the present disclosure, the element information includes the position area where the element is located.
[0016] In some embodiments of the present disclosure, identifying the text element information corresponding to the text element includes:
[0017] Determine the position area where the initial text element is located;
[0018] Determine the proportion of the overlapping part between the position area where the initial text element is located and the position area where the picture element is located;
[0019] Determine the initial text element with the proportion less than or equal to the preset threshold as the text element;
[0020] Identify the text information in the text element information corresponding to the text element.
[0021] In some embodiments of the present disclosure, the text information includes text content and text format information.
[0022] In some embodiments of the present disclosure, the element information further includes the layer order information of the layer where the element is located; the image processing method further includes:
[0023] Determine the layer order of each layer where the element is located according to the position area where the element is located.
[0024] In some embodiments of the present disclosure, if the position areas where each element is located do not overlap with each other, determining the layer order of each layer where the element is located according to the position area where the element is located includes:
[0025] Randomly generate the layer order of each layer where the element is located.
[0026] In some embodiments of the present disclosure, if there is an overlap between the position areas of any plurality of elements, determining the layer order of each layer where the element is located according to the position area where the element is located includes:
[0027] Obtain the depth of field map corresponding to the image to be processed;
[0028] Calculate the average gray value of the position area where each element is located in the depth of field map;
[0029] Determine the layer order based on the average gray value.
[0030] In some embodiments of the present disclosure, determining the background image in the image to be processed includes:
[0031] Remove the picture element and the text element in the image to be processed to obtain a background image to be completed;
[0032] Perform image redrawing processing or image completion processing on the background image to be completed, to obtain the background image.
[0033] In some embodiments of the present disclosure, generating a target image according to the element information and the background image includes:
[0034] Determine the layer information of the target image according to the layer order of each layer where the elements are located and the background image;
[0035] Generate an image configuration file according to the element information, the background image and the layer information of the target image to generate the target image. According to the second aspect of the embodiments of the present disclosure, there is provided an image processing apparatus, where the image processing apparatus includes:
[0036] An acquisition module, which is configured to acquire an image to be processed;
[0037] An identification module, which is configured to identify the element information corresponding to each element in the image to be processed;
[0038] A first determination module, which is configured to determine the background image in the image to be processed;
[0039] A generation module, which is configured to generate a target image according to the element information and the background image.
[0040] In some embodiments of the present disclosure, the elements in the image to be processed include picture elements and / or text elements;
[0041] The identification module is further configured to: sequentially identify the picture element information corresponding to the picture elements and the text element information corresponding to the text elements.
[0042] In some embodiments of the present disclosure, the identification module is further configured to: when the image to be processed includes picture elements, perform masking processing on the picture elements before identifying the text element information corresponding to the text elements.
[0043] In some embodiments of the present disclosure, the element information includes the position area where the element is located.
[0044] In some embodiments of the present disclosure, the identification module is further configured to: determine the position area where the initial text element is located;
[0045] Determine the proportion of the overlapping part between the position area where the initial text element is located and the position area where the picture element is located;
[0046] Determine the initial text element with the proportion less than or equal to a preset threshold as the text element;
[0047] Identify the text information in the text element information corresponding to the text element.
[0048] In some embodiments of the present disclosure, the text information includes text content and text format information.
[0049] In some embodiments of the present disclosure, the element information further includes layer order information of the layer where the element is located; the image processing device further includes a second determination module, and the second determination module is configured to determine the layer order of each layer where the element is located according to the position area where the element is located.
[0050] In some embodiments of the present disclosure, if the position areas where each element is located do not overlap with each other, the second determination module is further configured to: randomly generate the layer order of the layers where each element is located.
[0051] In some embodiments of the present disclosure, if there is an overlap between the position areas where any plurality of elements are located, the second determination module is further configured to:
[0052] Obtain a depth of field map corresponding to the image to be processed;
[0053] Calculate the average gray value of the position area where each element is located in the depth of field map;
[0054] Based on the average gray value, determine the layer order.
[0055] In some embodiments of the present disclosure, the first determination module is further configured to:
[0056] Remove the picture element and the text element from the image to be processed to obtain a background image to be complemented;
[0057] Perform image redrawing processing or image complementing processing on the background image to be complemented to obtain a background image.
[0058] In some embodiments of the present disclosure, the generation module is further configured to:
[0059] Determine the layer information of the target image according to the layer order of each layer where the element is located and the background image;
[0060] Generate an image configuration file according to the element information, the background image, and the layer information of the target image to generate the target image.
[0061] In some embodiments of the present disclosure, the picture element includes one or a combination of a commodity element, an identification element, and a control element.
[0062] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, and the electronic device includes:
[0063] Processor;
[0064] A memory for storing processor-executable instructions;
[0065] Wherein, the processor is configured to execute the image processing method described in the first aspect.
[0066] According to a fourth aspect of the embodiments of the present disclosure, a non-transitory computer-readable storage medium is provided. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the image processing method described in the first aspect.
[0067] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: By identifying the element information corresponding to each element in the image to be processed and determining the background image in the image to be processed, a target image can be generated based on the element information and the background image, achieving accurate identification and acquisition of different elements, as well as the combination of each element and the background image in the target image, facilitating subsequent separate editing and processing of different elements, improving the image processing efficiency and reducing the image processing difficulty.
[0068] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.
[0070] Figure 1 is a flowchart of an image processing method shown according to an exemplary embodiment.
[0071] Figure 2 is a schematic diagram of an image to be processed shown according to an exemplary embodiment.
[0072] Figure 3 is a flowchart of identifying text element information corresponding to a text element shown according to an exemplary embodiment.
[0073] Figure 4 is a flowchart of determining the layer order of each element's layer according to the position area where the element is located shown according to an exemplary embodiment.
[0074] Figure 5 is a schematic diagram of an image to be processed shown according to another exemplary embodiment.
[0075] Figure 6 is a schematic diagram of a depth-of-field map shown according to an exemplary embodiment.
[0076] Figure 7 is a flowchart for determining a background image in an image to be processed shown according to another exemplary embodiment.
[0077] Figure 8 is a flowchart for generating a target image according to element information and a background image shown according to an exemplary embodiment.
[0078] Figure 9 is a flowchart for an image processing method shown according to another exemplary embodiment.
[0079] Figure 10 is a flowchart for identifying text element information corresponding to a text element shown according to another exemplary embodiment.
[0080] Figure 11 is a flowchart for determining a layer order of each element shown according to another exemplary embodiment.
[0081] Figure 12 is a block diagram of an image processing apparatus shown according to an exemplary embodiment.
[0082] Figure 13 is a block diagram of an electronic device shown according to an exemplary embodiment.
[0083] In the figure:
[0084] 11 - Product picture; 12 - Function button; 13 - Platform logo; 14 - Copywriting; 15 - Title; 20 - Acquisition module; 30 - Recognition module; 40 - First determination module; 50 - Generation module; 101 - Processing component; 102 - Memory; 103 - Power component; 104 - Multimedia component; 105 - Audio component; 106 - Input / output interface; 107 - Sensor component; 108 - Communication component; 109 - Processor. Detailed implementation manners
[0085] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0086] In recent years, with the increasing diversification of layout design and the rapid development of image editing technology, images, as planar materials, have gradually become an important part of various commercial platforms and applications. The image materials can include elements such as product names, product prices, product pictures, function buttons, platform logos, etc. During the process of putting the image materials into use, it is usually necessary to perform secondary editing and processing on the elements in the image materials, such as changing the product names, product prices, and product pictures in the image materials, so as to realize the secondary utilization of the image materials.
[0087] However, for single-layer image files in formats such as JPG and PNG, when converting to the dedicated format (PSD format) of image processing software such as Photoshop, even if the PSD format supports multi-layer images, it is impossible to convert the originally single-layer image into a multi-layer image. When using related technologies to perform secondary editing and processing on single-layer image materials, it is difficult to accurately identify and obtain elements such as text, product pictures, platform logos, etc. in the image materials and edit each element separately, resulting in low image processing efficiency and great image processing difficulty.
[0088] Based on this, the exemplary embodiments of the present disclosure provide an image processing method. By identifying the element information corresponding to each element in the image to be processed and determining the background image of the image to be processed, a target image can be generated according to the element information and the background image, realizing the accurate identification and acquisition of different elements, as well as the combination of each element and the background image in the target image, facilitating subsequent separate editing and processing of different elements, improving the image processing efficiency and reducing the image processing difficulty.
[0089] In an exemplary embodiment, an image processing method is provided. Referring to Figure 1 as shown, the image processing method includes:
[0090] S100. Obtain the image to be processed.
[0091] In step S100, the image to be processed can be a planar material image that is about to be put into use or has already been put into use. The image to be processed can be obtained, for example, by downloading, copying, or intercepting.
[0092] S200. Identify the element information corresponding to each element in the image to be processed.
[0093] In step 200, the image to be processed includes at least one type of element. When there are multiple types of elements in the image to be processed, it is necessary to separately identify the element information corresponding to each element, so as to divide different types of elements in the image to be processed according to the element information corresponding to each element, providing a basis for generating the target image subsequently. When there is only one type of element in the image to be processed, due to the need to edit the element independently from the background, it is also necessary to identify the element information corresponding to the element, so that the layer corresponding to the element and the layer corresponding to the background image can be generated separately subsequently.
[0094] The element information corresponding to each element may include, for example, information on the position area corresponding to each element and / or information required to generate the layer corresponding to each element. For different elements, different detection methods or detection models can be used to extract and determine the element information. Exemplarily, for example, the element information corresponding to the text element can be determined through a text detection model, and the element information corresponding to the picture element can be determined through a picture recognition model and a picture segmentation model.
[0095] S300. Determine the background image in the image to be processed;
[0096] In step S300, it can be understood that in addition to each element, the image to be processed also includes the background of the image to be processed. However, the background of the image to be processed is not complete because it is covered by each element. If the elements are modified and moved during the subsequent editing process of the target image, it will cause the background to be missing. Before generating the target image, a complete background image is also required. The determination of the background image can be achieved, for example, through image redrawing or image completion technology.
[0097] S400. Generate a target image based on the element information and the background image.
[0098] In step S300, after identifying the element information corresponding to each element and determining the background image, a target image can be generated according to the element information corresponding to each element and the background image. The target image can be an image in which each element can be edited.
[0099] After generating the target image, subsequent independent editing of each element can be achieved by editing the target image. For example, when wanting to modify the product name in the picture material, the text element in the target image can be edited separately. Using existing image processing software, the product name can be modified by editing the text. Compared with directly processing the image to be processed in the related art, processing the target image is more convenient and fast and can be reused multiple times.
[0100] In this embodiment, by identifying the element information corresponding to each element in the image to be processed and determining the background image in the image to be processed, a target image can be generated based on the element information and the background image, achieving accurate identification and acquisition of different elements, as well as the combination of each element and the background image in the target image, facilitating subsequent separate editing of different elements, improving the image processing efficiency and reducing the image processing difficulty.
[0101] In some embodiments, the elements in the image to be processed include picture elements, or the elements in the image to be processed include text elements, or alternatively, the elements in the image to be processed include picture elements and text elements.
[0102] Exemplarily, referring to Figure 2 As shown, for the image to be processed as a planar material, the product picture 11, function button 12, platform logo 13, etc. in the image are all picture elements. The image to be processed may include multiple picture elements, and each picture element can be edited when generating the target image. When editing the target image subsequently, picture editing can be performed on the picture elements to achieve separate editing of the picture elements. The copywriting 14 such as the product name, product price, promotional message, etc. and the title 15 in the image are all text elements. Each text element can be edited when generating the target image. When editing the target image subsequently, text editing can be directly performed on the text elements to achieve separate editing of the text elements.
[0103] Identifying the element information corresponding to each element in the image to be processed includes: sequentially identifying the picture element information corresponding to the picture elements and the text element information corresponding to the text elements.
[0104] In an exemplary embodiment of the present disclosure, the element information corresponding to each element in the image to be processed can be detected and recognized in any order as needed. For example, in an exemplary embodiment, the element information corresponding to each element in the image to be processed can be determined by first detecting and recognizing the picture element information corresponding to the picture element, and then detecting and recognizing the text element information corresponding to the text element. It can be understood that for picture elements such as product picture 11, since there may be text printed on the surface of the product in product picture 11, the picture element will also carry text. Generally, the text printed on the surface of the product is not allowed or does not need to be edited and changed, and the text printed on the surface of the product needs to be recognized as an integral picture element together with the product itself. If the picture element and the text element are recognized synchronously or the text element is recognized first, the original text carried in the picture element will be recognized as the text element, which will interfere with the determination of the text element and the element information corresponding to the picture element, and affect the generation of the target image. Therefore, after first recognizing the picture element information corresponding to the picture element, and then recognizing the text element information corresponding to the text element, it is possible to avoid recognizing the text in the picture element as the text element again on the basis of determining the picture element.
[0105] Exemplarily, the text element information corresponding to the text element can be recognized by combining the Character Region Awareness For Text Detection (CRAFT) model with the Optical Character Recognition (OCR) model. The Character Region Awareness For Text Detection model can perform image segmentation at the character scale to effectively detect the text region. The Optical Character Recognition model can select the open-source PaddleOCR_v4.0 model, and the PaddleOCR_v4.0 model can recognize the text in multiple languages. For example, the position area where the text element is located can be recognized by the Character Region Awareness For Text Detection model, and then the text in the area can be recognized by the PaddleOCR_v4.0 model to recognize the text element information corresponding to the text element.
[0106] When using the method of combining the CRAFT model with the PaddleOCR_v4.0 model to recognize the text element information, the processing speed of nearly 2 seconds and 1.5 seconds can be achieved respectively in the Central Processing Unit (CPU) and Graphics Processing Unit (GPU) environments, and the authenticity of the detection results can be guaranteed.
[0107] In this embodiment, the elements in the image to be processed include picture elements and text elements. When identifying the element information corresponding to each element, first identify the picture element information corresponding to the picture elements, and then identify the text element information corresponding to the text elements, which can avoid identifying the text that is not allowed to be edited or does not need to be edited in the picture elements as text elements, prevent the misidentification of text elements from interfering with the identification of element information, and ensure the accuracy of the subsequent generated target image.
[0108] In some embodiments, the picture elements include one or a combination of multiple types such as product elements, logo elements, and control elements.
[0109] As mentioned above, for the image to be processed as a flat material, the product picture 11, function button 12, platform logo 13, etc. in the image can all be used as picture elements. Different picture elements have different detection and recognition difficulties and editing requirements, so the picture elements are further divided. Exemplarily, the product picture 11 is a product element among the picture elements, the function button 12 is a control element among the picture elements, and the platform logo 13 is a logo element among the picture elements.
[0110] Exemplarily, the picture element information corresponding to the picture elements of the logo element and control element types can be identified through the open-source YOLOv8 (You Only Look Once version 8) detection model. The YOLOv8 detection model supports image classification, object detection, and instance segmentation tasks, and can identify the categories and positions of objects in the image. For example, the YOLOv8 detection model can be trained using a sufficient number of sample pictures with various logo elements and control elements, so that the YOLOv8 detection model can identify the logo elements and control elements in the image to be processed, in order to identify the picture element information corresponding to the logo elements and control elements.
[0111] The commodity elements are different from the identification elements and control elements. Restricted by the shape of the commodity itself in the commodity elements, the boundaries of the commodity elements are generally irregular, and their recognition difficulty is higher than that of the identification elements and control elements. For the picture element information corresponding to the picture elements of the commodity element type, it can be recognized by the method of combining the Contrastive Language Image Pre-training (CLIP) model and the Salient Object Detection (SOD) model through contrastive learning. The DenseCLIP model can be selected as the contrastive learning multi-modal model. The DenseCLIP model is used for image segmentation and object detection, and can better utilize natural language supervision for high-quality visual representation learning. The U2Net model can be selected as the salient object detection model. The U2Net model adopts a nested basic structure and is used for salient object detection and semantic segmentation. For example, the commodity in the image to be processed can be recognized through the DenseCLIP model, and then the commodity picture 11 can be located and segmented through the U2Net model to recognize the picture element information corresponding to the commodity element.
[0112] In this embodiment, the picture elements include a combination of one or more of the commodity elements, identification elements, and control elements, realizing the reclassification of the picture elements. Different methods can be used to recognize the corresponding picture elements according to different recognition difficulties, improving the accuracy of recognizing the picture element information. In addition, the commodity elements, identification elements, and control elements can respectively correspond to different layers in the target image, enabling separate editing of different picture elements when editing the target image subsequently.
[0113] In some embodiments, the image processing method further includes: when the image to be processed includes picture elements, performing a masking process on the picture elements before recognizing the text element information corresponding to the text elements.
[0114] When the image to be processed includes picture elements, in order to sequentially recognize the picture element information corresponding to the picture elements and the text element information corresponding to the text elements, and to avoid re-recognizing the text in the picture elements as text elements, it is necessary to perform a masking process on the picture elements before recognizing the text element information corresponding to the text elements.
[0115] Exemplarily, after identifying the picture element information corresponding to the picture element, the picture element in the image to be processed can be covered or masked according to the picture element information to obtain the image after the masking process. At this time, the picture element no longer participates in the subsequent detection and recognition process of the text element information. The text element information can then be identified in the image after the masking process through the above-described method for identifying text element information. Since the image after the masking process no longer contains picture elements, the text originally carried in the picture elements cannot be recognized, which can avoid recognizing the text in the picture elements as text elements again.
[0116] In this embodiment, when the image to be processed includes picture elements, by performing a masking process on the picture elements before identifying the text element information corresponding to the text elements, it can be ensured that the identified picture elements no longer participate in the subsequent detection and recognition process. The text originally carried in the picture elements cannot be recognized, avoiding recognizing the text that is not allowed to be edited or does not need to be edited in the picture elements as text elements, preventing the misidentification of text elements from interfering with the determination of element information, and ensuring the accuracy of the subsequent generated target image.
[0117] In some embodiments, the element information includes the position area where the element is located.
[0118] The element information includes the position area where the element is located. When determining the element information corresponding to each element in the image to be processed, the position area where each element is located is determined, which is convenient for subsequent image segmentation of each element according to the position area where the element is located to accurately identify and obtain different elements in the image to be processed.
[0119] The position area where the element is located includes the shape and size of the area. The position area where the text element is located can be, for example, a rectangle. The position areas where picture elements such as identification elements and control elements are located can be relatively regular shapes such as rectangles, circles, ellipses, etc. The position area where the picture element of the commodity element category is determined according to the actual shape of the commodity. Exemplarily, the position area where the text element is located can be characterized by means of plane coordinates. When the position area where the text element is located is a rectangle, the position area where the text element is located can be characterized by the plane coordinates of the four corner points of the rectangular area, or can also be characterized by the plane coordinates of one of the corner points and the length and width of the rectangular area.
[0120] In this embodiment, the element information includes the position area where the element is located. By determining the position area where the element is located, it is convenient for subsequent image segmentation of each element according to the position area where the element is located, providing a basis for generating the target image. In addition, after determining the position area where the picture element is located, the masking process of the picture element can be realized according to the position area where the picture element is located, providing a prerequisite for sequentially identifying the picture element information and the text element information.
[0121] As described above, in the process of sequentially identifying picture element information and text element information, the text in the picture element can be prevented from being recognized as a text element again by masking the picture element. In some other embodiments, referring to Figure 3 as shown, the text element information corresponding to the text element is recognized, including:
[0122] S210. Determine the position area where the initial text element is located.
[0123] In step S210, the position area where the initial text element is located can be recognized by the CRAFT model as described above. Since the picture element has not been masked in advance, the text in the picture element and other texts are recognized as the initial text element together, that is, the position area where the initial text element is located may include the position area where the picture element is located.
[0124] S220. Determine the proportion of the overlapping part between the position area where the initial text element is located and the position area where the picture element is located.
[0125] In step S220, the position area where the picture element is located is included in the picture element information. For each initial text element, the proportion of the overlapping part between the position area where the initial text element is located and the position area where the picture element is located in the position area where the initial text element is located can be determined.
[0126] S230. Determine the initial text elements with a proportion less than or equal to the preset threshold as text elements.
[0127] In step S230, for any initial text element, if the corresponding proportion is too large, it means that the position area where the initial text element is located highly overlaps with the position area where the picture element is located, and the text in the position area where the initial text element is located is very likely to be the text originally carried by the picture element. Therefore, this initial text element is actually part of the picture element and cannot be used as the final text element. It is necessary to determine the initial text elements with a proportion less than or equal to the preset threshold as text elements.
[0128] S240. Recognize the text information in the text element information corresponding to the text element.
[0129] In step S240, after the selection of the text elements is completed, the text information of the text elements can be recognized by the PaddleOCR_v4.0 model as described above. Since the position areas where each text element is located have been determined in step S210, the recognition of the text element information corresponding to the text element is achieved after the text information is recognized.
[0130] In this embodiment, by determining the proportion of the overlapping part between the position area where the initial text element is located and the position area where the picture element is located, and comparing the proportion with a preset threshold, the final text element can be selected from the initial text elements, realizing the determination of the text element and the recognition of the text element information. Since the initial text elements with too large a proportion are removed, it is avoided that the text that is not allowed to be edited or does not need to be edited in the picture element is recognized as a text element, preventing the wrong recognition of the text element from interfering with the determination of the element information and ensuring the accuracy of the subsequent generated target image.
[0131] In some embodiments, the text information includes text content and text format information.
[0132] After determining the initial text elements with a proportion less than or equal to the preset threshold as text elements, the text information within the position area where the text elements are located can be obtained. The text information includes text content information representing the text content and text format information representing the text format. Exemplarily, the text format information may include relevant information on text formats such as font style, font size, font color, whether it is bold, whether it is italic, the number of text lines, line spacing, and character spacing.
[0133] In this embodiment, the text information includes text content and text format information. By determining the text information in the text element information corresponding to the text element, the text content and text format in the image to be processed can be restored as much as possible in the target image according to the text information, which is beneficial to the subsequent editing process of the text element.
[0134] In some embodiments, the element information further includes the layer order information of the layer where the element is located; the image processing method further includes: determining the layer order of each element according to the position area where the element is located.
[0135] In an exemplary embodiment, each element and the background image can be identified and marked according to the corresponding layer order, so as to generate a target image according to the position relationship between the layers subsequently. Exemplarily, the image to be processed may be a single-layer image in formats such as JPG and PNG, and the image to be processed cannot be directly converted into a multi-layer image through format conversion. The element information includes the layer information corresponding to the element, and the layer information may include the layer order information of the layer where the element is located, so as to be able to represent the position relationship between the layers corresponding to each element and be used to generate a target image according to the position relationship between the layers subsequently.
[0136] Multiple layers of the target image, for example, may include layers corresponding to each element and a layer of the background image. Exemplarily, for example, layers corresponding to each element may be generated respectively according to the element information, and the layers corresponding to each element and the layer of the background image are combined into the target image according to the layer information in the element information. When generating the target image, each element can have a corresponding layer, and when the target image is edited subsequently, any layer corresponding to an element can be edited to achieve separate editing of the element.
[0137] Since there may be partial overlaps between the position areas where each element is located, when generating the target image, it is necessary to clarify the covering relationship between the layers where each element is located to ensure that the overall visual effect of the multi-layer target image can be consistent with the image to be processed. The element information includes the layer order information of the layer where the element is located, so that the element information can represent the position relationship between the layers corresponding to each element. When determining the element information corresponding to each element, it is necessary to determine the layer order of the layer where each element is located according to the position area where the element is located. When generating the target image subsequently, the layers where each element is located are arranged according to the determined layer order.
[0138] In this embodiment, the element information includes the layer order information of the layer where the element is located, and the layer order of the layer where each element is located is determined according to the position area where the element is located, realizing the determination of the layer order information. The determination of the layer order information clarifies the position relationship between the layers where each element is located, enables the layers where each element is located to generate a multi-layer target image according to the layer order, ensures that the overall visual effect of the target image can be consistent with the image to be processed, provides a basis for the arrangement of the layers where each element is located, and improves the accuracy of generating the target image.
[0139] In some embodiments, if there are no overlaps between the position areas where each element is located, determining the layer order of the layer where each element is located according to the position area where the element is located includes: randomly generating the layer order of the layer where each element is located.
[0140] If there are no overlaps between the position areas where the elements are located, it means that no matter how the layers where each element is located are arranged, the position areas where each element is located in the target image will not be covered by the position areas where other elements are located, so that each element can completely appear in the target image. Therefore, the layer order of the layer where each element is located can be randomly generated, and the layer order of the background image is set at the bottommost layer. Exemplarily, if there are no overlaps between the position areas of each picture element and text element, the layer order can be set as the layer where the picture element is located - the layer where the text element is located - the background image, or it can also be set as the layer where the text element is located - the layer where the picture element is located - the background image.
[0141] In this embodiment, if the position areas where each element is located do not overlap with each other, the layer order can be determined by randomly generating the layer order of the layers where each element is located. When the position areas where each element is located in the target image will not be covered by the position areas where other elements are located, that is, each element can completely appear in the target image, the process of determining the layer order is simplified.
[0142] In some other embodiments, if there is an overlap between the position areas where any number of elements are located, it means that the layer order representing the layers where each element is located will determine the covering relationship of the position areas where the overlapping elements are located, that is, whether the element can completely appear in the target image. Refer to Figure 4 As shown, according to the position area where the element is located, determining the layer order of the layers where each element is located includes:
[0143] S500. Obtain the depth map corresponding to the image to be processed.
[0144] In step S500, the depth value of each pixel point in the image to be processed can be estimated through a depth estimation algorithm to obtain the depth map corresponding to the image to be processed. Exemplarily, for the image to be processed as shown in Figure 5 As shown, for example, the DPT-Hybrid depth estimation algorithm can be used to obtain the depth map corresponding to the image to be processed as shown in Figure 6 As shown. The DPT-Hybrid depth estimation algorithm is a deep learning method that combines direct prediction and a hybrid network, and can utilize the advantages of direct prediction and the hybrid network to improve the accuracy and robustness of depth estimation.
[0145] S600. Calculate the average gray value of the position area where each element is located in the depth map.
[0146] In step S600, the position areas where each element is located have been determined in the previous steps. The average gray value of the position area where each element is located in the depth map can be calculated. As mentioned above, the position area where the text element is located can be, for example, a rectangle, and the position areas where the picture elements of the identification element and the control element class can be relatively regular shapes such as rectangles, circles, ellipses, etc. The position area where the picture element of the commodity element class is determined according to the actual shape of the commodity. The gray values of the pixel points within the position area where each element is located can be averaged to obtain the average gray value of the position area where each element is located in the depth map.
[0147] S700. Determine the layer order based on the average gray value.
[0148] In step S700, the larger the average gray value of the position area where the element is located in the depth map, the closer the position of the element is to the top in the image to be processed. When the position area where the element with a larger average gray value overlaps with the position area where the element with a smaller average gray value, the element with a larger average gray value covers the element with a smaller average gray value in the image to be processed. Therefore, the layer order of each element can be determined according to the average gray value. The larger the average gray value of the element, the higher the layer order of the layer where the element is located is determined.
[0149] Exemplarily, in the depth of field map as Figure 6 shown, the average gray value of the position area where the commodity element is located is 232, the average gray value of the position area where the text element is located is 222, and the average gray value of the position area where the logo element is located is 42. Then the layer order can be determined as the layer where the commodity element is located - the layer where the text element is located - the layer where the logo element is located - the background image. When generating the target image, sort the layers where each element is located according to this layer order, so that the commodity element is located on top of the text element, and the commodity element that partially overlaps with the text element can appear completely in the target image, thus being consistent with the image to be processed.
[0150] In this embodiment, by obtaining the depth of field map corresponding to the image to be processed and calculating the average gray value of the position area where each element is located in the depth of field map, the layer order can be determined according to the average gray value. Determining the layer order according to the average gray value ensures that the overall visual effect of the target image can be consistent with the image to be processed, provides a basis for the arrangement of the layers where each element is located, and improves the accuracy of generating the target image.
[0151] In some embodiments, as shown in Figure 7 , determining the background image in the image to be processed includes:
[0152] S310. Remove the picture elements and text elements in the image to be processed to obtain the background image to be completed.
[0153] In step S310, after determining the element information corresponding to each element, the picture elements and text elements in the image to be processed can be removed by means such as masking or positioning and erasing to obtain the partially missing background image to be completed. Since both the picture elements and text elements in the image to be processed are removed at the same time, even if the position areas where the picture elements and text elements are located overlap locally, it will not affect the background image to be completed.
[0154] S320. Perform image redrawing processing or image completion processing on the background image to be completed to obtain the background image.
[0155] In step S320, the obtained background image to be completed is subjected to image redrawing processing or image completion processing to obtain a complete background image. When generating the target image, the background image can be used as one of the layers of the target image, ensuring the integrity of the background image even when subsequent editing is performed on other layers in the target image.
[0156] Exemplarily, for example, the background image to be completed can be AI redrawn by combining an artificial intelligence image generation model with a neural network model. The artificial intelligence image generation model can be the SD (Stable Diffusion) model, and the neural network model can be the ControlNet model. As an open-source plugin, the ControlNet model can learn texture structures and control the SD model to first compress the image to be completed into the latent space and then perform diffusion in the latent space to achieve efficient image generation.
[0157] The background image to be completed can also be subjected to image completion processing through image completion technology. For a simple, solid-color background image to be completed, for example, the image can be completed by extracting the theme color. For a complex background image to be completed, for example, the LAMA (Large Mask Inpainting) model can be used to repair the area with a relatively small missing area in the background image to be completed through its feature extraction network, and then the SD model as described above can be used to repair the area with a relatively large missing area in the background image to be completed. Finally, the repaired images are integrated to obtain the final complete background image.
[0158] In this embodiment, by removing the picture elements and text elements in the image to be processed, the background image to be completed is obtained, and the background image to be completed is subjected to image redrawing processing or image completion processing, achieving the acquisition of a complete background image, enabling the background image to be used as one of the layers of the target image, and ensuring the integrity of the background image when subsequent editing is performed on other layers in the target image. In some embodiments, as shown in Figure 8 According to the element information and the background image, a target image is generated, including: S410. Determine the layer information of the target image according to the layer order of each element's layer and the background image.
[0159] In step S410, after determining the layer order of each element's layer and the background image, the layer information of the multi-layer target image can be determined according to the layer order and the background image. The layer information of the target image can characterize the positional relationship between the layer corresponding to each element and the layer corresponding to the background image, and is used to combine the layer where each element is located and the layer corresponding to the background image into the target image.
[0160] S420. Generate an image configuration file based on the element information, the background image, and the layer information of the target image to generate the target image.
[0161] In step S420, an image configuration file can be generated based on the element information, the background image, and the layer information of the target image, and a multi-layer target image can be drawn and generated by virtue of the image configuration file.
[0162] Exemplarily, a configuration file in json format can be generated according to the element information, the background image, and the layer information. The configuration file in json format can describe the element information such as the position area where the element is located and the layer information, as well as the background image. When generating the target image at the front end, it can be rendered through the configuration file in json format, and finally a multi-layer target image in PSD format is obtained.
[0163] In this embodiment, by determining the layer information and generating an image configuration file according to the element information, the background image, and the layer information, the output storage of the element information and the generation of the target image are realized, so that the output target image can be a multi-layer image, and the individual elements can be independently edited by separately editing each layer of the target image in the subsequent process, which improves the image processing efficiency and reduces the image processing difficulty.
[0164] In an exemplary embodiment, refer to Figure 9 as shown, a method for image processing is provided. The method for image processing includes:
[0165] S1. Obtain the image to be processed;
[0166] S2. Identify the picture element information corresponding to the picture elements;
[0167] S3. Identify the text element information corresponding to the text elements;
[0168] S4. Remove the position areas where the picture elements are located and the position areas where the text elements are located in the image to be processed to obtain the background image to be complemented;
[0169] S5. Perform image redrawing processing or image complementing processing on the background image to be complemented to obtain the background image;
[0170] S6. Determine the layer order of each element;
[0171] S7. Generate an image configuration file according to the element information, the background image, and the layer order;
[0172] S8. Generate a multi-layer target image according to the image configuration file.
[0173] In this exemplary embodiment, after determining the element information of each element, the background image of the image to be processed is complemented, and then the layer order of each element is determined to generate a multi-layer target image. This does not limit the image processing sequence of the present disclosure. In other exemplary embodiments of the present disclosure, after determining the element information of each element, the layer order of each element can be determined, and then the background image of the image to be processed is complemented to generate a multi-layer target image.
[0174] In this embodiment, by identifying the element information corresponding to each element in the image to be processed and determining the background image in the image to be processed, a target image can be generated according to the element information and the background image, realizing the accurate identification and acquisition of different elements, as well as the combination of each element and the background image in the target image, facilitating subsequent separate editing of different elements, improving the image processing efficiency and reducing the image processing difficulty.
[0175] For the above steps S2 to S3, that is, sequentially identifying the picture elements corresponding to the picture elements and the text element information corresponding to the text elements, the picture elements can be masked after identifying the picture element information corresponding to the picture elements, and then the text element information corresponding to the text elements is identified. It is also possible, after identifying the picture element information, as Figure 10 shown, to perform the steps of identifying the text element information corresponding to the text elements in the following steps:
[0176] S31. Determine the position area where the initial text element is located;
[0177] S32. Determine the proportion of the overlapping part between the position area where the initial text element is located and the position area where the picture element is located;
[0178] S33. Determine the initial text elements with a proportion less than or equal to the preset threshold as text elements;
[0179] S34. Identify the text information in the text element information corresponding to the text elements.
[0180] In this embodiment, by determining the proportion of the overlapping part between the position area where the initial text element is located and the position area where the picture element is located and comparing the proportion with the preset threshold, the final text elements can be selected from the initial text elements, realizing the determination of the text elements and the identification of the text element information. Since the initial text elements with too large a proportion are removed, it is avoided that the text that is not allowed to be edited or does not need to be edited in the picture element is identified as a text element, preventing the wrong identification of the text element from interfering with the determination of the element information and ensuring the accuracy of the subsequent generated target image.
[0181] For the above step S6, that is, determining the layer order of each element, it can be asFigure 11 As shown in the figure, the following steps are carried out:
[0182] S61. Determine whether there is an overlap between the position areas where each element is located. If there is an overlap, execute S62; otherwise, execute S65.
[0183] S62. Obtain the depth map corresponding to the image to be processed.
[0184] S63. Calculate the average gray value of the position area where each element is located in the depth map.
[0185] S64. Determine the layer order based on the average gray value.
[0186] S65. Randomly generate the layer order of the layers where each element is located.
[0187] In this embodiment, if there is no overlap between the position areas where each element is located, the determination of the layer order can be achieved by randomly generating the layer order of the layers where each element is located. When the position areas where each element is located in the target image are not covered by the position areas where other elements are located, that is, each element can completely appear in the target image, the process of determining the layer order is simplified. If there is an overlap between the position areas where each element is located, the layer order is determined according to the average gray value, ensuring that the overall visual effect of the target image can be consistent with the image to be processed, providing a basis for the arrangement of the layers where each element is located, and improving the accuracy of generating the target image.
[0188] In an exemplary embodiment, as shown in the figure, an image processing apparatus is provided. The image processing apparatus includes an acquisition module 20, an identification module 30, a first determination module 40, and a generation module 50. The acquisition module 20 is used to acquire the image to be processed. The identification module 30 is used to identify the element information corresponding to each element in the image to be processed. The first determination module 40 is used to determine the background image in the image to be processed. The generation module 50 is used to generate a multi-layer target image according to the element information. Figure 12 In this embodiment, by using the identification module 30 to identify the element information corresponding to each element in the image to be processed acquired by the acquisition module 20, and using the first determination module 40 to determine the background image in the image to be processed, the generation module 50 can generate a target image according to the element information and the background image, realizing the accurate identification and acquisition of different elements, as well as the combination of each element and the background image in the target image, facilitating subsequent separate editing and processing of different elements, improving the image processing efficiency and reducing the image processing difficulty. In one embodiment, the elements in the image to be processed include picture elements and / or text elements, and the identification module 30 is further used to: sequentially identify the picture element information corresponding to the picture elements and the text element information corresponding to the text elements.
[0189]
[0190] In one embodiment, the recognition module 30 is further configured to: when the image to be processed includes picture elements, perform a masking process on the picture elements before recognizing the text element information corresponding to the text elements.
[0191] In one embodiment, the element information further includes the position area where the element is located.
[0192] In one embodiment, the recognition module 30 is further configured to: determine the position area where the initial text element is located; determine the proportion of the overlapping part between the position area where the initial text element is located and the position area where the picture element is located; determine the initial text element with a proportion less than or equal to a preset threshold as a text element; and recognize the text information in the text element information corresponding to the text element.
[0193] In one embodiment, the text information includes text content and text format information.
[0194] In some embodiments, the element information further includes the layer order information of the layer where the element is located, and the image processing device further includes a second determination module, which is configured to determine the layer order of each element according to the position area where the element is located.
[0195] In one embodiment, if there is no overlap between the position areas where the elements are located, the second determination module is further configured to: randomly generate the layer order of each element.
[0196] In one embodiment, if there is overlap between the position areas of any multiple elements, the second determination module is further configured to: obtain the depth map corresponding to the image to be processed; calculate the average gray value of the position areas where the elements are located in the depth map; and determine the layer order based on the average gray value.
[0197] In one embodiment, the first determination module 40 is further configured to: remove the picture elements and text elements in the image to be processed to obtain a background image to be filled; and perform image redrawing processing or image filling processing on the background image to be filled to obtain a background image.
[0198] In one embodiment, the generation module 50 is further configured to: determine the layer information of the target image according to the layer order of each element and the background image. Generate an image configuration file according to the element information, the background image, and the layer information of the target image to generate a target image.
[0199] In one embodiment, the picture elements include one or a combination of commodity elements, identification elements, and control elements.
[0200] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0201] In an exemplary embodiment, an electronic device is provided. Referring to Figure 13 as shown, the electronic device may include one or more of the following components: a processing component 101, a memory 102, a power component 103, a multimedia component 104, an audio component 105, an input / output (I / O) interface 106, a sensor component 107, and a communication component 108.
[0202] The processing component 101 generally controls the overall operation of the electronic device, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 101 may include one or more processors 109 to execute instructions to complete all or part of the steps of the above-described methods. In addition, the processing component 101 may include one or more modules to facilitate the interaction between the processing component 101 and other components. For example, the processing component 101 may include a multimedia module to facilitate the interaction between the multimedia component 104 and the processing component 101.
[0203] The memory 102 is configured to store various types of data to support the operation of the electronic device. Examples of such data include instructions for any application or method operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc. The memory 102 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0204] The power component 103 provides power to various components of the electronic device. The power component 103 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the electronic device.
[0205] The multimedia component 104 includes a screen that provides an output interface between the electronic device and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 104 includes a front camera and / or a rear camera. When the electronic device is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0206] The audio component 105 is configured to output and / or input audio signals. For example, the audio component 105 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 102 or transmitted via the communication component 108. In some embodiments, the audio component 105 further includes a speaker for outputting audio signals.
[0207] The I / O interface 106 provides an interface between the processing component 101 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0208] The sensor component 107 includes one or more sensors for providing an assessment of the various aspects of the status of the electronic device. For example, the sensor component 107 can detect the on / off state of the electronic device, the relative positioning of components, such as the display and the keypad of the electronic device. The sensor component 107 can also detect a change in the position of the electronic device or a component of the electronic device, the presence or absence of user contact with the electronic device, the orientation or acceleration / deceleration of the electronic device, and the temperature change of the electronic device. The sensor component 107 can include a proximity sensor that is configured to detect the presence of nearby objects without any physical contact. The sensor component 107 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 107 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0209] The communication component 108 is configured to facilitate communication between the electronic device and other devices in a wired or wireless manner. The device can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 108 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 108 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0210] In an exemplary embodiment, the electronic device can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described image processing method applied to the electronic device.
[0211] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as the memory 102 including instructions, is also provided. The above instructions can be executed by the processor 109 of the electronic device to complete the above-described image processing method applied to the electronic device. For example, the non-transitory computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc. When the instructions in the storage medium are executed by the processor 109 of the electronic device, the electronic device is enabled to execute the image processing method shown in the above embodiments.
[0212] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed in this disclosure. The specification and embodiments are only to be considered exemplary, and the true scope and spirit of the present invention are pointed out by the following claims.
[0213] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. An image processing method, characterized in that, the image processing method includes: obtaining an image to be processed; identifying the element information corresponding to each element in the image to be processed; determining the background image in the image to be processed; generating a target image according to the element information and the background image.
2. The image processing method according to claim 1, characterized in that, the elements in the image to be processed include picture elements and / or text elements; the identifying the element information corresponding to each element in the image to be processed includes: successively identifying the picture element information corresponding to the picture elements and the text element information corresponding to the text elements.
3. The image processing method according to claim 2, characterized in that, the image processing method further includes: when the image to be processed includes picture elements, before identifying the text element information corresponding to the text elements, performing a masking process on the picture elements.
4. The image processing method according to claim 2, characterized in that, the element information includes the position area where the element is located.
5. The image processing method according to claim 4, characterized in that, the identifying the text element information corresponding to the text elements includes: determining the position area where the initial text element is located; determining the proportion of the overlapping part between the position area where the initial text element is located and the position area where the picture elements are located; determining the initial text elements with the proportion less than or equal to a preset threshold as text elements; identifying the text information in the text element information corresponding to the text elements.
6. The image processing method according to claim 5, characterized in that, the text information includes text content and text format information.
7. The image processing method according to claim 4, characterized in that, the element information further includes the layer order information of the layer where the element is located; the image processing method further includes: determining the layer order of each layer where the elements are located according to the position area where the elements are located.
8. The image processing method according to claim 7, characterized in that, if there is no overlap between the position areas where each element is located, the determining the layer order of each layer where the elements are located according to the position area where the elements are located includes: randomly generating the layer order of each layer where the elements are located.
9. The image processing method according to claim 7, characterized in that, if there is overlap between the position areas of any plurality of elements, the determining the layer order of each layer where the elements are located according to the position area where the elements are located includes: obtaining the depth of field map corresponding to the image to be processed; calculating the average gray value of the position area where each element is located in the depth of field map; determining the layer order based on the average gray value.
10. The image processing method according to any one of claims 1-9, characterized in that, the determining the background image in the image to be processed includes: removing the picture elements and the text elements in the image to be processed to obtain a background image to be complemented; Perform image redrawing processing or image completion processing on the background image to be completed, to obtain a background image.
11. The image processing method according to claim 7, wherein, the generating a target image according to the element information and the background image includes: determining layer information of the target image according to the layer order of each layer where the elements are located and the background image; generating an image configuration file according to the element information, the background image and the layer information of the target image to generate the target image.
12. The image processing method according to claim 2, wherein, the picture elements include one or a combination of multiple of commodity elements, identification elements and control elements.
13. An image processing apparatus, wherein, the image processing apparatus includes: an acquisition module configured to acquire an image to be processed; an identification module configured to identify element information corresponding to each element in the image to be processed; a first determination module configured to determine a background image in the image to be processed; a generation module configured to generate a target image according to the element information and the background image.
14. An electronic device, wherein, the electronic device includes: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the image processing method according to any one of claims 1-12.
15. A non-transitory computer-readable storage medium, wherein, when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the image processing method according to any one of claims 1-12.