Image processing method and apparatus, electronic device, storage medium, and program product

By acquiring and analyzing the differences between the original object and its simulated object in the image, the appropriate parts are determined and combined, thus solving the problem of poor quality control in stylized image technology and achieving higher quality stylized transfer effects.

WO2026016624A1PCT designated stage Publication Date: 2026-01-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/095965
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-16
Filing Date
2025-05-20
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing stylized image technologies cannot ensure that stylized images manually combined by designers accurately reflect the content of the original images, resulting in poor quality control.

Method used

By acquiring the original object and its multiple simulated objects in the image, the difference information between the simulated object and the original object at the same location is determined. Based on the difference information, the matching parts for each part of the original object are determined from the multiple simulated objects, and finally the target simulated object is obtained by combining them.

Benefits of technology

It improves the quality of content stylization transfer, making the generated target simulated object closer to the original object, ensuring that its core features and style are preserved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025095965_22012026_PF_FP_ABST
    Figure CN2025095965_22012026_PF_FP_ABST
Patent Text Reader

Abstract

An image processing method, executed by an electronic device, and comprising: acquiring an original object in an image and a plurality of simulated objects corresponding to the original object, wherein the simulated objects are objects obtained by reproducing the original object by using reference elements, and arrangement modes of the reference elements in different simulated objects are different (101); determining difference information between the simulated objects and the original object at the same part (102); for corresponding parts at positions between the original object and the simulated objects, on the basis of the difference information, determining, from among the plurality of simulated objects, matched parts for the parts of the original object (103); and combining the matched parts determined for the parts of the original object to obtain a target simulated object (104).
Need to check novelty before this filing date? Find Prior Art

Description

Image processing methods, apparatuses, electronic devices, storage media, and program products

[0001] Related applications

[0002] This application claims priority to Chinese patent application filed on July 16, 2024, with application number 202410950917.1, entitled "Image Processing Method, Apparatus, Electronic Device, Storage Medium and Program Product", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of computers, and specifically to an image processing method, apparatus, electronic device, storage medium, and program product. Background Technology

[0004] Stylized images reinterpret the structure and texture of an original image by selecting, adding, and transforming basic shape elements, such as circles, rectangles, or triangles. This technique involves operations such as translating, rotating, and scaling these shape elements to reimagine and simplify the visual content of the original image. Through these transformations and combinations, a new visual style can be created, giving the image an artistic feel and visual effect distinctly different from the original.

[0005] Current stylized images are created by designers manually combining specific elements based on original images. However, due to differences in designers' technical skills and aesthetic preferences, it is difficult to ensure that stylized images accurately reflect the content of the original images, thus affecting the controllability of their quality. Summary of the Invention

[0006] This application provides an image processing method, apparatus, electronic device, storage medium, and program product that can improve the quality of content stylization transfer.

[0007] This application provides an image processing method, executed by an electronic device, including:

[0008] The original object in the image is obtained, as well as multiple simulated objects corresponding to the original object. The simulated objects are objects that reproduce the original object using reference elements. The arrangement of reference elements in different simulated objects is different.

[0009] Determine the differences between the simulated object and the original object at the same location;

[0010] For each part corresponding to the position between the original object and the simulated object, based on the difference information, the corresponding fitting parts are determined from multiple simulated objects for each part of the original object; and

[0011] The target simulation object is obtained by combining the adapted parts determined for each part of the original object.

[0012] This application also provides an image processing apparatus, including:

[0013] The acquisition unit is used to acquire the original object in the image and multiple simulated objects corresponding to the original object. The simulated object is the object obtained by reproducing the original object using reference elements. The arrangement of reference elements in different simulated objects is different.

[0014] The difference determination unit is used to determine the difference information between the simulated object and the original object at the same location;

[0015] The part determination unit is used to determine the appropriate parts for each part of the original object from multiple simulated objects based on the difference information, for each part corresponding to the position between the original object and the simulated object; and

[0016] The part combination unit is used to combine the adaptable parts determined for each part of the original object to obtain the target simulation object.

[0017] This application also provides an electronic device, including a processor and a memory, wherein the memory stores multiple instructions; the processor loads instructions from the memory to execute steps in any of the image processing methods provided in this application.

[0018] This application also provides a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute steps in any of the image processing methods provided in this application.

[0019] This application also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of any of the image processing methods provided in this application.

[0020] Details of one or more embodiments of this application are set forth in the following drawings and description. Other features, objects, and advantages of this application will become apparent from the specification, drawings, and claims. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the published drawings without creative effort.

[0022] Figure 1a is a schematic diagram of a scene of the image processing method provided in an embodiment of this application;

[0023] Figure 1b is a schematic flowchart of the image processing method provided in an embodiment of this application;

[0024] Figure 2a is a flowchart of the image processing method provided in this application being applied to a scene where the filling element is a geometric element;

[0025] Figure 2b is a simplified scene diagram provided in an embodiment of this application;

[0026] Figure 2c is a schematic diagram of image style transfer provided in an embodiment of this application;

[0027] Figure 2d is a schematic diagram of the scene for generating the target simulation object provided in an embodiment of this application;

[0028] Figure 2e is a schematic diagram of the fill position update scenario provided in the embodiment of this application;

[0029] Figure 2f is a schematic diagram of a scenario for generating and reconstructing simulated objects provided in an embodiment of this application;

[0030] Figure 2g is a schematic diagram of the transformation of a triangle into a right triangle according to an embodiment of this application;

[0031] Figure 2h is a schematic diagram of the triangle approximation optimization provided in the embodiment of this application;

[0032] Figure 2i is a schematic diagram of the effect of updating and reconstructing the simulated object provided in the embodiment of this application;

[0033] Figure 3 is a schematic diagram of the image processing apparatus provided in an embodiment of this application;

[0034] Figure 4 is a schematic diagram of the structure of the electronic device provided in an embodiment of this application. Detailed Implementation

[0035] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0036] This application provides an image processing method, apparatus, electronic device, storage medium, and program product.

[0037] Specifically, the image processing device can be integrated into an electronic device, such as a terminal or server. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, or personal computer (PC); the server can be a single server or a server cluster consisting of multiple servers.

[0038] In some embodiments, the image processing apparatus may also be integrated into multiple electronic devices, such as multiple servers, with the image processing method of this application being implemented by the multiple servers.

[0039] In some embodiments, the server may also be implemented as a terminal.

[0040] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0041] For example, referring to Figure 1a, the electronic device can acquire the original object in the image and multiple simulated objects corresponding to the original object. The simulated objects are objects obtained by reproducing the original object using reference elements, and the arrangement of the reference elements in different simulated objects is different. The device can determine the difference information between the simulated object and the original object at the same location. Based on the difference information, for each location corresponding to the position between the original object and the simulated object, the device can determine the matching parts for each part of the original object from multiple simulated objects. The device can combine the matching parts determined for each part of the original object to obtain the target simulated object.

[0042] Specifically, for each part corresponding to the position between the original object and the simulated object, based on the differences between the simulated object and the original object in the same part, suitable parts can be determined from multiple simulated objects for each part of the original object. This allows for the targeted retention of parts in each simulated object that are similar to the original object. The suitable parts determined for each part of the original object are then combined so that the final target simulated object is closer to the original object as a whole compared to multiple simulated objects, ensuring that the generated target simulated object accurately retains the core features and style of the original object. This improves the quality of content stylization transfer.

[0043] In this embodiment, a computer vision-based image processing method involving artificial intelligence is provided, as shown in Figure 1b. The specific process of this image processing method is as follows:

[0044] 101. Obtain the original object in the image, and multiple simulated objects corresponding to the original object. The simulated objects are objects obtained by reproducing the original object using reference elements. The arrangement of reference elements in different simulated objects is different.

[0045] The images are those awaiting style transfer. For example, the images can be raw, unprocessed images containing background and noise, or they can be images that have been segmented by the background, retaining only the main subject. The main subject refers to the specific entity that dominates the image, including people, objects, etc.

[0046] The original object refers to the object in the image content that requires style transfer. For example, the original object can be a person, object, anime character, etc. in the image content.

[0047] Reference elements are elements used to transfer the style of an image. For example, reference elements can be geometric elements, such as polygons, circles, and ellipses, or elements with complex structures, such as flowers, trees, and furniture.

[0048] A simulated object is an object obtained by reproducing the original object using reference elements. Different simulated objects corresponding to the original object have different arrangements of reference elements.

[0049] Multiple simulated objects can be obtained from different servers or from the same server. Here, the server refers to the server used to generate the simulated object corresponding to the original object.

[0050] In some embodiments, in order to reduce the influence of the background on the subject object during image style transfer, obtaining an image includes: obtaining an image to be processed; and performing original object extraction processing on the image to be processed to obtain an image.

[0051] Here, the image to be processed refers to an image that contains background content. For example, the image to be processed can be an unprocessed image, or an image that has been partially processed but still contains background content, and so on.

[0052] In some embodiments, the image may be an image obtained after the image to be processed has been processed by object separation methods such as region segmentation and foreground segmentation.

[0053] In some embodiments, in order to obtain simulated objects similar to the original object, before obtaining multiple simulated objects corresponding to the original object, the method further includes: obtaining an image and reference elements, wherein the image content of the image includes the original object; and generating multiple simulated objects based on the original object in the image and using the reference elements.

[0054] Image content refers to the information conveyed by the specific visual elements contained in an image. For example, image content can include a scene and the main objects within that scene.

[0055] The method for determining the original objects in an image can be selected based on the image's characteristics and processing requirements, or a combination of methods can be used. Specific methods are as follows:

[0056] 1) Determine the feature distribution of pixels in an image using a Gaussian distribution, such as color or texture information. The specific steps are as follows: First, calculate the color or texture feature value of each pixel in the image, and then construct a Gaussian distribution model based on these values. By setting an appropriate threshold, pixels are classified into categories belonging to the background or the original object, thus distinguishing between the background and the original object in the image, as well as different parts of the original object. Specifically, multiple experiments can be conducted to calculate the segmentation accuracy of the background and the original object at different thresholds, and the threshold with the highest accuracy can be selected as the appropriate threshold. For example, a range from 0.1 to 0.9 can be iterated with a step size of 0.1, calculating the segmentation accuracy at each threshold. This method is suitable for images with relatively obvious color or texture features.

[0057] 2) Edge detection algorithms can detect intensity changes in different regions of an image, thereby distinguishing between the background and the original object, as well as different parts of the original object. Taking the Canny edge detection algorithm as an example, the image is first Gaussian smoothed to reduce noise, then the gradient magnitude and direction are calculated, followed by non-maximum suppression to refine the edges, and finally, double thresholding and edge concatenation are used to determine the true edges. Based on the detected edges, the boundaries between the background and the original object, as well as the boundaries of different parts of the original object, can be defined. In double thresholding, the low and high thresholds can be set according to the characteristics of the image. Generally, the high threshold can be set to the median of the image gradient magnitude, and the low threshold can be set to 0.4 times the high threshold. For example, if the median of the image gradient magnitude is 50, then the high threshold is set to 50 and the low threshold is set to 20. This method is suitable for situations where the edges of objects in the image are relatively clear.

[0058] 3) Image semantic detection can obtain semantic information of different regions in an image, thereby identifying the region where the original object is located and the regions where different parts of the original object are located, thus identifying the various parts of the original object. Taking a deep learning-based semantic segmentation model (such as U-Net) as an example, the image is first input into the model. The model extracts image features through structures such as convolutional layers and pooling layers, then uses an upsampling layer to restore the feature map to the same size as the input image. Finally, each pixel is classified to obtain semantic labels for different regions, thereby identifying the region where the original object is located and the regions where different parts of the original object are located. The U-Net model can be trained using publicly available image segmentation datasets, such as Cityscapes and Pascal VOC. Regarding training parameters, the learning rate can be set to 0.001, the batch size to 16, and the number of training epochs to 50. This method is suitable for situations where high requirements are placed on the semantic information of objects in the image.

[0059] In practical applications, edge detection algorithms can be used to quickly determine the approximate boundaries, and then image semantic detection can be used to further accurately identify objects and their parts; or the Gaussian distribution method can be combined to perform preliminary classification of the image, and then other methods can be used for refinement.

[0060] In some embodiments, in order to regularly combine reference elements to obtain multiple simulated objects similar to the original object, the method further includes: generating a primary simulated object based on the original object in the image using reference elements; and adjusting the reference elements constituting the primary simulated object based on the original object in the image to obtain a secondary simulated object.

[0061] The initial simulation object is a simulation object that is similar to the original object, obtained by laying out multiple reference elements for the first time. The sizes of the multiple padding elements involved in the layout may be different.

[0062] A secondary simulation object is a simulation object obtained by adjusting the reference elements that constitute the primary simulation object based on the original object.

[0063] For example, a secondary simulated object can be obtained by adjusting parameters such as the layout position, size, or rotation angle of the reference elements that constitute the primary simulated object, based on the original object in the image.

[0064] Secondary simulated objects can also be obtained by adjusting methods such as adding reference elements, deleting reference elements, merging reference elements, and splitting reference elements.

[0065] In some embodiments, including but not limited to using genetic algorithms and particle swarm optimization algorithms, the reference elements constituting the initial simulated object are adjusted. Taking a genetic algorithm as an example, the position, size, shape, and other parameters of the reference elements are first encoded into chromosomes, and then an initial population is randomly generated. The fitness (e.g., similarity) between the simulated object corresponding to each individual (chromosome) and the original object is calculated. Based on the fitness, superior individuals are selected for crossover and mutation operations to generate a new population. The above process is repeated until a preset stopping condition is met, resulting in the optimal combination of reference element parameters, thereby adjusting the reference elements constituting the initial simulated object.

[0066] For example, using genetic algorithms and particle swarm optimization (PSO) algorithms, the reference elements constituting the primary simulated object can be adjusted—specifically, the position, size, and shape of these reference elements—to make the resulting secondary simulated object more similar in appearance to the original object compared to the primary one. Taking PSO as an example, a swarm of particles is first initialized, each particle representing a set of reference element parameters (position, size, shape, etc.). Each particle has its own velocity and position, and its fitness is evaluated based on a fitness function (such as the similarity between the simulated object and the original object). Particles update their velocity and position based on their historical best position and global best position, iterating continuously until the optimal combination of reference element parameters is found, thereby adjusting the reference elements constituting the primary simulated object. Particle swarm optimization is an algorithm used to adjust the reference elements constituting the primary simulated object. It first initializes a swarm of particles, each particle representing a set of reference element parameters (position, size, shape, etc.). Each particle has its own velocity and position, and its fitness is evaluated based on a fitness function (such as the similarity between the simulated object and the original object). Particles update their velocity and position based on their historical best position and global best position, iterating continuously until the optimal combination of reference element parameters is found.

[0067] In some embodiments, to avoid the infinite generation of simulation objects, the method further includes: obtaining a preset simulation stop condition; and stopping the adjustment processing of the reference elements constituting the primary simulation object according to the preset simulation stop condition.

[0068] Among them, the preset simulation stopping condition is a pre-set condition that controls the stopping of the generation of the simulation object. For example, the preset simulation stopping condition can be the iteration time, reaching a preset goal, etc. The preset goal can be a preset goal that the generated simulation object needs to achieve, specifically, the similarity between the final generated target simulation object and the original object meets the preset condition.

[0069] In some embodiments, in order to generate multiple simulated objects similar to the original object, so that a final simulated object that is more similar to the original object can be obtained through comparison and selection in the future, the reference elements constituting the primary simulated object are adjusted based on the original object in the image to obtain the secondary simulated object. This includes: generating a new arrangement based on the original object in the image and the arrangement of reference elements in the primary simulated object; and adjusting the reference elements constituting the primary simulated object according to the new arrangement to obtain the secondary simulated object.

[0070] The layout refers to the arrangement used by the reference elements to reproduce the original object. For example, the layout can include the position, size, and rotation angle of the reference elements in reproducing the original object.

[0071] The new layout is based on the original objects in the image, and the layout of the reference elements that make up the primary simulated objects is adjusted to obtain the new layout.

[0072] For example, based on the original object in the image, the position, size, rotation angle, etc. of the reference elements that constitute the primary simulated object are adjusted to obtain the secondary simulated object.

[0073] The purpose of the original object in the image is to ensure that reference elements do not exceed specific boundaries or ranges when constructing the simulated object.

[0074] In some embodiments, in order to combine multiple reference elements to reproduce the original object, a primary simulated object is generated based on the original object in the image using reference elements, including: determining the element layout positions on the original object in the image; and generating the primary simulated object based on the element layout positions and reference elements.

[0075] The following methods can be used to determine the layout position of elements:

[0076] Using Gaussian distribution: Calculate the color value or texture feature value of each pixel within the original object region of the image to construct a Gaussian distribution model. Set a probability threshold, and select pixels with a probability higher than this threshold as representative pixels for layout. For example, setting the probability threshold to 0.8 selects pixels with a probability greater than 0.8. Alternatively, the original object region can be divided into different layout regions based on the standard deviation of the Gaussian distribution; for example, regions with smaller standard deviations can be grouped into a single layout region.

[0077] Edge detection algorithms: Taking the Canny edge detection algorithm as an example, edge detection is performed on the original object region to obtain edge information. The curvature of each point on the edge is calculated, and points with larger curvatures are selected as key points for layout. For example, a curvature threshold of 0.5 is set, and points with curvature greater than 0.5 are selected as key points. Alternatively, the edge can be divided into different line segments based on its connectivity, and the area enclosed by each line segment is considered as a layout area.

[0078] Image semantic detection: Image semantic detection uses a deep learning-based semantic segmentation model (such as U-Net) to perform semantic segmentation on the original object, obtaining semantic labels for different parts. For each semantic label, the midpoint of its boundary is calculated as a layout point; or the four vertices of the bounding rectangle of the semantic label region are used as layout points. Alternatively, each semantic label region can be used as a layout region. Based on the determined element layout positions and reference elements, a preliminary simulated object is generated. If the element layout position is a layout point, the vertices, center points, or preset position points of the reference elements can be aligned with the layout point to arrange reference elements at each layout point, forming a preliminary simulated object similar to the original object. The preset position point can be a specific position of the preset reference element so that the reference element is aligned with the layout point, thereby realizing the layout of reference elements when reproducing the original object. If the element layout position is a layout region, reference elements are arranged in each layout region to reproduce the original object. In some embodiments, in order to enable the combination of reference elements to form a simulated object similar to the original object, a primary simulated object is generated based on the element layout position and the reference element, including: aligning the vertices of the reference element with the element layout position to generate the primary simulated object.

[0079] In this context, element layout position refers to the alignment position within the area where the original object is located in the image, guiding the combination of reference elements to form a preliminary simulated object. For example, element layout position can be a layout point or layout area that guides the combination of reference elements to form a preliminary simulated object.

[0080] In some embodiments, if the element layout position is based on the layout point obtained from the original object in the image, the vertex, center point or preset position point of the reference element can be aligned with the layout point to realize the arrangement of reference elements at each layout point to form a preliminary simulated object similar to the original object. The preset position point can be a specific position of the reference element set in advance so that the reference element is aligned with the layout point, thereby realizing the layout of reference elements when reproducing the original object.

[0081] In some embodiments, if the element layout position is a layout region obtained based on the original object in the image, reference elements are arranged in each layout region to reproduce the original object. In some embodiments, in order to enable the combination of reference elements to form a simulated object similar to the original object, a primary simulated object is generated based on the element layout position and the reference elements, including: aligning the vertices of the reference elements with the element layout position to generate the primary simulated object.

[0082] In this context, the vertex of a reference element is the point where two sides of the reference element intersect. For example, if the reference element is a polygon (such as a triangle, quadrilateral, pentagon, etc.), then the vertex is the point where two sides of the polygon intersect.

[0083] In some embodiments, Gaussian distribution, edge detection algorithms, and image semantic detection can be used to determine the element layout positions on the original object in the image.

[0084] For example, when the reference element is a polygon, first determine the element layout position on the original object in the image, then align the vertices of the reference element with the element layout position to obtain the reference element at the element layout position, thus realizing a simulated object that reproduces the original object.

[0085] In some embodiments, when the reference element is a polygon, the vertices of the reference element are aligned with the element layout position to generate a primary simulation object; when the reference element is a circle, the center of the circle is aligned with the element layout position; when the reference element is an ellipse, the center point of the ellipse is aligned with the element layout position, etc., to generate a primary simulation object.

[0086] In some embodiments, in order to enable the simulated object to have color information similar to the original object, so as to distinguish the difference between the simulated object and the original object based on the color of the simulated object, a primary simulated object is generated based on the element layout position and the reference element, including: obtaining the target color information of the original object at the element layout position; and performing color filling processing on the reference element at the element layout position according to the target color information to obtain the primary simulated object.

[0087] Target color information refers to the color characteristics of the original object in the image at its element layout location. For example, target color information could be the exact color, representative color, composite color, average color, etc. of the original object at its element layout location.

[0088] In some embodiments, the target color information of the original object at the element layout position can be obtained by sampling the pixel color at the element layout position and calculating the average value.

[0089] In some embodiments, in order to obtain a secondary simulated object similar to the primary simulated object, the primary simulated object includes multiple reference elements; based on the original object in the image, the reference elements constituting the primary simulated object are adjusted to obtain the secondary simulated object, including: based on the original object in the image, updating the element layout positions aligned to each reference element in the primary simulated object to obtain updated element layout positions; determining the reference elements at the updated element layout positions; and filling the reference elements at the updated element layout positions with color information from the original object, to obtain the secondary simulated object.

[0090] Among them, updating the element layout position refers to updating the positions that need to be aligned when multiple reference elements are formed into a preliminary simulated object, based on the original object in the image.

[0091] For example, if the reference element is a polygon, and the vertices of the polygon constituting the primary simulation object are aligned with the element layout position, within the range of the original object in the image, the vertices of the reference element can be aligned with the updated element layout position by adjusting the element layout position aligned with the polygon constituting the primary simulation object, thus obtaining the reference element at the updated element layout position. The shape obtained by combining multiple reference elements at the updated element layout positions is similar to the shape of the primary simulation object. Then, based on the color information of the original object at the updated element layout position, the reference element at the updated element layout position is color-filled to obtain the secondary simulation object.

[0092] In some embodiments, when the similarity between some parts of the primary simulated object and the original object is insufficient, in order to update some parts of the primary simulated object and retain parts similar to the original object, so that the subsequently obtained simulated object is more similar to the original object than the primary simulated object, a new arrangement is generated based on the original object in the image and the arrangement of reference elements in the primary simulated object. This includes: determining the parts of the primary simulated object to be adjusted; and adjusting the arrangement of reference elements of the parts of the primary simulated object to be adjusted based on the original object in the image to obtain a new arrangement.

[0093] The parts to be adjusted refer to the parts in the initial simulation object that do not match the original object.

[0094] For example, when adjusting the reference elements that make up the part to be adjusted, if the reference element is a polygon, the position of the vertices of the polygon can be adjusted within the range of the part to be adjusted. If the reference element is a circle, the radius, diameter, rotation angle, etc. of the circle can be adjusted. If the reference element is an ellipse, the center point, semi-axis length, rotation angle, etc. of the ellipse can be adjusted. If the reference element is a structurally complex element, the position, size, rotation angle, etc. of the structurally complex element can be adjusted, which is beneficial to obtaining a secondary simulation object that is closer to the original object.

[0095] In some embodiments, the difference information between the primary simulated object and the original object at the same location can be compared. If the difference information does not meet the preset difference conditions, the location is determined as the location to be adjusted.

[0096] 102. Determine the differences between the simulated object and the original object at the same location.

[0097] The same part refers to the area where the simulated object and the original object are in the same location.

[0098] Difference information indicates the differences between the simulated object and the original object. For example, difference information can be loss value, similarity, etc.

[0099] In some embodiments, the method further includes: dividing the simulated object and the original object into multiple parts; and determining the same part corresponding to the simulated object and the original object based on the position of each part in the simulated object and the position of each part in the original object.

[0100] The location of a part within the simulated object refers to its position within the simulated object. For example, the location of a part can be the area it occupies within the simulated object, its coordinate position within the simulated object, or the semantic meaning of the part obtained through semantic segmentation, such as left hand, right hand, or body part.

[0101] The position of a part in the original object refers to the position of that part within the original object. The method for determining the position of a part in the original object is the same as the method for determining the position of a part in the simulated object.

[0102] In some embodiments, the difference information includes loss values ​​or similarity between the simulated object and the original object in terms of shape, size, position, color, material, texture, etc., corresponding to the same location. For example, for shape differences, the Hausdorff distance between the contours of the simulated object and the original object at the same location can be calculated; for size differences, the ratio of their areas or volumes at the same location can be calculated; for position differences, the Euclidean distance between their centroid coordinates at the same location can be calculated; for color differences, the Bach distance between their color histograms at the same location can be calculated; for material and texture differences, texture feature extraction algorithms (such as LBP, GLCM) can be used to extract features, and then the distance between the features can be calculated.

[0103] In some embodiments, in order to analyze the differences between the simulated object and the original object at different locations, so as to update the parts of the simulated object that are not similar to the original object and retain the parts that are similar to the original object, and to determine the difference information between the simulated object and the original object at the same location, the following steps are taken: determining the target difference information between the primary simulated object and the original object at the same location; and determining the parts of the primary simulated object to be adjusted, including: determining the parts to be adjusted from the primary simulated object based on the target difference information.

[0104] The target difference information refers to the differences between the primary simulated object and the original object at the same location. In some embodiments, the target difference information can be determined by calculating the loss value or similarity of the primary simulated object and the original object at the same location in terms of shape, size, position, color, etc. For example, for shape target difference information, the Hausdorff distance between the contours of the primary simulated object and the original object at the same location is calculated; the larger the Hausdorff distance, the greater the shape difference. For size target difference information, the ratio of the area or volume of the two objects at the same location is calculated; the greater the ratio deviates from 1, the greater the size difference. For position target difference information, the Euclidean distance between the centroid coordinates of the two objects at the same location is calculated; the greater the distance, the greater the position difference. For color target difference information, the Barcol distance between the color histograms of the two objects at the same location is calculated; the greater the distance, the greater the color difference.

[0105] For example, if the target difference information reflects that the difference between the primary simulated object and the original object in the same part does not meet the preset difference conditions, then that part is designated as the part to be adjusted. The preset difference conditions are preset settings used to measure the difference between the primary simulated object and the original object in the same part.

[0106] 103. For each part corresponding to the position between the original object and the simulated object, based on the difference information, determine the matching part for each part of the original object from multiple simulated objects.

[0107] In one embodiment, step 103 includes: for each part of the original object, comparing the difference information (such as loss value or similarity) of this part in each simulated object, and selecting the corresponding part in the simulated object with the optimal difference information (the smallest loss value or the largest similarity) as the adapted part. For example, for part A of the original object, the difference information in simulated object 1 is loss value L1, and the difference information in simulated object 2 is loss value L2. If L1 < L2, then select part A in simulated object 1 as the adapted part.

[0108] Among them, the adapted part refers to the part in the simulated object that is similar to the original object. For example, multiple parts corresponding in position between the original object and the simulated object include part 1 and part 2, and the difference information between the simulated object and the original object includes difference information 1 corresponding to part 1 and difference information 2 corresponding to part 2. According to difference information 1 and difference information 2, determine the adapted part from part 1 and part 2 of the simulated object. Specifically, for part 1, compare the values of difference information 1 in each simulated object, and select part 1 in the simulated object with the optimal difference information 1 (the smallest loss value or the largest similarity) as the adapted part; for part 2, compare the values of difference information 2 in each simulated object, and select part 2 in the simulated object with the optimal difference information 2 (the smallest loss value or the largest similarity) as the adapted part.

[0109] In some embodiments, the multiple simulated objects may include primary simulated objects and secondary simulated objects, or may only include secondary simulated objects, and so on.

[0110] 104. Combine the adapted parts determined for each part of the original object to obtain a target simulated object.

[0111] Among them, the target simulated object is a simulated object obtained by combining the adapted parts that are similar to each part of the original object in multiple simulated objects.

[0112] In some embodiments, in order to ensure that the generated simulated object is similar to the original object and can also be adjusted according to the user's needs, it further includes: obtaining a specified element; determining a reference element according to the type corresponding to the specified element, and the subclass of the reference element includes the specified element; after combining the adapted parts determined for each part of the original object to obtain a target simulated object, it further includes: based on the specified element, performing an element conversion process on the reference elements constituting the target simulated object to obtain a new target simulated object, and the new target simulated object is an object obtained by reproducing the original object using the specified element.

[0113] Among them, the specified element is used to specify that the finally generated simulated object is composed of it, and the type it has corresponds to a specific reference element.

[0114] For example, when the specified element is a right-angled triangle, isosceles triangle, or obtuse triangle, its corresponding type is triangle, and the triangle is used as a reference element. When the specified element is a flower, grass, or tree, its corresponding type is a complex element, and the bounding box of the flower, grass, or tree, or a preset shape that can well summarize the overall outline of the flower, grass, or tree (such as an ellipse, polygon, etc.) is used as a reference element. When the specified element is a parallelogram, rectangle, square, etc., its corresponding type is quadrilateral, and the quadrilateral is used as a reference element, and so on.

[0115] A new target simulation object is a new object that is similar to the original object, obtained by transforming the reference elements that constitute the target simulation object into specified elements.

[0116] For example, when the reference element is a triangle and the specified element is a right triangle, the triangles that make up the target simulation object are converted into right triangles, resulting in an object composed of multiple right triangles that looks similar to the original object.

[0117] In some embodiments, the conversion methods for different types of elements are as follows:

[0118] Triangle Transformation: Taking the transformation from a triangle to a right triangle as an example, the initial simulated object formed by the triangulation algorithm is composed of multiple arbitrary triangles, which need to be post-processed to transform into right triangles or isosceles triangles, etc. By definition, a right triangle can be obtained by drawing a perpendicular line to any triangle. For an obtuse triangle, the perpendicular lines drawn from the two acute angle vertices are not inside the triangle; therefore, drawing a perpendicular line from the vertex of the largest angle can divide the arbitrary triangle into two right triangles, obtaining the specified elements that meet the requirements. To transform a triangle into an isosceles triangle, any shorter side can be padded to be the same length as the longer side. The triangulation algorithm is an algorithm that constructs multiple non-intersecting triangular meshes based on the element layout positions (layout points) on the original object in the image. That is, the vertices of the triangles are aligned with the layout points on the original object in the image, realizing the reproduction of the original object through reference elements to obtain the simulated object.

[0119] Circle Conversion: If the reference element is a circle and the specified element is an ellipse, the circle's radius can be appropriately scaled based on its radius and position information, giving it different semi-axis lengths in different directions, thus converting it into an ellipse. Simultaneously, the center position remains unchanged.

[0120] Conversion of Structurally Complex Elements: When the specified element is a structurally complex element such as flowers, plants, or trees, and the reference element is its bounding box or a preset shape (such as an ellipse or polygon), it can be mapped to the corresponding flower, plant, or tree element based on the position and size information of the bounding box or preset shape. For example, if the reference element is an elliptical bounding box, the image of the flowers, plants, or trees can be cropped and deformed to match the elliptical bounding box. After setting preset simulation stopping conditions (such as iteration time), a primary simulated object, similar to the original object and formed by combining reference elements, is created. For cases where some specified elements differ from the reference elements, the reference elements constituting the target simulated object are converted into specified elements using the above method.

[0121] In some embodiments, considering that there may be redundant specified elements after the target simulation object is converted into a new target simulation object, in order to optimize the new target simulation object, after performing element transformation processing on the reference elements constituting the target simulation object based on the specified elements to obtain the new target simulation object, the method further includes: performing element overlap detection on the specified elements constituting the new target simulation object to obtain elements to be expanded and elements to be deleted, wherein the elements to be expanded and elements to be deleted overlap; performing expansion processing on the elements to be expanded in the new target simulation object according to the elements to be deleted to obtain expanded elements, wherein the expanded elements include the elements to be deleted; and deleting the elements to be deleted in the new target simulation object according to the expanded elements to obtain the updated new target simulation object.

[0122] Among them, the element to be expanded refers to the specified element among the overlapping elements that needs to be expanded to cover the overlapping area.

[0123] The expansion process can be achieved by translating the boundary of the element to be expanded towards the direction of the element to be deleted. The translation distance can be determined based on the size of the overlapping area; for example, translating the boundary of the element to be expanded towards the direction of the element to be deleted by half the width of the overlapping area.

[0124] The element to be deleted refers to the specified element that needs to be deleted among the overlapping specified elements.

[0125] For example, the element to be expanded can be the specified element on top of the overlapping specified elements, or the specified element below, or the specified element whose overlapping area occupies a smaller portion of its total area, or the specified element whose overlapping area occupies a larger portion of its total area, and so on.

[0126] For example, if the element to be expanded is the uppermost specified element among overlapping specified elements, then the element to be deleted is the lower specified element, and vice versa. If the element to be expanded is a specified element whose overlapping area occupies a smaller portion of its total area, then the element to be deleted is a specified element whose overlapping area occupies a larger portion of its total area, and vice versa, and so on.

[0127] An expanding element is an element that, after expansion of the to-be-expanded element, covers the to-be-deleted element.

[0128] The updated new target simulation object refers to the simulation object obtained after deleting redundant specified elements in the new target simulation object.

[0129] In some embodiments, element overlap detection can be performed by determining whether the bounding boxes of the specified elements intersect. When the bounding boxes intersect, these specified elements are considered to overlap. The specific determination method is as follows: For the bounding boxes of two specified elements, obtain their upper left coordinates (x1, y1), lower right coordinates (x2, y2) and the upper left coordinates (x3, y3), lower right coordinates (x4, y4) of the other bounding box respectively. The judgment condition max(x1, x3) < min(x2, x4) and max(y1, y3) < min(y2, y4) here is based on the positional relationship of the bounding boxes in the two-dimensional plane. max(x1, x3) < min(x2, x4) means that in the x-axis direction, there is an overlapping part between the two bounding boxes, that is, the right boundary (the smaller of x2 or x4) of one bounding box is greater than the left boundary (the larger of x1 or x3) of the other bounding box; similarly, max(y1, y3) < min(y2, y4) means that in the y-axis direction, there is an overlapping part between the two bounding boxes. If both of these conditions are satisfied, it is considered that the two bounding boxes intersect, that is, these two specified elements overlap.

[0130] In some embodiments, in order to accurately remove redundant specified elements in the new target simulation object, according to the expanding element, the to-be-deleted elements in the new target simulation object are deleted to obtain the updated new target simulation object, including: determining the color information of the expanding element based on the color information of the to-be-expanded element and the color information of the to-be-deleted element; determining the reference part corresponding to the expanding element from the original object; determining the element retention information of the expanding element according to the color information of the reference part and the color information of the expanding element; and deleting the to-be-deleted elements in the new target simulation object according to the element retention information of the expanding element to obtain the updated new target simulation object.

[0131] Among them, the color information of the to-be-expanded element is used to represent the color of the corresponding part in the original object.

[0132] The color information of the to-be-deleted element is used to represent the color of the corresponding part in the original object.

[0133] The color information of the expanding element is used to represent the color of the corresponding part of the original object to be expanded and the element to be deleted. In some embodiments, the color information of the expanding element can be determined by weighted averaging of the color information of the expanding element and the element to be deleted. The specific weighting method can be determined based on the proportion of the overlapping area to the total area of ​​each element. Let the color information of the element to be expanded be C1, and the proportion of its overlapping area to the total area be w1; let the color information of the element to be deleted be C2, and the proportion of its overlapping area to the total area be w2, and w1 + w2 = 1. Then the color information of the expanding element C1 is... out C can be calculated using the following formula: out = w1C1 + w2C2. For example, if the overlapping area accounts for 60% of the total area of ​​elements to be expanded and 40% of the total area of ​​elements to be deleted, then w1 = 0.6 and w2 = 0.4.

[0134] Color information can include exact color, representative color, composite color, average color, etc.

[0135] The reference location is the part of the original object that corresponds to the expanding element.

[0136] The color information of the reference part is the color of the reference part in the original object.

[0137] Element retention information is used to indicate whether to retain extended elements.

[0138] In some embodiments, if the color information of the extended element matches the color information of the reference portion, the element retention information indicates that the extended element should be retained.

[0139] In some embodiments, if the color information of the extended element and the color information of the reference part do not match, the element retention information indicates that the extended element should not be retained.

[0140] In some embodiments, in order to present a virtual object obtained from a new target simulation object in a specified scene, the method further includes: obtaining an element list of elements constituting the target simulation object and a data format of the scene to be transitioned to, wherein the element list includes element parameters of each element constituting the target simulation object; performing format conversion processing on the element parameters of each element in the element list according to the data format to obtain an updated element list; and sending the updated element list to the scene to be transitioned to, so as to present a virtual object composed of each element in the updated element list in the scene to be transitioned to.

[0141] The element list records the relevant element parameters of each element that constitutes the target simulation object.

[0142] For example, the element list could be a list recording the relevant element parameters of each element that constitutes the target simulation object, where each element can be a reference element. It could also be a list recording the relevant element parameters of each element that constitutes the new target simulation object, where each element can be a specified element. Alternatively, it could be a list recording the relevant element parameters of each element that constitutes the updated new target simulation object, where each element can be an optimized specified element.

[0143] The relevant element parameters for each element in the element list include the element's vertex coordinates, center point coordinates, size, rotation angle, etc.

[0144] A scene to be transitioned to is a scene waiting to introduce the target simulation object. For example, a scene to be transitioned to can be a game scene, a simulation scene, a virtual scene, etc.

[0145] The data format is the format required for the data to be transferred to the scene. For example, the data format may include data processing format, parameter mapping, coordinate system transformation, etc.

[0146] Update the element list to a new list obtained by converting the element list format so that the scene to be transitioned to can read and load it.

[0147] In some embodiments, based on the data format requirements of the scene to be transitioned to, operations such as coordinate transformation, unit conversion, and data type conversion are performed on the element parameters in the element list to obtain an updated element list. For example, for coordinate transformation, if the coordinate system used by the scene to be transitioned to is different from the coordinate system in the current element list, the coordinates of the elements can be transformed to the coordinate system of the scene to be transitioned to through transformations such as translation, rotation, and scaling; for unit conversion, if the units in the current element list are different from the units required by the scene to be transitioned to, conversion can be performed according to unit conversion relationships, such as converting the length unit from centimeters to meters; for data type conversion, if the data type required by the scene to be transitioned to is different from the data type in the current element list, corresponding functions can be used for conversion, such as converting integer types to floating-point types.

[0148] For coordinate transformation, if the coordinate system used by the scene to be transformed into is different from the coordinate system in the current element list, assuming the current coordinate system is O1-x1y1 and the coordinate system of the scene to be transformed into is O2-x2y2, a translation transformation is first performed, translating O1 to O2 with a translation vector of (Δx, Δy). The relationship between the element's coordinates (x2, y2) in the new coordinate system and its coordinates (x1, y1) in the original coordinate system is x2 = x1 + Δx, y2 = y1 + Δy. Then a rotation transformation is performed, assuming a rotation angle of θ. The rotated coordinates (x′2, y′2) can be calculated using the formulas x′2 = x2cosθ - y2sinθ, y′2 = x2sinθ + y2cosθ. Finally, a scaling transformation is performed with a scaling factor of s. x and s y Then the final coordinates (x″2, y″2) are x″2 = s x x′2,y″2=s y y′2.

[0149] Virtual objects are the visual effects that the updated element list presents in the scene to be transitioned to. For example, virtual objects can be virtual props, virtual scenes, virtual characters, etc.

[0150] The method provided in this application embodiment can obtain the original object in an image and multiple simulated objects corresponding to the original object. The simulated object is an object obtained by reproducing the original object using reference elements. The arrangement of reference elements in different simulated objects is different. The method can determine the difference information between the simulated object and the original object at the same location. Based on the difference information, for each location corresponding to the position between the original object and the simulated object, the method can determine the matching part for each part of the original object from multiple simulated objects. The method can combine the matching parts determined for each part of the original object to obtain the target simulated object.

[0151] As can be seen from the above, in this embodiment, the original object corresponds to multiple simulated objects. For each part corresponding to the position between the original object and the simulated objects, based on the differences between the simulated objects and the original object at the same part, suitable parts can be determined from the multiple simulated objects for each part of the original object. This achieves targeted retention of parts in each simulated object that are similar to the original object. The suitable parts determined for each part of the original object are combined so that the final target simulated object is closer to the original object overall compared to the multiple simulated objects, ensuring that the generated target simulated object accurately retains the core features and style of the original object. This improves the quality of content stylization transfer.

[0152] The method described in the above embodiments will be further described in detail below.

[0153] In this embodiment, the method of this application embodiment will be described in detail using the fill element as a geometric element as an example.

[0154] As shown in Figure 2a, the specific process of an image processing method is as follows:

[0155] 201. Get the image and specified elements. The image content includes the original object.

[0156] In some embodiments, obtaining an image includes: obtaining an image to be processed; and performing original object extraction processing on the image to be processed to obtain an image.

[0157] In some embodiments, when extracting the original objects from the image to be processed, operations such as region segmentation and foreground segmentation can be performed according to the image to be processed and user requirements to meet the needs. Technically, the image to be processed is simplified to be more conducive to the input of subsequent algorithms. Techniques employed include, but are not limited to, color clustering and simplified models. Color clustering is an image simplification processing method based on algorithms such as K-means, which can cluster complex gradient light and shadow to avoid an increase in the number of fitted components caused by gradient light and shadow.

[0158] When extracting raw objects from an image to be processed:

[0159] 1) Object Segmentation. Object segmentation refers to performing operations such as region cropping and foreground segmentation on an image according to user requirements to extract the object that the user desires to stylize. This object may be the main subject of the image, such as a person, an anime character, or another object in a scene. The object that the user desires to stylize is not necessarily the complete input image; it may be the main subject of the image, such as a person, an anime character, or another object in a scene. For this type of requirement, this application includes optional image region cropping and foreground segmentation modules, and also supports extracting transparent background regions based on the user-input red-blue-green transparency channels (RGBA channels). The foreground segmentation module typically uses relevant deep learning models.

[0160] 2) Image Simplification. Image simplification refers to processing complex images to preserve the main subject and simplify details. This operation may include multiple processing flows, including but not limited to traditional image processing schemes (such as color clustering, opening and closing operations) and deep learning models (such as generative adversarial networks, diffusion models), to avoid the negative impact of details in the image (such as details of human hair, landscape vegetation, furniture textures, object gradations and lighting, noise in the image, etc.) on the abstract expression of the image content. The comparison effect before and after image simplification is shown in Figure 2b. The details of complex images to be processed are often not conducive to the abstract expression of the image to be processed by filling elements. For example, details of human hair, landscape vegetation, furniture textures, object gradations and lighting, noise in the image, etc., will all have a negative impact on the abstract expression of the image content in the image to be processed. Therefore, this application includes an optional image simplification module to preserve the main subject and simplify details of complex images to be processed. The simplification module may include multiple processing flows, including but not limited to traditional image processing schemes and deep learning models. Color clustering, based on algorithms such as K-means, can cluster complex gradient colors and lighting to avoid increasing the number of fitted components due to gradient lighting. Traditional image processing schemes, such as opening and closing operations, can simplify details and noise. Deep learning models treat image simplification as a computer vision task of image transformation and style transfer, and their models are not limited to Generative Adversarial Networks (GANs) or Diffusion Models (DMs). Among them, multimodal image-to-image models can edit images through multimodal control. Specifically, they can perform image-to-image editing based on text control tasks and new style descriptions. Through multiple control network modules (ControlNet modules), the semantic content of the image to be processed is kept unchanged, while only the style is transferred, based on the reference style map, the semantic segmentation map, and the edge extraction map of the image to be processed. Multimodal image-to-image models are a type of model that can edit images through multimodal control. Specifically, they can perform image-to-image editing based on text control tasks and new style descriptions. Through multiple control network modules (ControlNet modules), the semantic content of the image to be processed is kept unchanged, while only the style is transferred, based on the reference style map, the semantic segmentation map, and the edge extraction map of the image to be processed.

[0161] 202. Determine the reference element based on the type corresponding to the specified element. The subclass of the reference element includes the specified element.

[0162] Based on the type of the specified element, determine the reference element. If the specified element is a right triangle, isosceles triangle, or obtuse triangle, its corresponding type is triangle, and the triangle is used as the reference element. When the specified element is a flower, grass, or tree, its corresponding type is a complex element, and the bounding box of the flower, grass, or tree, or a preset shape that can better summarize the overall outline of the flower, grass, or tree (such as an ellipse, polygon, etc.) is used as the reference element. When the specified element is a parallelogram, rectangle, square, etc., its corresponding type is quadrilateral, and the quadrilateral is used as the reference element.

[0163] 203. Based on the original object in the image, multiple simulated objects are generated using reference elements. The simulated objects are objects that reproduce the original object using reference elements, and the arrangement of reference elements in different simulated objects is different.

[0164] In some embodiments, the construction of simulated objects includes, but is not limited to, the use of approximate optimization algorithms, data-driven methods based on reinforcement learning and large language models.

[0165] Approximate optimization algorithms are used when constructing simulated objects. For example, a greedy algorithm starts with an initial layout of reference elements and, at each step, selects an adjustment operation that minimizes the difference between the simulated object and the original object until a preset stopping condition is met. Approximate optimization algorithms can employ a greedy approach, starting with an initial layout of reference elements and, at each step, selecting an adjustment operation that minimizes the difference between the simulated object and the original object until a preset stopping condition is met.

[0166] Data-driven methods based on reinforcement learning are used when constructing simulated objects. For example, using a Deep Q-Network (DQN), the layout of reference elements is treated as the state, adjustments (such as changing the position, size, or rotation angle of the reference elements) as actions, and the differences between the simulated object and the original object (such as calculating the Hausdorff distance between their contours, the ratio of their areas or volumes, or the Euclidean distance between their centroid coordinates) as rewards. The optimal layout of the reference elements is found through iterative training. This data-driven method based on reinforcement learning can use a Deep Q-Network (DQN), treating the layout of reference elements as the state, adjustments (such as changing the position, size, or rotation angle of the reference elements) as actions, and the differences between the simulated object and the original object (such as calculating the Hausdorff distance between their contours, the ratio of their areas or volumes, or the Euclidean distance between their centroid coordinates) as rewards. The DQN network parameters are initialized with a learning rate of 0.001 and a discount factor of 0.9. In each training step, an action is selected based on the current state. Executing this action yields a new state and reward. This experience is stored using an experience replay mechanism. A batch of experiences is randomly sampled from the experience pool to update the parameters of the DQN network. Training is iterated continuously until the optimal layout of reference elements is found. The experience replay mechanism is a mechanism used in reinforcement learning training based on Deep Q-Networks (DQNs) to store the experience gained by the agent in performing actions in the environment (including information such as the current state, action, reward, and next state), and to randomly sample a batch of experiences from the experience pool to update the parameters of the DQN network, thereby improving the stability and efficiency of training.

[0167] Deep Q-networks are neural networks used in a data-driven approach based on reinforcement learning. They take the layout of reference elements as the state, adjustment operations (such as changing the position, size, rotation angle, etc. of the reference elements) as actions, and simulated differences between the object and the original object (such as calculating the Hausdorff distance between their outlines, the ratio of their areas or volumes, the Euclidean distance between their centroid coordinates, etc.) as rewards. Through iterative training, they find the optimal layout of reference elements.

[0168] The data-driven approach based on large language models is used when constructing simulated objects. First, a pre-trained convolutional neural network is used to extract features from the image. Then, a mapping network transforms the visual feature vectors into textual descriptions. These textual descriptions are then input into the large language model. Based on the input textual descriptions, the large language model uses an attention mechanism to retrieve and infer relevant knowledge, generating layout suggestions for reference elements. The attention mechanism is a mechanism used in large language models for retrieving and inferring relevant knowledge. In the data-driven approach based on large language models, the large language model focuses on relevant knowledge based on the input image textual descriptions, thereby generating layout suggestions for reference elements.

[0169] Data-driven methods based on large language models can input image descriptions (such as the type, location, and color of objects in the image) into a large language model. First, a pre-trained convolutional neural network (such as ResNet) is used to extract visual feature vectors from the image. Next, a mapping network, which can be a multilayer perceptron, transforms the visual feature vectors into textual descriptions. This mapping network maps the visual feature vectors to a text embedding space, and then a vocabulary is used to convert the embedding vectors into specific textual descriptions, such as 'There is a red circular object in the center of the image'. Then, the generated textual descriptions are input into the large language model. During pre-training, the large language model learns a large amount of image-layout knowledge. Based on the input textual descriptions, it uses an attention mechanism to retrieve and infer relevant knowledge, generating layout suggestions for reference elements. For example, the large language model can output parameters such as the position of the reference element (e.g., 'The circular reference element is located in the center of the image, with coordinates (x, y)'), its size (e.g., 'The radius of the circular reference element is r'), and its rotation angle (e.g., 'The circular reference element is rotated by 0 degrees'). A multilayer perceptron is a mapping network used in data-driven approaches based on large language models. It maps image visual feature vectors extracted by convolutional neural networks to a text embedding space so as to transform them into specific textual descriptive information.

[0170] In constructing the simulated object, a greedy algorithm is used for approximate optimization. First, an initial layout of reference elements is randomly generated as the starting point. Then, a series of possible adjustments are made to the layout of the reference elements, such as changing their position, size, and rotation angle. For each adjustment, the difference between the adjusted simulated object and the original object is calculated (e.g., by calculating the Hausdorff distance between their outlines, the ratio of their areas or volumes, the Euclidean distance between their centroid coordinates, etc.). The adjustment that minimizes the difference is selected and implemented each time, and this process is repeated until a preset stopping condition is met, such as reaching the maximum number of iterations or the difference falling below a certain threshold.

[0171] In some embodiments, when constructing a simulated object, multiple polygonal partitions can be obtained by nesting image processing algorithms (such as performing edge detection first, and then performing region segmentation on the detection results). A traditional polygon fitting algorithm (such as least squares polygon fitting) can then be applied to each polygonal partition to further reduce the number of reference elements. Least squares polygon fitting is a traditional polygon fitting algorithm. When constructing a simulated object, multiple polygonal partitions obtained through nested image processing algorithms can be fitted to further reduce the number of reference elements.

[0172] In some embodiments, when constructing simulated objects, image complexity can be graded based on factors such as texture complexity and color richness (e.g., images with simple textures and uniform colors are defined as simple images, and vice versa). Simple images can be replaced with faster, less component-intensive traditional algorithms (such as linear fitting algorithms). Specifically, texture complexity indices, such as the contrast and entropy of the gray-level co-occurrence matrix, and color richness indices, such as the entropy of the color histogram, can be calculated. When the texture complexity index is below a certain threshold and the color richness index is below another threshold, the image is determined to be simple and a linear fitting algorithm is used; otherwise, a more complex algorithm, such as a polygon fitting algorithm, is used. For example, when the contrast of the texture complexity index is below 100 and the entropy of the color richness index is below 2, the image is determined to be simple. The entropy of the color histogram is an indicator used to measure the color richness of an image; the higher the entropy value, the richer the colors in the image, and it can be used as a measure of color richness when judging image complexity. The contrast ratio of the gray-level co-occurrence matrix is ​​an indicator used to measure the texture complexity of an image. A higher contrast value indicates a more complex texture, and it can be used as a criterion for judging image complexity. Linear fitting-based algorithms are a traditional algorithm that is faster and requires fewer components when constructing simulated objects for simple images (simple texture, single color).

[0173] In some embodiments, in addition to a one-time global fit, a user interaction process can be developed to further refine the specific important regions selected by the user to improve the results.

[0174] As shown in Figure 2c, this application can also be used for image-specific style transfer, transforming images into specific styles (such as pixel style and low-poly style) built upon basic elements, for creative display, content creation, etc. Low-poly style is an image-specific style that reinterprets the structure and texture of the original image with a smaller number of polygons by selecting, adding, and transforming basic polygonal elements (such as triangles, quadrilaterals, etc.). This style simplifies and abstracts the original image, emphasizing the use of simple polygonal combinations to present the approximate shape and outline of objects, discarding excessive details and complex lighting effects, thus forming an image style with a unique artistic feel and simple visual effect, often used in creative display, content creation, and other fields.

[0175] In some embodiments, multiple simulated objects are generated based on the original object in the image and using reference elements, including: determining the element layout position on the original object in the image; obtaining the target color information of the original object at the element layout position; performing color filling processing on the reference elements at the element layout position according to the target color information to obtain a primary simulated object; and adjusting the reference elements constituting the primary simulated object based on the original object in the image to obtain a secondary simulated object.

[0176] In some embodiments, to avoid the infinite generation of simulation objects, the method further includes: obtaining a preset simulation stop condition; and stopping the adjustment processing of the reference elements constituting the primary simulation object according to the preset simulation stop condition.

[0177] When converting an image into a simulated object composed of multiple specified elements, it is necessary to first design according to the type of the specified elements, specifically including:

[0178] For example, as shown in Figure 2d, when the reference element is determined to be a triangle based on the type of the specified element, the triangulation algorithm can construct multiple non-intersecting triangular meshes based on the element layout positions on the original object in the image (layout points created on the original object in the image). This means aligning the vertices of the triangles with the layout points on the original object in the image, thus reproducing the original object to obtain a simulated object. Therefore, a series of vertices can be defined, and the coordinates of each layout point can be optimized using an approximation algorithm to adjust the reference elements constituting the primary simulated object, thereby obtaining the secondary simulated object.

[0179] When the type of the specified element determines that the reference element is a circle, the center position and radius of the circle constituting the primary simulation object can be optimized to obtain a secondary simulation object. For example, when the reference element is an ellipse, parameters such as the center point, semi-axis length, and rotation angle of the ellipse constituting the primary simulation object can be optimized to obtain a secondary simulation object. When the type of the specified element determines that the reference element is a structurally complex element, such as flowers, trees, tables, chairs, etc., the structurally complex elements constituting the primary simulation object (such as position, size, and rotation parameters) can be optimized to obtain a secondary simulation object. It is worth mentioning that the specified element is not necessarily directly optimized. For example, the optimization effect of a right triangle is not as good as that of a regular triangle. Therefore, a triangle can be used as the fitting target, and then a regular triangle can be replaced by direct triangle post-processing and approximate optimization.

[0180] The reference elements constituting the primary simulation object are adjusted to obtain the secondary simulation object, which can be specifically:

[0181] Based on the original object in the image, the reference elements constituting the initial simulated object are adjusted. The quality is evaluated and further optimization is determined by the differences between the simulated object and the original object at the same location. Parameter initialization can use a random Gaussian distribution, or leverage prior knowledge such as object edge detection or image semantic detection to obtain the original object in the image, thus determining the element layout positions on the original object. The optimization objective is typically defined as semantic perception loss (LPIPS loss), pixel consistency loss (L1 / L2 loss), and other custom losses (such as edge loss, a custom loss function used to measure the difference in edge features between the simulated object and the original image when adjusting the reference elements constituting the initial simulated object). Semantic perception loss measures the semantic difference between the simulated object and the original image when adjusting the reference elements constituting the initial simulated object. Pixel consistency loss measures the pixel-level consistency between the simulated object and the original image when adjusting the reference elements constituting the initial simulated object, and can take forms such as L1 loss and L2 loss. The reference elements constituting the primary simulated object can be adjusted using population approximation search algorithms such as genetic algorithms and particle swarm optimization. These algorithms randomly perturb the element layout positions on the original object in the image to obtain updated element layout positions. The reference elements at these updated positions constitute the secondary simulated object. The perturbation result, i.e., whether to retain the updated element layout positions, is determined by comparing the perceptual loss change with the input image. The perturbation amplitude can be controlled according to rules. Element layout positions and / or updated element layout positions can be manually added according to manually designed rules, as shown in Figure 2e, including adding, deleting, merging, and splitting element layout positions.

[0182] 204. Determine the differences between the simulated object and the original object at the same location.

[0183] 205. For each part corresponding to the position between the original object and the simulated object, based on the difference information, determine the matching part for each part of the original object from multiple simulated objects.

[0184] 206. Combine the adapted parts determined for each part of the original object to obtain the target simulation object.

[0185] 207. Based on the specified elements, perform element transformation processing on the reference elements that constitute the target simulation object to obtain a new target simulation object. The new target simulation object is the object obtained by reproducing the original object using the specified elements.

[0186] For example, as shown in Figure 2f, after setting preset simulation stopping conditions (such as iteration time), a primary simulated object, similar to the original object and composed of reference elements, has been formed. In cases where certain specified elements differ from the reference elements, it is necessary to convert the reference elements constituting the target simulated object into specified elements.

[0187] As shown in Figure 2g, taking the transformation from a triangle to a right triangle as an example: the initial simulation object formed by combining triangulation algorithms consists of multiple arbitrary triangles, which need to be post-processed to transform into right triangles or isosceles triangles, etc. According to the definition, a right triangle can be obtained by drawing a perpendicular line to any triangle. For an obtuse triangle, the perpendicular lines drawn from the two acute angle vertices are not inside the triangle; therefore, drawing a perpendicular line from the vertex containing the largest angle can divide the arbitrary triangle into two right triangles, obtaining the specified element that meets the requirements.

[0188] In some embodiments, optimizing a new target simulation object involves significantly reducing the number of specified elements with minimal loss of accuracy. Depending on the application scenario, specified elements may be adapted, approximated, deleted, split, or merged in certain situations.

[0189] In some embodiments, after performing element transformation processing on the reference elements constituting the target simulation object based on specified elements to obtain a new target simulation object, the method further includes: performing element overlap detection on the specified elements constituting the new target simulation object to obtain elements to be expanded and elements to be deleted, wherein the elements to be expanded and elements to be deleted overlap; performing expansion processing on the elements to be expanded in the new target simulation object according to the elements to be deleted to obtain expanded elements, wherein the expanded elements include the elements to be deleted; and deleting the elements to be deleted in the new target simulation object according to the expanded elements to obtain an updated new target simulation object.

[0190] The elements to be expanded and deleted can be determined by calculating the proportion of the overlapping area to the total area of ​​each specified element. Specified elements with a smaller overlapping area are designated as elements to be expanded, and those with a larger overlapping area are designated as elements to be deleted. For example, if elements A and B overlap, with the overlapping area occupying 20% ​​of element A's total area and 80% of element B's total area, then element A is designated as the element to be expanded, and element B is designated as the element to be deleted.

[0191] In some embodiments, the process of deleting the element to be deleted in the new target simulation object based on the extended element to obtain an updated new target simulation object includes: determining the color information of the extended element based on the color information of the element to be extended and the color information of the element to be deleted; determining the reference part corresponding to the extended element from the original object; determining the element retention information of the extended element based on the color information of the reference part and the color information of the extended element; and deleting the element to be deleted in the new target simulation object based on the element retention information of the extended element to obtain an updated new target simulation object.

[0192] For example, transforming the reference elements constituting the target simulation object into specified elements may cause the number of specified elements to increase exponentially, and direct optimization may produce redundant specified elements. Therefore, approximate optimization alternatives can be designed. Taking the transformation of a triangle into a right triangle as an example: when component overlap is allowed, in addition to segmentation, approximate expansion can also be used. Specifically, any acute-angled side can be expanded into a right-angled side, or obtuse-angled sides can be supplemented to right-angled sides. Taking the transformation of a triangle into an isosceles triangle as an example: any shorter side can be supplemented to be the same length as the longer side. The triangle approximate optimization is shown in Figure 2h. In addition to triangles, approximate optimization can also include some traditional image processing schemes, such as ellipse recognition, and attempt to replace polygonal regions of the same color with approximate ellipses.

[0193] Since updating and reconstructing the simulation object will generate an expanded region, which may lead to a decrease in the perceptual accuracy of the updated simulation object, rules can be manually designed to determine whether approximate optimization is feasible. For example, the difference between the color information of the expanded fill element and the color information of the reference part in the original object can be used to determine whether to proceed. A comparison of the filled element adaptation and the approximate optimized image is shown in Figure 2i.

[0194] In some embodiments, the method further includes: obtaining an element list of elements constituting a target simulation object and a data format of the scene to be transitioned to, the element list including element parameters of each element constituting the target simulation object; performing format conversion processing on the element parameters of each element in the element list according to the data format to obtain an updated element list; and sending the updated element list to the scene to be transitioned to, so as to present a virtual object composed of each element in the updated element list in the scene to be transitioned to.

[0195] For example, an element list constituting the target simulation object has already been obtained, including the two-dimensional position of each element. In some application scenarios, such as the automated construction of user-generated content games (UGC games), a conversion interface also needs to be designed to perform format conversion processing on the element parameters of each element in the element list, transforming them into element parameters that the game can recognize. For example, calculating the center point coordinates, side lengths, scaling scales and rotation angles of relatively standard fill elements from vertex coordinates, etc., to complete the integration. User-generated content games refer to a type of game that allows users to create, edit, and share game content. In this type of game, some users expect to add advanced semantic content (such as anime characters, landscape photos, artistic creations, etc.) to the game.

[0196] This application is applicable to all application scenarios that require converting images to a specific element style, such as UGC game element construction and image style transfer.

[0197] As shown in Figure 2f, in UGC games, some users expect to incorporate advanced semantic content, such as anime characters, landscape photos, and artistic creations. This requires manually building and meticulously editing dozens or even hundreds of components to obtain a complex image, which is time-consuming and has a high creative threshold. The image element-based intelligent construction system allows players to upload original images and selectable actions to approximate them into a new target simulated object composed of multiple specified elements within seconds to minutes.

[0198] The beneficial effects of this application are:

[0199] (1) Compared with the block pixelation scheme, it has higher precision, can obtain better image approximation visualization effect with fewer elements, and retain more details;

[0200] (2) It is time-efficient and controllable, several orders of magnitude more efficient than manual editing by users, which speeds up user creation and design, and can be personalized and adjusted over time, with better results as time and complexity increase;

[0201] (3) It has good applicability and is suitable for a variety of filling elements. The scheme can be designed and expanded in detail. It is not limited to simple geometric shapes in traditional image processing. It can also handle more complex target shapes. It is also suitable for a variety of input images, including simple abstract patterns and extremely complex landscape photos. It reduces the number of components while retaining the effect well.

[0202] (4) Iterative optimization avoids the error caused by the failure of object detection in one go in traditional algorithms and gradually obtains better results.

[0203] As shown above, this application selectively retains the parts of each simulated object that are similar to the original object, and determines the matching parts combination for each part of the original object. This ensures that the final target simulated object is closer to the original object as a whole compared to multiple simulated objects, guaranteeing that the generated target simulated object can accurately retain the core features and style of the original object. This improves the quality of content stylization transfer.

[0204] To better implement the above methods, this application also provides an image processing apparatus, which can be integrated into an electronic device, such as a terminal or server. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, or personal computer; the server can be a single server or a server cluster composed of multiple servers.

[0205] For example, in this embodiment, the method of this application embodiment will be described in detail by taking the image processing device specifically integrated into an electronic device as an example.

[0206] For example, as shown in FIG3, the image processing device may include an acquisition unit 301, a difference determination unit 302, a part determination unit 303, and a part combination unit 304, as follows:

[0207] (I) Acquisition Unit 301.

[0208] The acquisition unit 301 is used to acquire the original object in the image and multiple simulated objects corresponding to the original object. The simulated objects are objects obtained by reproducing the original object using reference elements. The arrangement of reference elements in different simulated objects is different.

[0209] In some embodiments, the apparatus further includes an object generation unit, configured to: generate a primary simulated object based on an original object in an image using reference elements; and adjust the reference elements constituting the primary simulated object based on the original object in the image to obtain a secondary simulated object.

[0210] In some embodiments, the object generation unit is configured to: generate a new arrangement based on the original object in the image and the arrangement of reference elements in the primary simulated object; and adjust the reference elements constituting the primary simulated object according to the new arrangement to obtain a secondary simulated object.

[0211] In some embodiments, the object generation unit is configured to: determine the element layout positions on the original object in the image;

[0212] A basic mock object is generated based on the element's layout position and reference element.

[0213] In some embodiments, the object generation unit is configured to: obtain target color information of the original object at the element layout position; and perform color filling processing on the reference element at the element layout position according to the target color information to obtain a preliminary simulated object.

[0214] In some embodiments, the object generation unit is configured to: determine the parts of a primary simulated object to be adjusted; and adjust the arrangement of reference elements of the parts to be adjusted in the primary simulated object based on the original object in the image to obtain a new arrangement.

[0215] (II) Difference Determination Unit 302.

[0216] The difference determination unit 302 is used to determine the difference information between the simulated object and the original object at the same location.

[0217] In some embodiments, the difference determination unit is configured to: determine target difference information between the primary simulated object and the original object at the same location;

[0218] The object generation unit is used to: determine the parts to be adjusted from the primary simulation objects based on the target difference information.

[0219] (III) Location Determination Unit 303.

[0220] The part determination unit 303 is used to determine the appropriate parts for each part of the original object from multiple simulated objects based on the difference information, for each part corresponding to the position between the original object and the simulated object.

[0221] (iv) Part combination unit 304.

[0222] The part combination unit 304 is used to combine the adaptable parts determined for each part of the original object to obtain the target simulation object.

[0223] In some embodiments, the apparatus further includes an element determination unit and an object conversion unit:

[0224] The element determination unit is used to obtain a specified element; based on the type corresponding to the specified element, a reference element is determined, and the subclass of the reference element includes the specified element.

[0225] The object transformation unit is used to perform element transformation processing on the reference elements that constitute the target simulation object based on specified elements, so as to obtain a new target simulation object. The new target simulation object is the object obtained by reproducing the original object using the specified elements.

[0226] In some embodiments, the apparatus further includes an object updating unit, configured to: perform element overlap detection on specified elements constituting a new target simulation object to obtain elements to be expanded and elements to be deleted, wherein the elements to be expanded and elements to be deleted overlap; perform expansion processing on the elements to be expanded in the new target simulation object according to the elements to be deleted to obtain expanded elements, wherein the expanded elements include the elements to be deleted; and delete the elements to be deleted in the new target simulation object according to the expanded elements to obtain an updated new target simulation object.

[0227] In some embodiments, the object update unit is configured to: determine the color information of the expanding element based on the color information of the element to be expanded and the color information of the element to be deleted; determine the reference part corresponding to the expanding element from the original object; determine the element retention information of the expanding element based on the color information of the reference part and the color information of the expanding element; and delete the element to be deleted in the new target simulation object based on the element retention information of the expanding element to obtain the updated new target simulation object.

[0228] In some embodiments, the apparatus further includes an object introduction unit, configured to: obtain an element list of elements constituting a target simulation object and a data format of the scene to be transitioned to, the element list including element parameters of each element constituting the target simulation object; perform format conversion processing on the element parameters of each element in the element list according to the data format to obtain an updated element list; and send the updated element list to the scene to be transitioned to, so as to present a virtual object composed of each element in the updated element list in the scene to be transitioned to.

[0229] In practice, each of the above units can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units, please refer to the previous method embodiments, which will not be repeated here.

[0230] As can be seen from the above, the image processing device of this embodiment acquires the original object in the image by the acquisition unit, and multiple simulated objects corresponding to the original object. The simulated object is an object obtained by reproducing the original object using reference elements. The arrangement of the reference elements in different simulated objects is different. The difference determination unit determines the difference information between the simulated object and the original object at the same location. The location determination unit determines the appropriate location for each location corresponding to the position between the original object and the simulated object based on the difference information from multiple simulated objects. The location combination unit combines the appropriate locations determined for each location of the original object to obtain the target simulated object.

[0231] Therefore, this embodiment of the application selectively retains the parts of each simulated object that are similar to the original object, and determines the combination of adapted parts for each part of the original object. This ensures that the final target simulated object is closer to the original object as a whole compared to multiple simulated objects, thus ensuring that the generated target simulated object can accurately retain the core features and style of the original object. This improves the quality of content stylization transfer.

[0232] This application also provides an electronic device, which can be a terminal, a server, or other similar device. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, personal computer, etc.; the server can be a single server or a server cluster composed of multiple servers, etc.

[0233] In some embodiments, the image processing apparatus may also be integrated into multiple electronic devices, such as multiple servers, with the image processing method of this application being implemented by the multiple servers.

[0234] In this embodiment, a server will be used as an example for detailed description. For instance, as shown in Figure 4, which illustrates the structural diagram of the server involved in this application embodiment, specifically:

[0235] The server may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input module 404, and a communication module 405. Those skilled in the art will understand that the server structure shown in Figure 4 does not constitute a limitation on the server, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0236] The processor 401 is the control center of the server, connecting various parts of the server via various interfaces and lines. It performs various server functions and processes data by running or executing software programs and / or modules stored in the memory 402, and by calling data stored in the memory 402. In some embodiments, the processor 401 may include one or more processing cores; in some embodiments, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 401.

[0237] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the server, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0238] The server also includes a power supply 403 that supplies power to the various components. In some embodiments, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0239] The server may also include an input module 404, which can be used to receive input numeric or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0240] The server may also include a communication module 405. In some embodiments, the communication module 405 may include a wireless module, through which the server can perform short-range wireless transmission, thereby providing users with wireless broadband internet access. For example, the communication module 405 can be used to help users send and receive emails, browse web pages, and access streaming media.

[0241] Although not shown, the server may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the server loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402, thereby implementing the steps in the methods of the various embodiments of this application.

[0242] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0243] As shown above, it is possible to selectively retain the parts of each simulated object that are similar to the original object, and combine the adapted parts determined for each part of the original object. This makes the final target simulated object more closely resemble the original object as a whole compared to multiple simulated objects, ensuring that the generated target simulated object accurately retains the core features and style of the original object. This improves the quality of content stylization transfer.

[0244] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0245] To this end, embodiments of this application provide a computer-readable storage medium storing multiple instructions that can be loaded by a processor to execute steps in any of the image processing methods provided in embodiments of this application. For example, the instructions can execute the following steps: acquiring an original object in an image, and multiple simulated objects corresponding to the original object, wherein the simulated objects are objects obtained by reproducing the original object using reference elements, and the arrangement of the reference elements in different simulated objects is different; determining the difference information between the simulated objects and the original object at the same location; for each location corresponding to the position between the original object and the simulated objects, determining the appropriate locations for each location of the original object from the multiple simulated objects based on the difference information; and combining the appropriate locations determined for each location of the original object to obtain a target simulated object.

[0246] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0247] According to one aspect of this application, a computer program product or computer program is provided, comprising a computer program / instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer program / instructions from the computer-readable storage medium and executes the computer program / instructions, causing the electronic device to perform the methods provided in the various optional implementations of the image processing aspects provided in the above embodiments.

[0248] Since the instructions stored in the storage medium can execute the steps of any of the image processing methods provided in the embodiments of this application, the beneficial effects that any of the image processing methods provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.

[0249] In summary, this application provides an image processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. The electronic device first acquires an original object and its corresponding multiple simulated objects from an image. These simulated objects are obtained by reproducing the original object using reference elements, and the arrangement of the reference elements varies among the different simulated objects. Next, the electronic device meticulously compares the simulated objects and the original object at the same location to determine the differences between them. Then, based on the difference information, for each location corresponding to the original object and the simulated objects, it precisely selects suitable parts from the multiple simulated objects for each part of the original object. Finally, these suitable parts are combined to obtain the target simulated object. From a technical perspective, this method fully utilizes the advantages of the different arrangements of multiple simulated objects. By selecting suitable parts, it can selectively retain parts in each simulated object that are similar to the original object. In image processing, the features and style of an image are often scattered across different local areas. Through this method of combining suitable parts, these scattered similar features can be integrated, making the final target simulated object closer to the original object as a whole, accurately preserving the core features and style of the original object. During content stylization transfer, the style of the original object can be transferred to the target simulated object more accurately, avoiding the loss or deformation of some features that may be caused by overall fitting, thus significantly improving the quality of content stylization transfer.

[0250] Furthermore, based on the original object in the image, the electronic device first generates a primary simulated object using reference elements. Then, based on the original object, it adjusts the reference elements constituting the primary simulated object to obtain a secondary simulated object. In image processing, a simulated object generated in one step may not perfectly match all the features of the original object. By generating simulated objects in stages, first obtaining a primary simulated object as a foundation, and then adjusting it, the generation process of simulated objects becomes more flexible. Different adjustment strategies can optimize for different features of the original object, thereby obtaining multiple simulated objects similar to the original object. When subsequently selecting adaptable parts for various parts of the original object, more simulated objects mean more selection possibilities. It is possible to select the adaptable parts that are most similar to various parts of the original object, further improving the similarity between the target simulated object and the original object. This allows for a more accurate presentation of the style of the original object during content stylization transfer, improving the quality of content stylization transfer.

[0251] Furthermore, when generating secondary simulated objects, the electronic device generates a new arrangement based on the original object in the image and the arrangement of reference elements in the primary simulated object. Then, it adjusts the reference elements constituting the primary simulated object according to this new arrangement to obtain the secondary simulated object. The arrangement of reference elements in the primary simulated object is initially generated and may contain some discrepancies with the features of the original object. The electronic device generates a new arrangement by analyzing the features of the original object and the arrangement of the primary simulated object. This new arrangement optimizes the arrangement of reference elements based on the features of the original object. During the adjustment process, the distribution of reference elements better matches the structure and features of the original object, making the secondary simulated object closer to the original object than the primary simulated object. This improves the accuracy and efficiency of the simulated object generation process, reduces unnecessary adjustments, and thus enables more efficient and accurate transfer of the style of the original object to the simulated object during content stylization transfer, improving the quality of content stylization transfer.

[0252] Furthermore, the electronic device determines the element layout positions on the original object in the image, and generates a preliminary simulated object based on these positions and reference elements. In image processing, the combination of reference elements significantly impacts the similarity between the simulated object and the original object. Determining the element layout positions provides clear rules and frameworks for combining reference elements. The original object possesses specific structures and features, and the determination of element layout positions is based on these structures and features, enabling reference elements to be combined systematically according to the structure of the original object, thus more accurately reproducing it. The improved similarity between the preliminary simulated object and the original object lays a solid foundation for generating better simulated objects subsequently. Subsequent adjustments and optimizations can be performed on a basis closer to the original object, reducing the difficulty and workload of adjustments and improving the quality of content stylization transfer.

[0253] Furthermore, the electronic device acquires the target color information of the original object at its element layout location. Based on this target color information, it performs color filling processing on the reference elements at those locations, resulting in a preliminary simulated object. Color is one of the important features of an image, and in content stylization transfer, accurate color representation is crucial for preserving the style of the original object. By acquiring the target color information of the original object at its element layout location and filling the reference elements with color, the simulated object acquires color information similar to the original object. In subsequent processing, color information can serve as an important basis for distinguishing the differences between the simulated object and the original object. When analyzing the differences between the simulated object and the original object, color differences can be intuitively reflected, facilitating a more accurate identification of areas requiring adjustment. When adjusting the simulated object, adjustments based on color information can more precisely bring the simulated object closer to the original object, improving the quality of content stylization transfer.

[0254] Furthermore, when generating a new layout, the electronic device identifies the parts of the primary simulated object to be adjusted. Based on the original object in the image, it adjusts the arrangement of reference elements for the parts to be adjusted in the primary simulated object to obtain a new layout. The primary simulated object may have parts that are dissimilar to the original object. If the entire simulated object is adjusted uniformly, it may destroy parts that are already similar to the original object. By identifying the parts to be adjusted, adjustments can be made specifically to the dissimilar parts. During the adjustment process, parts similar to the original object are retained, and only dissimilar parts are updated, making the adjustment more precise and effective. This makes the subsequently obtained simulated objects more similar to the original object, improving the quality of the simulated objects. In content stylization transfer, more similar simulated objects can more accurately represent the style of the original object, improving the quality of content stylization transfer.

[0255] Furthermore, the electronic device determines the target difference information between the primary simulated object and the original object at the same location, and identifies the areas to be adjusted from the primary simulated object based on this target difference information. In image processing, accurately identifying the areas that need adjustment is crucial for improving the quality of the simulated object. By analyzing the target difference information, the electronic device can quantify the degree of difference between the primary simulated object and the original object at various locations. Based on the degree of difference, it can accurately determine which areas are dissimilar to the original object and require adjustment. This method of determination based on difference information avoids blind adjustments and improves the accuracy and efficiency of the adjustments. During the adjustment process, it can more effectively optimize dissimilar areas, allowing the simulated object to more quickly approach the original object and improving the quality of content stylization transfer.

[0256] Furthermore, the electronic device acquires specified elements, determines reference elements based on the type corresponding to the specified elements, and combines the adapted parts determined for each part of the original object to obtain the target simulated object. Then, based on the specified elements, it performs element transformation processing on the reference elements constituting the target simulated object to obtain a new target simulated object. In practical applications, users may have specific style requirements for the simulated object. By acquiring specified elements and performing element transformation processing, the style of the target simulated object can be adjusted according to the user's needs. Simultaneously, since the target simulated object is obtained through the combination of adapted parts, it already retains the core features of the original object. During element transformation, the style can be adjusted while preserving these core features. This method improves the applicability and flexibility of the simulated object, meeting the diverse needs of different users. During content stylization migration, the style of the original object can be blended with the style of the specified elements according to the user's requirements, improving the quality of content stylization migration.

[0257] Furthermore, the electronic device performs element overlap detection on the specified elements constituting the new target simulation object, obtaining elements to be expanded and elements to be deleted. Based on the elements to be deleted, the elements to be expanded are expanded to obtain expanded elements. Based on the expanded elements, the elements to be deleted in the new target simulation object are deleted, resulting in the updated new target simulation object. During the element conversion process, overlapping of specified elements may occur. Overlapping elements make the simulation object appear redundant and cluttered, affecting its quality. Element overlap detection can accurately identify the elements to be deleted and the elements to be expanded. Expanding the elements to be expanded fills the gaps left after the deletion of the elements to be deleted, making the structure of the simulation object more complete. After deleting the elements to be deleted, the new target simulation object is more concise and accurate, removing redundant information. During content stylization migration, a concise and accurate simulation object can more clearly present the style of the original object, improving the quality of content stylization migration.

[0258] Furthermore, when deleting an element to be deleted, the electronic device determines the color information of the expanded element based on the color information of the element to be expanded and the color information of the element to be deleted. It then determines the reference part corresponding to the expanded element from the original object. Based on the color information of the reference part and the color information of the expanded element, it determines the element retention information of the expanded element. Based on the element retention information of the expanded element, the element to be deleted is deleted from the new target simulation object, resulting in an updated new target simulation object. Color matching is crucial to the quality of the simulation object when deleting elements to be deleted and expanding them. By comprehensively considering the color information of the element to be expanded, the element to be deleted, and the reference part of the original object, the color information and element retention information of the expanded element can be determined. This allows for the precise removal of redundant specified elements in the new target simulation object while ensuring that the color of the expanded element matches the color of the original object. During content stylization migration, color-matched simulation objects can more realistically present the style of the original object, improving the quality of the new target simulation object and enhancing the quality of content stylization migration.

[0259] Furthermore, the electronic device acquires the element list constituting the target simulated object and the data format of the scene to be transitioned to. Based on the data format, it performs format conversion processing on the element parameters of each element in the element list to obtain an updated element list. This updated element list is then sent to the scene to be transitioned to, so that the virtual object composed of the elements in the updated element list can be presented in the scene. Different scenes may have different data format requirements. If the element parameter format of the target simulated object does not match the scene to be transitioned to, it will not be able to be presented correctly in that scene. By performing format conversion processing on the element parameters, the element parameters of the target simulated object can be made to conform to the data format requirements of the scene to be transitioned to. In this way, the target simulated object can be accurately presented in the specified scene, improving the practicality and applicability of the simulated object. During content stylization migration, the style of the original object can be accurately transferred to the specified scene, improving the quality of content stylization migration.

[0260] Furthermore, the use of various algorithms makes the construction of simulated objects more flexible and efficient. Approximate optimization algorithms, during the construction of simulated objects, start from an initial layout of reference elements and, at each step, select an adjustment operation that minimizes the difference between the simulated object and the original object until a preset stopping condition is met. This algorithm does not require extensive global searches and can quickly find a better layout of reference elements, reducing computation time. It can improve processing efficiency in large-scale image processing or scenarios with high real-time requirements. Data-driven methods based on reinforcement learning treat the layout of reference elements as the state, adjustment operations as actions, and the difference between the simulated object and the original object as the reward, finding the optimal layout of reference elements through iterative training. This method can adaptively adjust based on the features of the original object, continuously optimizing the layout of reference elements and improving the similarity between the simulated object and the original object. Data-driven methods based on large language models first use a pre-trained convolutional neural network to extract features from the image, then convert the visual feature vectors into textual description information through a mapping network. This textual description information is then input into a large language model, which, based on the input textual description, retrieves and infers relevant knowledge through an attention mechanism to generate layout suggestions for reference elements. Large language models possess powerful knowledge reasoning capabilities, enabling them to combine image features and related knowledge to generate reference element layout suggestions that better match image characteristics. The combined use of multiple algorithms allows for the selection of appropriate algorithms based on different image features and processing requirements, making the construction of simulated objects more flexible and efficient, thereby improving the quality of content stylization transfer.

[0261] Furthermore, selecting an appropriate algorithm based on the actual complexity of the image can improve processing efficiency while ensuring the quality of the simulated object. For simple images, the texture complexity and color richness are low, and the image features are relatively few. Using fast, low-component traditional algorithms (such as linear fitting algorithms) can generate the simulated object in a short time, saving computational resources and time. This can significantly improve processing efficiency when processing a large number of simple images. For complex images, the texture and color information are rich, and the image features are complex and diverse. Using more complex algorithms (such as polygon fitting algorithms) can better fit the image features, capture detailed information in the image, and improve the similarity between the simulated object and the original object. By selecting an appropriate algorithm based on the image complexity, processing efficiency can be improved while ensuring the quality of the simulated object. This allows for more efficient and accurate transfer of the style of the original object to the simulated object during content style transfer, thus improving the quality of content style transfer.

[0262] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0263] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

An image processing method, executed by an electronic device, comprises: obtaining an original object in an image, and a plurality of simulation objects corresponding to the original object, the simulation objects being objects obtained by reproducing the original object using reference elements, and the arrangement of the reference elements in different simulation objects being different; determining difference information of the simulation objects and the original object at the same part; for each part corresponding to the position between the original object and the simulation object, determining an adaptive part for each part of the original object from the plurality of simulation objects according to the difference information; and combining the adaptive parts determined for each part of the original object to obtain a target simulation object. The image processing method of claim 1, further comprising: generating a primary simulation object using the reference elements based on the original object in the image; adjusting the reference elements constituting the primary simulation object based on the original object in the image to obtain a secondary simulation object. The image processing method of claim 2, wherein the adjusting the reference elements constituting the primary simulation object based on the original object in the image to obtain a secondary simulation object comprises: generating a new arrangement of the reference elements based on the original object in the image and the arrangement of the reference elements in the primary simulation object; adjusting the reference elements constituting the primary simulation object according to the new arrangement to obtain a secondary simulation object. The image processing method of claim 2 or 3, wherein the generating a primary simulation object using the reference elements based on the original object in the image comprises: determining an element layout position on the original object in the image; generating a primary simulation object based on the element layout position and the reference elements. The image processing method of claim 4, wherein the generating a primary simulation object based on the element layout position and the reference elements comprises: obtaining target color information of the original object at the element layout position; performing color filling processing on the reference elements at the element layout position according to the target color information to obtain a primary simulation object. The image processing method of any one of claims 3 to 5, wherein the generating a new arrangement of the reference elements based on the original object in the image and the arrangement of the reference elements in the primary simulation object comprises: determining an adjustment part of the primary simulation object; adjusting the arrangement of the reference elements of the adjustment part of the primary simulation object based on the original object in the image to obtain a new arrangement. The image processing method of claim 6, wherein the determining difference information of the simulation objects and the original object at the same part comprises: determining target difference information of the primary simulation object and the original object at the same part; the determining an adjustment part of the primary simulation object comprises: determining an adjustment part from the primary simulation object according to the target difference information. The image processing method of any one of claims 1 to 7, further comprising: obtaining a specified element; ​ According to a type corresponding to the specified element, a reference element is determined, and a subclass of the reference element includes the specified element; After the target simulation object is obtained by combining the fitting parts determined for each part of the original object, the method further includes: Based on the specified element, element conversion processing is performed on the reference elements constituting the target simulation object to obtain a new target simulation object, and the new target simulation object is an object obtained by reproducing the original object using the specified element. The image processing method of claim 8, further comprising: Performing element overlap detection on the specified elements constituting the new target simulation object to obtain to-be-expanded elements and to-be-deleted elements, and the to-be-expanded elements and the to-be-deleted elements overlap; According to the to-be-deleted elements, performing expansion processing on the to-be-expanded elements in the new target simulation object to obtain expanded elements, and the expanded elements include the to-be-deleted elements; According to the expanded elements, deleting the to-be-deleted elements in the new target simulation object to obtain an updated new target simulation object. The image processing method of claim 9, wherein the deleting the to-be-deleted elements in the new target simulation object according to the expanded elements to obtain an updated new target simulation object includes: Based on color information of the to-be-expanded elements and color information of the to-be-deleted elements, determining color information of the expanded elements; Determining a reference part corresponding to the expanded elements from the original object; According to color information of the reference part and the color information of the expanded elements, determining element retention information of the expanded elements; According to the element retention information of the expanded elements, deleting the to-be-deleted elements in the new target simulation object to obtain an updated new target simulation object. The image processing method of any one of claims 1-10, further comprising: Obtaining an element list of elements constituting the target simulation object and a data format to be converted into a scene, and the element list includes element parameters of each element constituting the target simulation object; According to the data format, performing format conversion processing on the element parameters of each element in the element list to obtain an updated element list; Sending the updated element list to the scene to be converted into the scene to present a virtual object constituted by each element in the updated element list in the scene to be converted. An image processing apparatus, comprising: An acquisition unit configured to acquire an original object in an image and a plurality of simulation objects corresponding to the original object, the simulation objects being objects obtained by reproducing the original object using reference elements, and the arrangement of the reference elements being different in different simulation objects; A difference determination unit configured to determine difference information of the simulation objects and the original object at the same part; A part determination unit configured to determine, for each part corresponding in position between the original object and the simulation objects, a fitting part for each part of the original object from the plurality of simulation objects according to the difference information; and A part combination unit configured to combine the fitting parts determined for each part of the original object to obtain a target simulation object. An electronic device comprising a processor and a memory, the memory storing a plurality of instructions; the processor loading the instructions from the memory to perform the steps of the image processing method of any one of claims 1-11. A computer readable storage medium storing a plurality of instructions adapted to be loaded by a processor to perform the steps of the image processing method of any one of claims 1-11. A computer program product comprising a plurality of instructions which, when executed by a processor, implement the steps of the image processing method of any one of claims 1-11.

Citation Information

Patent Citations

  • Image processing method and device, computer readable medium and terminal equipment

    CN111652830A

  • Image processing method and device, electronic equipment and storage medium

    CN113706369A

  • Image simulation method and device

    CN115174893A

  • Image processing method and device, equipment, medium and program product

    CN115937020A

  • Parameter configuration method and device of coloring model, computer equipment and storage medium

    CN117635803A