Image generation method, computing device, electronic device, storage medium and program product

The method aligns object models with scene images using structural information to enhance the realism and fidelity of generated images, addressing the low fidelity issue in existing image generation methods.

CN120318357APending Publication Date: 2025-07-15ZHEJIANG TMALL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510407441.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the prior art, the target image generated by stitching the product image and scene image has low fidelity, resulting in loss of details and discord.

Method used

By obtaining the object image and object model of the target object, as well as the scene image of the target scene, the object model is placed into the scene image according to the model placement parameters, the conditional image corresponding to the initial image is determined based on the structural information presented by the object model in the initial image, and the conditional image and the object image are fused to generate the target image.

Benefits of technology

While retaining the details of the target object and the target scene, the harmony between the target object and the target scene is ensured, and the realistic and fidelity of the generated target image is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318357A_ABST
    Figure CN120318357A_ABST
Patent Text Reader

Abstract

The invention discloses an image generation method, computing equipment, electronic equipment, a storage medium and a program product. The method relates to the field of image processing, and comprises the steps that an object image and an object model of a target object and a scene image of a target scene are acquired, and the target object is to be added to the target scene; placing the object model in the scene image according to the model placement parameters to obtain an initial image; determining a condition image corresponding to the initial image based on structure information presented by the object model in the initial image; and fusing the condition image and the object image to generate a target image. According to the invention, the technical problem that the fidelity of the generated image is low in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and in particular, to an image generation method, a computing device, an electronic device, a storage medium, and a program product. Background Art

[0002] In modern e-commerce, in order to help users make more reasonable shopping decisions, many e-commerce platforms have begun to provide virtual preview functions for products. In the virtual preview function, users can select suitable products according to their own needs and preferences, and splice the product images of the selected products into a specified scene image, such as a living room, a bedroom, or an office. By generating a corresponding target image in this way, users can intuitively see the placement effect of the product in a specific environment, thereby enhancing the user's shopping experience and satisfaction. However, since the product image and the scene image may contain more details and the images may have different styles, simply splicing the product image and the scene image may result in detail loss and disharmony in the generated target image, resulting in a low fidelity of the actually generated image.

[0003] For the above problems, no effective solution has been proposed yet. Summary of the Invention

[0004] Embodiments of this application provide an image generation method, a computing device, an electronic device, a storage medium, and a program product to at least solve the technical problem of low fidelity of the generated image in related technologies.

[0005] According to one aspect of the embodiments of this application, an image generation method is provided, including: obtaining an object image and an object model of a target object, and a scene image of a target scene, where the target object is to be added to the target scene; placing the object model into the scene image according to the model placement parameters to obtain an initial image; determining a conditional image corresponding to the initial image based on the structural information presented by the object model in the initial image; and fusing the conditional image and the object image to generate a target image.

[0006] According to one aspect of the embodiments of this application, an image generation method is further provided, including: in response to an input instruction acting on an operation interface, displaying an object image and an object model of a target object, and a scene image of a target scene on the operation interface, where the target object is to be added to the target scene; and in response to an image generation instruction acting on the operation interface, displaying a target image on the operation interface, where the target image is obtained by fusing a conditional image and an object image, and the conditional image is determined by the structural information presented by the object model in the initial image, and the initial image is used to represent that the object model is added to the scene image based on the model placement parameters.

[0007] According to one aspect of the embodiments of the present application, there is also provided an image generation method, including: obtaining an object image and an object model of a target object, and a scene image of a target scene by calling a first interface, where the first interface includes a first parameter, and the parameter value of the first parameter includes the object image, the object model, and the scene image, and the target object is to be added to the target scene; placing the object model into the scene image according to the model placement parameter to obtain an initial image; determining a conditional image corresponding to the initial image based on the structural information presented by the object model in the initial image; fusing the conditional image and the object image to generate a target image; and outputting the target image by calling a second interface, where the second interface includes a second parameter, and the parameter value of the second parameter includes the target image.

[0008] According to another aspect of the embodiments of the present application, there is also provided an image generation device, including: an obtaining module, configured to obtain an object image and an object model of a target object, and a scene image of a target scene, where the target object is to be added to the target scene; a placing module, configured to place the object model into the scene image according to the model placement parameter to obtain an initial image; a determining module, configured to determine a conditional image corresponding to the initial image based on the structural information presented by the object model in the initial image; and a fusing module, configured to fuse the conditional image and the object image to generate a target image.

[0009] According to one aspect of the embodiments of the present application, there is also provided an image generation device, including: a first execution module, configured to display an object image and an object model of a target object, and a scene image of a target scene on an operation interface in response to an input instruction acting on the operation interface, where the target object is to be added to the target scene; a second execution module, configured to display a target image on the operation interface in response to an image generation instruction acting on the operation interface, where the target image is obtained by fusing a conditional image and an object image, the conditional image is determined by the structural information presented by the object model in the initial image, and the initial image is used to represent that the object model is added to the scene image based on the model placement parameter.

[0010] According to one aspect of the embodiments of the present application, there is also provided an image generation method, including: a first calling module, configured to obtain an object image and an object model of a target object, and a scene image of a target scene by calling a first interface, where the first interface includes a first parameter, and the parameter value of the first parameter includes the object image, the object model, and the scene image, and the target object is to be added to the target scene; a model placement module, configured to place the object model into the scene image according to model placement parameters to obtain an initial image; an image determination module, configured to determine a conditional image corresponding to the initial image based on the structural information presented by the object model in the initial image; an image fusion module, configured to fuse the conditional image and the object image to generate a target image; a second calling module, configured to output the target image by calling a second interface, where the second interface includes a second parameter, and the parameter value of the second parameter includes the target image.

[0011] According to another aspect of the embodiments of the present application, there is also provided a computing device, including: a memory storing an executable program; a processor configured to run the program, where when the program runs, it executes the methods in the various embodiments of the present application.

[0012] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including: a memory storing an executable program; a processor connected to the memory through a bus and configured to run the program, where when the program runs, it executes the methods in the various embodiments of the present application.

[0013] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, where the computer-readable storage medium includes a stored executable program, and when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the methods in the various embodiments of the present application.

[0014] According to another aspect of the embodiments of the present application, there is also provided a computer program product, including a computer program, where when the computer program is executed by a processor, it implements the methods in the various embodiments of the present application.

[0015] According to another aspect of the embodiments of the present application, there is also provided a computer program product, including a non-volatile computer-readable storage medium storing a computer program, where when the computer program is executed by a processor, it implements the methods in the various embodiments of the present application.

[0016] According to another aspect of the embodiments of the present application, there is also provided a computer program, where when the computer program is executed by a processor, it implements the methods in the various embodiments of the present application.

[0017] In an embodiment of the present application, an object image and an object model of a target object, as well as a scene image of a target scene are obtained; the object model is placed into the scene image according to model placement parameters to obtain an initial image; based on the structural information presented by the object model in the initial image, a conditional image corresponding to the initial image is determined; the conditional image and the object image are fused to generate a target image. By using the structural information presented by the object model in the initial image and fusing the appearance information of the target object in the object image and the environmental information of the target scene in the scene image, it is possible to retain the details of the target object and the target scene while ensuring the harmony between the target object and the target scene, improve the realism of the generated target image, and thus solve the technical problem of low fidelity of the generated images in the related art.

[0018] It is easy to notice that the above general description and the following detailed description are only for exemplifying and explaining the present application, and do not constitute a limitation to the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0020] Figure 1 is a schematic diagram of an application scenario of an image generation method shown according to an embodiment of the present application;

[0021] Figure 2 is a flowchart of an image generation method shown according to an embodiment of the present application;

[0022] Figure 3 is a schematic diagram of an image generation process shown according to an embodiment of the present application;

[0023] Figure 4 is a schematic diagram of an image fusion process shown according to an embodiment of the present application;

[0024] Figure 5 is a flowchart of another image generation method shown according to an embodiment of the present application;

[0025] Figure 6 is a flowchart of another image generation method shown according to an embodiment of the present application;

[0026] Figure 7 is a structural block diagram of an image generation device shown according to an embodiment of the present application;

[0027] Figure 8 is a structural block diagram of another image generation device shown according to an embodiment of the present application;

[0028] Figure 9 is a structural block diagram of another image generation device shown according to an embodiment of the present application;

[0029] Figure 10 is a structural block diagram of a computing device shown according to an embodiment of the present application;

[0030] Figure 11 is a structural block diagram of an electronic device shown according to an embodiment of the present application. Detailed implementation manners

[0031] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these process, method, product or device.

[0033] First, some nouns or terms that appear in the process of describing the embodiments of the present application are applicable to the following explanations:

[0034] GLB: GL Transmission Format Binary, a standard binary file format that can package data such as 3D models, materials, textures, and animations into a single binary file for fast loading and rendering in web pages, applications, and virtual reality environments.

[0035] Canny edge detection algorithm: A multi-stage algorithm designed to detect edges in an image, and the Canny line drawing is the result obtained after applying this algorithm to the image, which is usually used in image analysis and computer vision tasks to better understand and process the structure of the image.

[0036] According to an embodiment of the present application, an image generation method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0037] The above method provided by the embodiments of the present application can be applied to, for example, Figure 1 the application scenarios shown, but not limited thereto. Figure 1 FIG. is a schematic diagram of an application scenario of an image generation method shown according to an embodiment of the present application. In the application scenario shown, for example, Figure 1 the model can be deployed in the server 10. The server 10 can be connected to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. The client devices 20 here can include, but are not limited to: smart phones, tablet computers, laptop computers, palmtop computers, personal computers, smart home devices, in-vehicle devices, etc. The client device 20 can interact with the user through a graphical user interface to implement the call of the model, and further implement the method provided by the embodiments of the present application.

[0038] In the embodiments of the present application, the system composed of the client device and the server can execute the following steps: The client device executes to obtain the object image and object model of the target object, and the scene image of the target scene. The server executes to place the object model into the scene image according to the model placement parameters to obtain an initial image; determine the conditional image corresponding to the initial image based on the structural information presented by the object model in the initial image; and fuse the conditional image and the object image to generate a target image.

[0039] It should be noted that with the rapid development of high-performance computing units, in other application scenarios, the above method provided by the embodiments of the present application can also be applied to a model all-in-one machine. In an optional embodiment, multiple models are built in the model all-in-one machine. The user can select and adjust one model according to needs to obtain his own model. Thus, the high-performance computing unit built in the model all-in-one machine can directly call the adjusted model to execute the above method provided by the embodiments of the present application. In another optional embodiment, a trained model is built in the model all-in-one machine. Thus, the high-performance computing unit built in the model all-in-one machine can directly call the model to execute the above method provided by the embodiments of the present application.

[0040] Further, when the user needs to train their own model, they can also upload their own dataset through the client. This dataset is sent from the client to the server, enabling the server to adjust the pre-trained model with this dataset to obtain the user's own model, which is then deployed to the production environment. To facilitate the user's model adjustment requirements, the server can provide complete adjustment tools, development frameworks, and processes, and support multiple adjustment strategies, enabling the adjusted model to better adapt to different field applications and achieve high customization.

[0041] Under the above operating environment, this application provides an image generation method as Figure 2 shown below. Figure 2 It is a flowchart of an image generation method shown according to an embodiment of this application. As Figure 2 shown, this method may include the following steps:

[0042] Step S202, obtain an object image and an object model of the target object, and a scene image of the target scene.

[0043] Among them, the target object is to be added to the target scene.

[0044] The above-mentioned target object may refer to an object that needs to be placed and displayed in the target scene, and may include but are not limited to: objects such as people, vehicles, and seats. The corresponding target scene may include but are not limited to: scenes such as parks, parking lots, and living rooms. The above-mentioned object image may refer to an image obtained by photographing the detailed information presented by the target object, and may be a white-background image. Among them, the detailed information of the target object may include but are not limited to: detailed information such as the structure, texture, and posture presented by the target object.

[0045] In an alternative solution of this embodiment, in order to reasonably display to the user the visual effect presented by the target object and the overall target scene after adding the target object to the target scene, the image generation system may first obtain the object image of the above-mentioned target object to accurately understand the appearance information presented by the target object in the real scene, such as color, texture, shape and other information, and obtain the scene image of the target scene to understand the layout, style, lighting conditions and other information of the target scene, so as to ensure that in the generated target image, the visual content presented by the target object can better conform to the appearance information of the target object in the real scene, and at the same time ensure that the scene information presented by the target image is more natural and harmonious. Further, considering that different users may have different pose information for the target object, when adding the target object to the target scene, the pose of the target object required by the user in the target scene may be different from the pose presented by the target object in the object image. The visual effect presented by the target image generated only by the object image and the scene image may be relatively rigid and cannot meet the needs of the user. Therefore, in order to fully reflect the visual effect that the target object can present in different poses in the target image, while obtaining the object image and the scene image, the image generation system may further obtain the object model of the target object and introduce the object model into the process of generating the target image, so as to allow the user to freely place and adjust the pose information such as the position, orientation and scale of the target object in the target scene, improve the user experience, and at the same time enable the image generation system to fully understand the spatial structure and morphological information of the target object through the object model, so as to ensure the rationality of the size and proportion between the target object and other objects in the target scene when adding the target object to the target scene, and avoid distortion. Based on this, combining the object image and object model of the target object, and the scene image of the target scene, the image generation form can accurately simulate the real appearance that the target object can present in the target scene, including visual effects such as shadows, reflections and light transmission effects, thereby improving the fidelity of the generated target image.

[0046] Step S204, place the object model into the scene image according to the model placement parameters to obtain an initial image.

[0047] The above model placement parameters can be actively set by the user or determined by the image generation system according to the pose information of other objects in the scene image. The model placement parameters may include, but are not limited to: coordinate position, direction and pose, scale and size, etc.

[0048] In an alternative solution of this embodiment, to improve the efficiency of generating the target image, after obtaining the object model and the scene image, the image generation system can first place the object model into the scene image quickly according to the above model placement parameters, so as to reasonably generate an initial image that fuses the target object with the target scene. For example, the image generation system can first place the object model in front of the scene image according to the parameters such as the coordinate position, direction and pose, scale and size included in the model placement parameters, and then use rendering technology to render the object model into the scene image to obtain the above initial image, so as to ensure that the pose information presented by the target object in the generated image can meet the user's needs as much as possible.

[0049] Step S206: Determine the conditional image corresponding to the initial image based on the structural information presented by the object model in the initial image.

[0050] The above structural information may include but is not limited to: the edge position, texture structure, etc. of the target object in the initial image.

[0051] In an alternative solution of this embodiment, considering that the pose presented by the target object in the initial image may be different from the pose presented by the target object in the object image, only relying on the visual information presented in the object image, such as texture, lighting, etc., may not be able to meet the visual information presented by the target object in the initial image. For example, assuming that the target object presents a side view in the initial image, while the target object presents a front view in the object image, then without considering the structural information of the target object in the initial image, directly using the front texture information presented by the target object in the object image to render the texture information presented by the target object in the initial image may result in rendering errors, leading to an inconsistent situation in the finally generated target image. Therefore, to ensure the fidelity of the generated target image, after obtaining the initial image, the image generation system can further perform a structural detection on the initial image to determine the structural information presented by the object model in the initial image, and construct a corresponding conditional image based on this structural information, so that the image generation system can reasonably render the appearance information that the target object in the initial image can present based on the appearance information presented by the target object in the object image on the basis of the conditional image, so as to ensure the rationality of the visual effect presented by the target object in the finally generated target image.

[0052] Step S208: Fuse the conditional image and the object image to generate the target image.

[0053] In an alternative solution of this embodiment, after constructing the conditional image corresponding to the initial image, the image generation system can fuse the conditional image with the object image. Guided by the structural information of the target object in the conditional image, the appearance information of the target object in the object image is reasonably fused into the initial image to obtain the corresponding target image, thereby improving the accuracy of the visual effect presented by the target object in the generated target image, further ensuring the realism of the generated target image, and enhancing the user experience when viewing the target image.

[0054] In the embodiment of the present application, the method includes obtaining the object image and object model of the target object, and the scene image of the target scene; placing the object model into the scene image according to the model placement parameters to obtain the initial image; determining the conditional image corresponding to the initial image based on the structural information presented by the object model in the initial image; and fusing the conditional image and the object image to generate the target image. By using the structural information presented by the object model in the initial image to fuse the appearance information of the target object in the object image and the environmental information of the target scene in the scene image, it is possible to retain the details of the target object and the target scene while ensuring the harmony between the target object and the target scene, improving the realism of the generated target image, and further solving the technical problem of low fidelity of the generated image in the related art.

[0055] In the embodiment of the present application, determining the conditional image corresponding to the initial image based on the structural information presented by the object model in the initial image includes: generating a mask image corresponding to the initial image based on the area where the target object is located in the initial image; segmenting the initial image based on the mask image to obtain a foreground image and a background image, where the foreground image is used to represent the image corresponding to the area where the object model is located in the initial image, and the background image is used to represent the image corresponding to the other areas in the initial image except the foreground image; performing structural detection on the foreground image to obtain the structural image of the object model in the foreground image; and splicing the structural image and the background image to obtain the conditional image.

[0056] In an alternative solution of this embodiment, considering that in addition to the structural information presented by the target object in the target scene, such as posture and proportion, etc., which affect the appearance information presented by the target object, the environmental information in the target scene, such as the overall lighting conditions of the target scene, the spatial layout of other objects, etc., will also largely affect the visual effect presented by the target object in the target image. Therefore, in the constructed conditional image, in addition to including the structural information of the target object, the environmental information of the target scene can also be introduced, so as to ensure the accuracy of the image generation system when generating the target image guided by the conditional image. Based on this, when constructing the conditional image, the image generation system can first determine the area where the target object is located in the initial image, and generate a mask image corresponding to the initial image according to this area. Through the mask image, the image generation system can divide the initial image into two parts: a foreground image and a background image. Among them, the foreground image can refer to the image corresponding to the area where the object model is located in the initial image. Through the foreground image, the image generation system can accurately determine the structural information presented by the target object in the initial image. The background image can refer to the image corresponding to the other areas in the initial image except the foreground image. Through the background image, the image generation system can accurately obtain the environmental information of the target scene. Based on this, by performing structural detection on the foreground image, such as processing the foreground image through the Canny edge detection algorithm, the image generation system can obtain the structural image of the object model in the foreground image, such as a Canny line drawing. By splicing the structural image and the background image, the image generation system can further obtain a conditional image that can accurately highlight and display the edge and contour information of the target object, and at the same time reflect the environmental information of the target scene, thereby ensuring the accuracy of the constructed conditional image and the accuracy of the subsequent target image generated using the conditional image.

[0057] In the embodiment of the present application, the target image includes: the object layer where the target object is located after adding the target object to the target scene; generating the target image by fusing the conditional image and the object image, including: fusing the conditional image and the object image to obtain a fused image, where the fused image is used to represent the image obtained after adding the target object to the target scene, and the visual parameters presented by the fused image are the same as the visual parameters presented by the scene image; segmenting the fused image based on the mask image corresponding to the initial image to obtain a segmented image, where the segmented image includes the target object; adjusting the display parameters of different regions in the segmented image to obtain the object layer of the target object, where the different regions in the segmented image include: the region where the target object is located, and the other regions except the target object.

[0058] The above display parameters may include, but are not limited to, parameters such as transparency, contrast, saturation, etc. For the convenience of constructing layers, the following takes transparency as an example for explanation.

[0059] In an alternative solution of this embodiment, in order to ensure that the user can conveniently view the visual effects that the target object can present in the target scene and flexibly control parameters such as the position, size, and lighting of the target object in the target scene, the generated target image may at least include an object layer where the target object is located after adding the target object to the target scene. Through the object layer, the user can independently control parameters such as the position, size, angle, and lighting of the commodity without affecting the target scene, so that the user can easily adjust the placement of the commodity, compare the effects of different objects in the same scene, and at the same time, the image generation system does not need to regenerate the entire scene image, improving the efficiency and experience of user interaction.

[0060] Based on this, when generating the object layer, the image generation system can first fuse the constructed conditional image and object image to introduce the appearance information of the target object in the object image based on the structural information of the target object in the conditional image, obtaining a fused image that can reflect the appearance information matching the structural information of the target object and the environmental information of the target scene. Among them, the visual parameters presented by the fused image, such as lighting conditions, size ratios between objects, etc., are the same as those presented by the scene image. Considering that the fused image is the image obtained after adding the target object to the target scene, the fused image may not only include the target object but also other objects in the target scene. Therefore, in order to obtain an accurate object layer, after obtaining the fused image, the image generation system can further segment the fused image according to the mask image corresponding to the initial image to segment the area where the target object is located from the fused image, obtaining the above-mentioned segmented image, and then adjust the display parameters of different regions in the segmented image. For example, keep the transparency of the foreground area where the target object is located in the segmented image unchanged, and set the transparency of the background area except the target object to completely transparent to accurately obtain the object layer where the target object is located.

[0061] In the embodiment of the present application, fusing the conditional image and the object image to obtain a fused image includes: encoding the conditional image based on a variational encoder to obtain conditional features of the conditional image; encoding the initial image based on a visual encoder to obtain visual features of the target object; inputting the conditional features and visual features into a feature fusion model, and using the feature fusion model to fuse the conditional features and visual features to obtain fused features; decoding the fused features based on a variational decoder to obtain a fused image.

[0062] In an alternative solution of this embodiment, to ensure the fidelity of the fused image obtained by fusion, the image generation system may first encode the above-mentioned conditional image using a variational encoder, compress the conditional image into a latent space to become a low-dimensional latent variable, so as to obtain the conditional features of the conditional image, such as the structure and pose of the target object, the lighting conditions of the target scene, etc. At the same time, the visual encoder is used to encode the initial image to obtain the visual features of the target object, such as the texture, shape, color, etc. of the target object. Then, the conditional features and visual features are input into the feature fusion model to fuse the conditional features and visual features using the feature fusion model, so as to obtain relatively accurate fused features. Finally, the variational decoder is used to decode the fused features, and the image generation system can generate a fused image with relatively high fidelity.

[0063] In the embodiment of the present application, inputting the conditional features and visual features into the feature fusion model, and using the feature fusion model to fuse the conditional features and visual features to obtain fused features, including: encoding the visual features based on the model identifier of the feature fusion model to obtain the image encoding features corresponding to the visual features; adding random noise to the conditional features to obtain the noise-added features; performing cross-attention processing on the noise-added features, visual features, and image encoding features to obtain attention features; and denoising the attention features to obtain fused features.

[0064] In an alternative solution of this embodiment, when fusing the conditional features and visual features, the image generation system may first encode the visual features using the model identifier of the feature fusion model to convert the visual features into image encoding features that the feature fusion model can understand. At the same time, the image generation system may add random noise to the conditional features to obtain the above-mentioned noise-added features. After obtaining the image encoding features and noise-added features, the image generation system can perform cross-attention processing on the above-mentioned noise-added features, visual features, and image encoding features to obtain the corresponding attention features, and then denoise the attention features to obtain relatively accurate fused features.

[0065] In the embodiment of the present application, the above method further includes: obtaining multiple sample object images of the sample object from different perspectives, the sample scene image of the sample scene, and multiple placement images, where the multiple placement images are used to represent multiple images obtained by adding the sample object to the sample scene; based on the pose information presented by the sample object in the multiple placement images, adding the sample object to the sample scene to obtain multiple initial sample images; extracting features from the multiple sample condition images corresponding to the multiple initial sample images to obtain sample condition features corresponding to different sample condition images, extracting features from the multiple sample object images to obtain sample visual features corresponding to the sample object, and extracting features from the sample scene image to obtain sample scene features of the sample scene image; training an initial fusion model based on the sample condition features, sample visual features, and sample scene features to obtain a feature fusion model.

[0066] In an alternative solution of this embodiment, in order to ensure the accuracy of the trained feature fusion model, during the model training process, the image generation system can first obtain multiple sample object images of the sample object from different perspectives, the sample scene image of the sample scene, and multiple placement images, where these multiple placement images are multiple images obtained by adding the sample object to the sample scene. After obtaining these training samples, the image generation system can then follow the application process of the foregoing feature fusion model, and according to the pose information presented by the sample object in the multiple placement images, add the sample object to the sample scene to obtain multiple initial sample images, and then extract features from the multiple sample condition images corresponding to the multiple initial sample images to obtain sample condition features corresponding to different sample condition images, extract features from the multiple sample object images to obtain sample visual features corresponding to the sample object, extract features from the sample scene image to obtain sample scene features of the sample scene image, and finally use the sample condition features, sample visual features, and sample scene features to train the initial fusion model, so as to obtain a feature fusion model with higher accuracy.

[0067] In the embodiment of the present application, the above method further includes: in response to receiving an image generation instruction, identifying the instruction type of the image generation instruction; in response to the instruction type being a first preset type, reading model placement parameters from the image generation instruction; in response to the instruction type being a second preset type, obtaining the description text corresponding to the image generation instruction, and identifying the description text to obtain model placement parameters, where the description text is used to describe the position and pose of the placement object model.

[0068] The above first preset type may refer to the instruction type of an image generation instruction generated when the user actively adjusts the pose information of the target object in the target scenario. The above second preset type may refer to the instruction type of an image generation instruction triggered when the user cannot actively adjust the pose information of the target object in the target scenario, but instead inputs the pose information of the target object in the target scenario through other means, such as through a description text, a voice description, etc.

[0069] In an alternative solution of this embodiment, in order to accurately obtain the model placement parameters, when receiving an image generation instruction, the image generation system may identify the instruction type of the image generation instruction. When the instruction type is the above first preset type, the image generation system may directly read the corresponding model placement parameters from the image generation instruction to ensure the efficiency of obtaining the model placement parameters; when the instruction type is the above second preset type, the image generation system may first obtain the description text corresponding to the image generation instruction to determine information such as the pose and attitude of the placement object model through the description text, and then identify the description text to determine the above model placement parameters to ensure the accuracy of the obtained model placement parameters.

[0070] In the embodiment of the present application, the above method further includes: extracting features from the object image to obtain multiple key points of the target object; performing texture recognition on the object image to obtain the object texture of the target object; and constructing an object model based on the multiple key points and the object texture.

[0071] In an alternative solution of this embodiment, considering that the user providing the target object may not provide the object model of the target object at the same time, when obtaining the object model, the image generation system may directly extract features from the obtained object image to determine multiple key points of the target object, and at the same time perform texture recognition on the obtained object image to obtain the object texture of the target object, and then perform three-dimensional reconstruction on the target object based on the determined multiple key points and the object texture to construct the object model of the target object. In order to facilitate adding the object model of the target object to the target scenario, the image generation system may also convert the format of the constructed object model of the target object into a format that the rendering tool can process, such as GLB, FBX (Filmbox), etc., to improve the convenience of generating the target image.

[0072] For ease of understanding, Figure 3 is a schematic diagram of an image generation process shown according to an embodiment of the present application, as Figure 3As shown, when generating a target image, the image generation system can first construct a corresponding object model based on the obtained object image of the target object. The user can actively adjust the pose information of the object model in the scene image of the target scene to obtain a corresponding initial image. After obtaining the initial image, the image generation system can construct a conditional image and a mask image corresponding to the initial image, and based on the conditional image and the object image, fuse the conditional image and the object image through AI (Artificial Intelligence) technology to obtain a corresponding fused image. Finally, the fused image can be segmented using the mask image to segment out the object layer corresponding to the target object in the target scene.

[0073] Figure 4 FIG. is a schematic diagram of an image fusion process shown according to an embodiment of the present application. As Figure 4 shown, to ensure the accuracy when fusing the conditional image and the object image, the image generation system can first encode the conditional image using a variational encoder to obtain corresponding conditional features, and add random noise to the conditional features to obtain corresponding noise-added features. At the same time, the object image can be encoded using a visual encoder to obtain corresponding event features, and the visual features can be converted into image encoding features that the feature fusion model can accurately understand. Finally, the image generation system can input the noise-added features, visual features, and image encoding features into the feature fusion model for fusion to obtain corresponding fusion features, and decode the fusion features through a variational decoder to generate a corresponding fused image.

[0074] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.

[0075] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0076] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of the various embodiments of the present application.

[0077] According to an embodiment of the present application, another image generation method is also provided. Figure 5 It is a flowchart of another image generation method shown according to an embodiment of the present application, as Figure 5 shown. The method may include the following steps:

[0078] Step S502, in response to an input instruction acting on the operation interface, display an object image and an object model of a target object, and a scene image of a target scene on the operation interface.

[0079] Among them, the target object is to be added to the target scene.

[0080] Step S504, in response to an image generation instruction acting on the operation interface, display a target image on the operation interface.

[0081] Among them, the target image is obtained by fusing a conditional image and an object image. The conditional image is determined by the structural information presented by the object model in the initial image. The initial image is used to represent adding the object model to the scene image based on the model placement parameters.

[0082] In an alternative solution of this embodiment, when receiving an input instruction acting on the operation interface, the image generation system may first obtain and display an object image and an object model of the target object, and a scene image of the target scene on the operation interface. Among them, the target object may refer to an object to be added to the target scene. When receiving an image generation instruction acting on the operation interface, the image generation system may first add the object model to the scene image according to the model placement parameters to obtain a corresponding initial image, and construct a conditional image corresponding to the initial image according to the structural information presented by the object model in the initial image. Finally, the conditional image and the object image are fused to obtain a corresponding target image. The image generation system may display the target image in the operation interface for the user to view conveniently.

[0083] According to an embodiment of the present application, another image generation method is also provided. Figure 6 It is a flowchart of another image generation method shown according to an embodiment of the present application, asFigure 6 As shown in the figure, the method may include the following steps:

[0084] Step S602, obtaining the object image and object model of the target object, and the scene image of the target scene by calling the first interface.

[0085] Wherein, the first interface includes a first parameter, and the parameter value of the first parameter includes the object image, the object model, and the scene image, and the target object is to be added to the target scene.

[0086] Step S604, placing the object model into the scene image according to the model placement parameter to obtain an initial image.

[0087] Step S606, determining the conditional image corresponding to the initial image based on the structural information presented by the object model in the initial image.

[0088] Step S608, fusing the conditional image and the object image to generate a target image.

[0089] Step S610, outputting the target image by calling the second interface.

[0090] Wherein, the second interface includes a second parameter, and the parameter value of the second parameter includes the target image.

[0091] In an alternative solution of this embodiment, when generating the target image, the image generation system may first obtain the parameter value of the first parameter by calling the first interface, that is, obtain the object image and object model of the target object, and the scene image of the target scene, where the target object may refer to the object to be added to the target scene. Then, place the object model into the scene image according to the model placement parameter to obtain the corresponding initial image, and construct the conditional image corresponding to the initial image according to the structural information presented by the object model in the initial image. Then, fuse the conditional image and the object image to obtain the corresponding target image. Finally, output the parameter value of the second parameter by calling the second structure, that is, output the target image for the user to view.

[0092] According to an embodiment of the present application, there is also provided an image generation device for implementing the above image generation method. Figure 7 It is a structural block diagram of an image generation device shown according to an embodiment of the present application. As Figure 7 shown, the device includes: an acquisition module 702, a placement module 704, a determination module 706, and a fusion module 708.

[0093] Among them, the acquisition module 702 is used to acquire the object image and object model of the target object, as well as the scene image of the target scene, where the target object is to be added to the target scene; the placement module 704 is used to place the object model into the scene image according to the model placement parameters to obtain an initial image; the determination module 706 is used to determine the conditional image corresponding to the initial image based on the structural information presented by the object model in the initial image; the fusion module 708 is used to fuse the conditional image and the object image to generate a target image.

[0094] In the embodiment of the present application, the determination module 706 includes: a mask generation unit, configured to generate a mask image corresponding to the initial image based on the area where the target object is located in the initial image; a first segmentation unit, configured to segment the initial image based on the mask image to obtain a foreground image and a background image, where the foreground image is used to represent the image corresponding to the area where the object model is located in the initial image, and the background image is used to represent the image corresponding to the other areas in the initial image except the foreground image; a structure detection unit, configured to perform structure detection on the foreground image to obtain the structure image of the object model in the foreground image; an image splicing unit, configured to splice the structure image and the background image to obtain a conditional image.

[0095] In the embodiment of the present application, the target image includes: an object layer where the target object is located after adding the target object to the target scene; the fusion module 708 includes: an image fusion unit, configured to fuse the conditional image and the object image to obtain a fusion image, where the fusion image is used to represent the image obtained after adding the target object to the target scene, and the visual parameters presented by the fusion image are the same as the visual parameters presented by the scene image; a second segmentation unit, configured to segment the fusion image based on the mask image corresponding to the initial image to obtain a segmented image, where the segmented image includes the target object; an image adjustment unit, configured to adjust the display parameters of different regions in the segmented image to obtain the object layer of the target object, where the different regions in the segmented image include: the region where the target object is located, and the other regions except the target object.

[0096] In the embodiment of the present application, the image fusion unit is further configured to: encode the conditional image based on a variational encoder to obtain the conditional features of the conditional image; encode the initial image based on a visual encoder to obtain the visual features of the target object; input the conditional features and the visual features into a feature fusion model, and use the feature fusion model to fuse the conditional features and the visual features to obtain fusion features; decode the fusion features based on a variational decoder to obtain a fusion image.

[0097] In the embodiment of the present application, the image fusion unit is further configured to: encode the visual features based on the model identifier of the feature fusion model to obtain image encoding features corresponding to the visual features; add random noise to the conditional features to obtain noise-added features; perform cross-attention processing on the noise-added features, visual features, and image encoding features to obtain attention features; and perform denoising processing on the attention features to obtain fusion features.

[0098] In the embodiment of the present application, the above device further includes: a sample acquisition module, configured to acquire multiple sample object images of a sample object from different perspectives, a sample scene image of a sample scene, and multiple placement images, where the multiple placement images are used to represent multiple images obtained by adding the sample object to the sample scene; a sample addition module, configured to add the sample object to the sample scene based on the pose information presented by the sample object in the multiple placement images to obtain multiple initial sample images; a sample extraction module, configured to extract features from multiple sample conditional images corresponding to the multiple initial sample images to obtain sample conditional features corresponding to different sample conditional images, extract features from the multiple sample object images to obtain sample visual features corresponding to the sample object, and extract features from the sample scene image to obtain sample scene features of the sample scene image; and a model training module, configured to train an initial fusion model based on the sample conditional features, sample visual features, and sample scene features to obtain a feature fusion model.

[0099] In the embodiment of the present application, the above device further includes: an instruction recognition module, configured to recognize the instruction type of an image generation instruction in response to receiving the image generation instruction; a parameter reading module, configured to read model placement parameters from the image generation instruction in response to the instruction type being a first preset type; and a parameter recognition module, configured to obtain a description text corresponding to the image generation instruction and recognize the description text to obtain model placement parameters in response to the instruction type being a second preset type, where the description text is used to describe the position and pose of a placement object model.

[0100] In the embodiment of the present application, the above device further includes: a key point extraction module, configured to extract features from an object image to obtain multiple key points of a target object; a texture recognition module, configured to perform texture recognition on the object image to obtain the object texture of the target object; and a model construction module, configured to construct an object model based on the multiple key points and the object texture.

[0101] It should be noted here that the above-mentioned acquisition module 702, placement module 704, determination module 706, and fusion module 708 correspond to steps S202 to S208 in the above-mentioned embodiment. The examples and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the content disclosed in the above-mentioned embodiment. It should be noted that the above-mentioned module or unit may be a hardware component or a software component stored in the memory and processed by one or more processors. The above-mentioned module may also be part of the device and may run in the server 10 provided in the above-mentioned embodiment.

[0102] According to an embodiment of the present application, there is also provided another image generation device for implementing the above-mentioned image generation method. Figure 8 It is a structural block diagram of another image generation device shown according to an embodiment of the present application. As Figure 8 shown, the device includes: a first execution module 802 and a second execution module 804.

[0103] Among them, the first execution module 802 is used to respond to an input instruction acting on the operation interface, and display an object image and an object model of a target object, and a scene image of a target scene on the operation interface, where the target object is to be added to the target scene; the second execution module 804 is used to respond to an image generation instruction acting on the operation interface, and display a target image on the operation interface, where the target image is obtained by fusing a conditional image and an object image, and the conditional image is determined by the structural information presented by the object model in the initial image, and the initial image is used to represent adding the object model to the scene image based on the model placement parameter.

[0104] It should be noted here that the above-mentioned first execution module 802 and second execution module 804 correspond to steps S502 to S502 in the above-mentioned embodiment. The examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the content disclosed in the above-mentioned embodiment. It should be noted that the above-mentioned module or unit may be a hardware component or a software component stored in the memory and processed by one or more processors. The above-mentioned module may also be part of the device and may run in the server 10 provided in the above-mentioned embodiment.

[0105] According to an embodiment of the present application, there is also provided another image generation device for implementing the above-mentioned image generation method. Figure 9 It is a structural block diagram of another image generation device shown according to an embodiment of the present application. As Figure 9 shown, the device includes: a first call module 902, a model placement module 904, an image determination module 906, an image fusion module 908, and a second call module 910.

[0106] Among them, the first calling module 902 is used to obtain the object image and object model of the target object, and the scene image of the target scene by calling the first interface. The first interface includes a first parameter, and the parameter value of the first parameter includes the object image, the object model, and the scene image. The target object is to be added to the target scene. The model placement module 904 is used to place the object model into the scene image according to the model placement parameters to obtain an initial image. The image determination module 906 is used to determine the conditional image corresponding to the initial image based on the structural information presented by the object model in the initial image. The image fusion module 908 is used to fuse the conditional image and the object image to generate a target image. The second calling module 910 is used to output the target image by calling the second interface. The second interface includes a second parameter, and the parameter value of the second parameter includes the target image.

[0107] It should be noted here that the above first calling module 902, model placement module 904, image determination module 906, image fusion module 908, and second calling module 910 correspond to steps S602 to S610 in the above embodiment. The instances and application scenarios implemented by the five modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment. It should be noted that the above modules or units may be hardware components or software components stored in the memory and processed by one or more processors. The above modules may also be part of a device and may run in the server 10 provided in the above embodiment.

[0108] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in the above embodiments, but are not limited to the schemes provided in the above embodiments.

[0109] An embodiment of the present application may provide a computing device. Figure 10 It is a structural block diagram of a computing device shown according to an embodiment of the present application. As shown in the figure, the computing device 1000 may include: one or more (only one is shown in the figure) processors 1002, a memory 1004, a storage controller, and a peripheral interface.

[0110] The above computing device can be understood as an integrated intelligent terminal, including but not limited to a server, a desktop computer, a PC (Personal Computer), a model all-in-one machine, etc. And, the above model in the above embodiments of the present application may be preset in the computing device.

[0111] Specifically, the computing device can pre - set multiple types of models, including but not limited to models in the fields of natural language processing, visual processing, speech processing, code processing, multi - modal task processing, etc., so as to provide diverse model selection. In different product forms, the computing device can support one or more model usage methods, including but not limited to model training, model invocation, model fine - tuning, model deployment, model inference and application, etc. In some product forms, the computing device also supports model management, including but not limited to multi - type model management (supporting the management of discriminative, generative and other types of models), model version control (supporting the control of different model versions), model evaluation (evaluating the performance and effect of the model based on model evaluation tools), etc. In other product forms, the computing device can also create applications based on models, provide API invocation capabilities, and can call the model into the created application through the API interface. At the same time, an application management tool is provided to realize the management and monitoring of the application.

[0112] Furthermore, the computing device can also include data management (supporting the creation and management of model tuning data sets), a training center (providing rich training resources to help users learn and master AI technologies), and basic control capabilities (providing enterprise - level basic control capabilities to ensure the security and efficient operation of the system). Through the above functions, a comprehensive and integrated AI development, training, deployment, and application device is provided.

[0113] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, to implement the methods in the above - mentioned embodiments. The memory can include high - speed random access memory, and can also include non - volatile memory, such as one or more magnetic storage devices, flash memory, or other non - volatile solid - state memories. In some instances, the memory can further include memories remotely set relative to the processor, and these remote memories can be connected to terminal A through a network. Examples of the above - mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.

[0114] The processor can call the executable program stored in the memory through the transmission device to execute the method of any one of the above - mentioned embodiments.

[0115] Embodiments of the present application can provide an electronic device. Figure 11 It is a structural block diagram of an electronic device shown according to an embodiment of the present application. As Figure 11As shown in the figure, the electronic device may include: an input / output device 1102; a memory 1104 and a processor 1106, wherein the processor 1106 is connected to the input / output device 1102 and the memory 1104 through a bus 1108.

[0116] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the methods in the above embodiments. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely set relative to the processor, and these remote memories can be connected to the terminal A through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0117] The processor can call the executable program stored in the memory through a transmission device to execute the method of any one of the above embodiments.

[0118] Those of ordinary skill in the art can understand that the structure shown Figure 11 is only schematic. The computing device can also be a terminal device such as a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and a mobile Internet device (Mobi leInternet Devices, MID), a PAD, etc. The Figure 11 does not limit the structure of the above computing device. For example, the computing device 1000 may further include more or fewer components (such as a network interface, a display device, etc.) than those shown Figure 11 in the figure, or have a different configuration from that shown Figure 11 in the figure.

[0119] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a read-only memory (Read-Only Memory, ROM), a random access memory (RandomAccess Memory, RAM), a magnetic disk, or an optical disc, etc.

[0120] The embodiments of the present application also provide a computer-readable storage medium. Optionally, in this embodiment, the above computer-readable storage medium can be used to save the program code executed by the method provided in the above embodiment.

[0121] Optionally, in this embodiment, the above storage medium may be located in a computing device.

[0122] Optionally, in this embodiment, the computer-readable storage medium is configured to store an executable program, and when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the method according to any one of the above embodiments.

[0123] An embodiment of the present application further provides a computer program product. Optionally, in this embodiment, the above computer program product may include a computer program, and when the computer program is executed by a processor, it implements the method provided by the above embodiment.

[0124] An embodiment of the present application further provides a computer program product. Optionally, the above computer program product may include a non-volatile computer-readable storage medium, which can be used to store a computer program, and when the computer program is executed by a processor, it implements the method provided by the above embodiment.

[0125] An embodiment of the present application further provides a computer program. Optionally, in this embodiment, when the above computer program is executed by a processor, it implements the method provided by the above embodiment.

[0126] In the above embodiments of the present application, the descriptions of each embodiment have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0127] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0128] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0129] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0130] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0131] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. An image generation method, characterized in that, Including: Obtaining an object image and an object model of a target object, and a scene image of a target scene, where the target object is to be added to the target scene; Placing the object model into the scene image according to model placement parameters to obtain an initial image; Determining a conditional image corresponding to the initial image based on the structural information presented by the object model in the initial image; Fusing the conditional image and the object image to generate a target image.

2. The method according to claim 1, wherein The determining the conditional image corresponding to the initial image based on the structural information presented by the object model in the initial image includes: Generating a mask image corresponding to the initial image based on the area where the target object is located in the initial image; Segmenting the initial image based on the mask image to obtain a foreground image and a background image, where the foreground image is used to represent the image corresponding to the area where the object model is located in the initial image, and the background image is used to represent the image corresponding to the other areas in the initial image except the foreground image; Performing structure detection on the foreground image to obtain a structure image of the object model in the foreground image; Stitching the structure image and the background image to obtain the conditional image.

3. The method according to claim 1, wherein The target image includes: an object layer where the target object is located after adding the target object to the target scene; the fusing the conditional image and the object image to generate a target image includes: Fusing the conditional image and the object image to obtain a fused image, where the fused image is used to represent the image obtained after adding the target object to the target scene, and the visual parameters presented by the fused image are the same as the visual parameters presented by the scene image; Segmenting the fused image based on the mask image corresponding to the initial image to obtain a segmented image, where the segmented image includes the target object; Adjusting the display parameters of different regions in the segmented image to obtain the object layer corresponding to the target object, where the different regions in the segmented image include: the region where the target object is located, and the other regions except the target object.

4. The method according to claim 3, wherein The fusing the conditional image and the object image to obtain a fused image includes: Encoding the conditional image based on a variational encoder to obtain conditional features of the conditional image; Encoding the initial image based on a visual encoder to obtain visual features of the target object; Inputting the conditional features and the visual features into a feature fusion model, and using the feature fusion model to fuse the conditional features and the visual features to obtain fused features; Decoding the fused features based on a variational decoder to obtain the fused image.

5. The method according to claim 4, characterized in that The inputting the conditional features and the visual features into a feature fusion model, and using the feature fusion model to fuse the conditional features and the visual features to obtain fused features includes: Encoding the visual features based on the model identification of the feature fusion model to obtain the image encoding features corresponding to the visual features; Adding random noise to the conditional features to obtain noise-added features; Performing cross-attention processing on the noise-added features, the visual features, and the image encoding features to obtain attention features; Performing denoising processing on the attention features to obtain the fusion features.

6. The method according to claim 4, characterized in that, The method further includes: Obtaining multiple sample object images of the sample object from different perspectives, a sample scene image of the sample scene, and multiple placement images, where the multiple placement images are used to represent multiple images obtained by adding the sample object to the sample scene; Based on the pose information presented by the sample object in the multiple placement images, adding the sample object to the sample scene to obtain multiple sample initial images; Performing feature extraction on multiple sample conditional images corresponding to the multiple sample initial images to obtain sample conditional features corresponding to different sample conditional images, performing feature extraction on the multiple sample object images to obtain sample visual features corresponding to the sample object, and performing feature extraction on the sample scene image to obtain sample scene features of the sample scene image; Training an initial fusion model based on the sample conditional features, the sample visual features, and the sample scene features to obtain the feature fusion model.

7. The method according to any one of claims 1-6, characterized in that, The method further includes: Responding to receiving an image generation instruction, identifying the instruction type of the image generation instruction; Responding to the instruction type being a first preset type, reading the model placement parameters from the image generation instruction; Responding to the instruction type being a second preset type, obtaining the description text corresponding to the image generation instruction, and identifying the model placement parameters from the description text, where the description text is used to describe the position and pose of placing the object model.

8. The method according to any one of claims 1-6, characterized in that, The method further includes: Performing feature extraction on the object image to obtain multiple key points of the target object; Performing texture recognition on the object image to obtain the object texture of the target object; Constructing the object model based on the multiple key points and the object texture.

9. An image generation method, characterized in that, Including: Responding to an input instruction acting on the operation interface, displaying an object image and an object model of the target object, and a scene image of the target scene on the operation interface, where the target object is to be added to the target scene; Responding to an image generation instruction acting on the operation interface, displaying a target image on the operation interface, where the target image is obtained by fusing a conditional image and the object image, and the conditional image is determined by the structural information presented by the object model in an initial image, and the initial image is used to represent adding the object model to the scene image based on the model placement parameters.

10. An image generation method, characterized in that, Including: Obtain the object image and object model of the target object, as well as the scene image of the target scene by invoking a first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter includes the object image, the object model, and the scene image, and the target object is to be added to the target scene; Place the object model into the scene image according to the model placement parameter to obtain an initial image; Determine the conditional image corresponding to the initial image based on the structural information presented by the object model in the initial image; Fuse the conditional image and the object image to generate a target image; Output the target image by invoking a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the target image.

11. A computing device, characterized in that, Comprising: A memory storing an executable program; A processor for running the program, wherein when the program runs, it executes the method according to any one of claims 1 to 10.

12. An electronic device, characterized in that, Comprising: A memory storing an executable program; A processor connected to the memory through a bus for running the program, wherein when the program runs, it executes the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the method according to any one of claims 1 to 10.

14. A computer program product, characterized in that, Comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 10.