Image generation method and device, program product, storage medium and electronic equipment
The method improves the accuracy of virtual model image generation by using a vision-language model to generate contextually appropriate poses through text descriptions and refined masks, addressing the unnatural poses in existing AI model platforms.
Patent Information
- Application Number
- CN202510407017.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-15
AI Technical Summary
The existing AI model drawing platform has the problem of low image accuracy when generating images, especially when the user-specified model poses exceed the scope of the pose library, the generated image may be unnaturally inconsistent with the original image content.
By obtaining the initial image, the posture image and attribute information of the virtual model, the visual language model and the image generation model are used for processing, and the target image containing the virtual model is generated. The specific steps include generating a mask image, laminating the mask image and the initial image, combining the posture image and attribute information, using the image generation model to repair the image, and finally generating a natural and reasonable model pose.
The compatibility between model poses and initial image content is improved, and the unreasonable pose generation problem caused by lack of context understanding in traditional methods is avoided, thereby achieving higher image generation accuracy.
Smart Images

Figure CN120318356A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and more specifically, to an image generation method, device, program product, storage medium and electronic device. Background Art
[0002] With the continuous development of artificial intelligence technology, artificial intelligence (AI) models are being used more and more widely, such as in commodity sales and demonstration of product usage methods. AI technology can efficiently generate visual content that is both beautiful and in line with product characteristics through intelligent means, thereby reducing production costs compared to traditional display methods that rely on real-life photography.
[0003] At present, the existing AI model drawing platform provides a limited model pose library. In the process of AI model drawing, users rely on the poses provided in the model pose library to generate AI models. When the model pose specified by the user exceeds the range of the pose library, the generated image may have an unnatural pose and be inconsistent with the original image content, resulting in the problem of low accuracy in generating images containing AI models.
[0004] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0005] The embodiments of the present application provide a method, device, program product, storage medium and electronic device for generating an image, so as to at least solve the technical problem of low image generation accuracy when generating an image containing a virtual model in the related art.
[0006] According to one aspect of an embodiment of the present application, a method for generating an image is provided, comprising: obtaining an initial image, a posture image of a virtual model to be drawn in the initial image, and attribute information of the virtual model, wherein the virtual model is used to assist in displaying a display object in the initial image; processing the initial image, the posture image, and the attribute information through a visual language model to obtain text information, wherein the text information is used to describe the posture of the virtual model; processing the text information, the attribute information, the posture image, and the initial image through an image generation model to obtain a target image, wherein the target image includes the virtual model.
[0007] Furthermore, the initial image, the posture image and the attribute information are processed by the visual language model to obtain the text information, including: generating a first mask image according to the posture image, wherein the first mask image is used to represent the regional image of the virtual model in the initial image; fitting the first mask image and the initial image to obtain the first image; and inputting the first image and the attribute information into the visual language model to obtain the text information.
[0008] Further, generating the first mask image based on the pose image includes: supplementing the morphological features of the virtual model in the pose image based on a preset morphological processing rule to obtain an object morphology map; performing dilation processing on the image content in the object morphology map to obtain the first mask image.
[0009] Further, the pose image of the virtual model is obtained through the following steps: obtaining a pose reference image, where the pose reference image includes a reference virtual model presenting an initial pose; extracting bone joint points from the reference virtual model in the pose reference image to obtain bone joint point information; determining the pose image of the virtual model according to the bone joint point information.
[0010] Further, determining the pose image of the virtual model according to the bone joint point information includes: rendering the bone joint points in the initial image displayed on the human-computer interaction interface according to the bone joint point information; responding to the position adjustment operation acting on the human-computer interaction interface to adjust the positions of the bone joint points to obtain target bone joint points; determining the pose image of the virtual model according to the target bone joint point information of the target bone joint points.
[0011] Further, processing the text information, attribute information, pose image, and initial image through an image generation model to obtain a target image includes: generating a first mask image according to the pose image, where the first mask image is used to represent the regional image of the virtual model in the initial image; inputting the text information, attribute information, pose image, first mask image, and initial image into the image generation model to obtain a second image, where the second image includes an initial virtual model; determining the target image according to the second image.
[0012] Further, determining the target image according to the second image includes: performing segmentation processing on the regional image of the initial virtual model in the second image to obtain a second mask image, where the second mask image is used to represent the target regional image of the virtual model in the initial image; processing the text information, attribute information, pose image, second mask image, and initial image through the image generation model to obtain the target image.
[0013] Further, processing the text information, attribute information, pose image, second mask image, and initial image through the image generation model to obtain the target image includes: inputting the text information, attribute information, pose image, second mask image, and initial image into the image generation model to obtain a third image; performing restoration processing on the target regional image in the third image through an image restoration model to obtain an updated target regional image, where the target regional image refers to the regional image including the virtual model; superimposing the updated target regional image on the initial image to obtain the target image.
[0014] According to another aspect of the embodiments of the present application, there is also provided a method for generating an image, including: obtaining an initial image uploaded by a client, a pose image of a virtual model to be drawn in the initial image, and attribute information of the virtual model, where the virtual model is used to assist in displaying a display object in the initial image; processing the initial image, the pose image, and the attribute information through a vision-language model in a cloud server to obtain text information, where the text information is used to describe the pose of the virtual model; processing the text information, the attribute information, the pose image, and the initial image through an image generation model to obtain a target image, where the target image includes the virtual model; and feeding back the target image to the client.
[0015] According to another aspect of the embodiments of the present application, there is also provided an image generation device, including: a first acquisition unit, configured to obtain an initial image, a pose image of a virtual model to be drawn in the initial image, and attribute information of the virtual model, where the virtual model is used to assist in displaying a display object in the initial image; a first processing unit, configured to process the initial image, the pose image, and the attribute information through a vision-language model to obtain text information, where the text information is used to describe the pose of the virtual model; and a second processing unit, configured to process the text information, the attribute information, the pose image, and the initial image through an image generation model to obtain a target image, where the target image includes the virtual model.
[0016] Further, the first processing unit includes: a first generation subunit, configured to generate a first mask image according to the pose image, where the first mask image is used to represent the regional image of the virtual model in the initial image; a first processing subunit, configured to perform a fitting process on the first mask image and the initial image to obtain a first image; and a second processing subunit, configured to input the first image and the attribute information into the vision-language model to obtain text information.
[0017] Further, the first generation subunit includes: a first processing module, configured to supplement the morphological features of the virtual model in the pose image based on a preset morphological processing rule to obtain an object morphology map; and a second processing module, configured to perform a dilation process on the image content in the object morphology map to obtain the first mask image.
[0018] Further, the image generation device further includes: a second acquisition unit, configured to obtain a pose reference image, where the pose reference image includes a reference virtual model presenting an initial pose; an extraction unit, configured to extract skeletal joint point information from the reference virtual model in the pose reference image; and a determination unit, configured to determine the pose image of the virtual model according to the skeletal joint point information.
[0019] Further, the determination unit includes: a third processing subunit, configured to render bone joint points in the initial image displayed on the human-computer interaction interface according to the bone joint point information; an adjustment subunit, configured to respond to a position adjustment operation applied to the human-computer interaction interface and adjust the positions of the bone joint points to obtain target bone joint points; and a first determination subunit, configured to determine a pose image of the virtual model according to the target bone joint point information of the target bone joint points.
[0020] Further, the second processing unit includes: a second generation subunit, configured to generate a first mask image according to the pose image, where the first mask image is used to represent the regional image of the virtual model in the initial image; a fourth processing subunit, configured to input the text information, attribute information, pose image, first mask image, and initial image into an image generation model to obtain a second image, where the second image includes an initial virtual model; and a second determination subunit, configured to determine a target image according to the second image.
[0021] Further, the second determination subunit includes: a third processing module, configured to perform segmentation processing on the regional image of the initial virtual model in the second image to obtain a second mask image, where the second mask image is used to represent the target regional image of the virtual model in the initial image; and a fourth processing module, configured to process the text information, attribute information, pose image, second mask image, and initial image through the image generation model to obtain a target image.
[0022] Further, the fourth processing module includes: a first processing sub-module, configured to input the text information, attribute information, pose image, second mask image, and initial image into the image generation model to obtain a third image; a second processing sub-module, configured to perform restoration processing on the target regional image in the third image through an image restoration model to obtain an updated target regional image, where the target regional image refers to the regional image including the virtual model; and a third processing sub-module, configured to superimpose the updated target regional image on the initial image to obtain a target image.
[0023] According to another aspect of the embodiments of the present invention, an electronic device is further provided, including: a memory storing an executable program; and a processor configured to run the program, where when the program runs, it executes the image generation method of any one of the above.
[0024] According to another aspect of the embodiments of the present invention, a computer-readable storage medium is further provided, where the storage medium stores a program, and when the program runs, it controls the device where the storage medium is located to execute the image generation method of any one of the above.
[0025] According to another aspect of the embodiments of the present invention, there is also provided a computer program product, including a computer program which, when executed by a processor, implements the method for generating an image according to any one of the above.
[0026] In the embodiments of the present application, by obtaining an initial image, a pose image of a virtual model to be drawn in the initial image, and attribute information of the virtual model, wherein the virtual model is used to assist in displaying a display object in the initial image; by processing the initial image, the pose image, and the attribute information through a vision-language model to obtain text information, wherein the text information is used to describe the pose of the virtual model; by processing the text information, the attribute information, the pose image, and the initial image through an image generation model to obtain a target image, wherein the target image includes the virtual model, the method of using the vision-language model to process the initial image, the pose image of the virtual model, and the attribute information of the virtual model to obtain text information realizes the effective determination of text conditions for intelligently inferring the action poses that the virtual model should present in a specific scenario. These text conditions can accurately describe the action intentions of the virtual model and its interaction relationship with the content of the initial image, providing an accurate guiding basis for generating natural and reasonable model poses. Therefore, when using the image generation model to process the text information, the attribute information, the pose image, and the initial image to obtain the target image, the degree of fit between the model pose and the content of the initial image can be improved, avoiding the problem of generating unreasonable poses due to the lack of context understanding in traditional methods, thereby effectively improving the accuracy of image generation. The present application achieves the purpose of generating an image including a virtual model by combining text conditions related to the model pose, thereby realizing the technical effect of improving the accuracy of generating an image including a virtual model, and further solving the technical problem of low accuracy in generating an image including a virtual model in the related art. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0028] Figure 1 is a schematic diagram of a computer terminal according to Embodiment 1 of the present application;
[0029] Figure 2 is a flowchart of the method for generating an image according to Embodiment 1 of the present application Figure 1 ;
[0030] Figure 3 is a schematic diagram of generating text information according to Embodiment 1 of the present application;
[0031] Figure 4Schematic diagram of the target image generated in the prior art;
[0032] Figure 5 Schematic diagram of the target image provided in the first embodiment of the present application Figure 1 ;
[0033] Figure 6 Schematic diagram of generating the first mask image provided in the first embodiment of the present application;
[0034] Figure 7 Schematic diagram of generating the pose image provided in the first embodiment of the present application;
[0035] Figure 8 Schematic diagram of the pose estimation network provided in the first embodiment of the present application;
[0036] Figure 9 Schematic diagram of the editing area provided in the first embodiment of the present application;
[0037] Figure 10 Schematic diagram of determining the second image provided in the first embodiment of the present application;
[0038] Figure 11 Schematic diagram of determining the second mask image provided in the first embodiment of the present application;
[0039] Figure 12 Schematic diagram of determining the third image provided in the first embodiment of the present application;
[0040] Figure 13 Schematic diagram of the target image provided in the first embodiment of the present application Figure 2 ;
[0041] Figure 14 Schematic diagram of the target image provided in the first embodiment of the present application Figure 3 ;
[0042] Figure 15 Schematic diagram of determining the target image provided in the first embodiment of the present application;
[0043] Figure 16 Flowchart of the method for generating an image provided in the first embodiment of the present application Figure 2 ;
[0044] Figure 17 Flowchart of the method for generating an image provided in the second embodiment of the present application;
[0045] Figure 18 Schematic diagram of the image generation device provided in the third embodiment of the present application;
[0046] Figure 19It is a structural block diagram of an electronic device provided according to Embodiment 4 of the present application. Detailed implementation manners
[0047] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0048] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data used in appropriate cases can be interchanged so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0049] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards in the relevant regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0050] First, some nouns or terms that appear during the description of the embodiments of the present application are applicable to the following explanations:
[0051] Vision-Language Model (VLM): A multi-modal artificial intelligence model that combines computer vision and natural language processing technologies. It can simultaneously process visual information such as images and videos, as well as language information such as text and speech, so as to achieve cross-modal understanding and generation tasks.
[0052] Diffusion model: A continuous-time Markov process that learns the data distribution by gradually adding random noise to a clear image, and then learns how to reverse this process, that is, gradually restore the clear image from the noise.
[0053] Mask: In image processing, a mask refers to a marking map used to specify the area in an image that needs to be processed or modified. In this solution, the mask is used to determine the redrawing area of the model.
[0054] Human Segmentation Model: A technical model specifically used to extract human parts from images, capable of accurately separating the human body area and providing a basis for subsequent processing.
[0055] Inpainting Model: A model used to repair or supplement missing areas in an image. In this solution, it is used to generate the model and subsequent image repair in the image.
[0056] Embodiment 1
[0057] According to an embodiment of the present application, a method for generating an image is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0058] The method embodiment provided by the first embodiment of the present application can be executed in a mobile terminal, a computer terminal, or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing the method for generating an image is shown. As Figure 1 shown, the computer terminal (or mobile device) 10 may include a set of processors 102 (the set of processors 102 may include, but is not limited to, processing devices such as a microcontroller unit (MCU) or a field programmable gate array (FPGA), and the set of processors 102 may include a set of processors, Figure 1 where 102a, 102b,..., 102n are used to show), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that, Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown, or have a structure different from Figure 1The different configurations shown.
[0059] It should be noted that one or more of the above-mentioned processors 102 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computer terminal 10 (or mobile device).
[0060] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage devices corresponding to the image generation method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned image generation method. The memory 104 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 can further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0061] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network can include the wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0062] The display can be a touch-screen liquid crystal display, which enables the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0063] With the continuous development of artificial intelligence technology, the application of artificial intelligence (AI) models is becoming more and more widespread. For example, for product sales, product usage method demonstration, etc. AI technology can efficiently generate visually appealing and product-feature-compliant visual content through intelligent means, reducing production costs compared to traditional display methods that rely on physical photography.
[0064] At present, the existing AI model drawing platform provides a limited model pose library. In the process of AI model drawing, users rely on the poses provided in the model pose library to generate AI models. When the model pose specified by the user exceeds the range of the pose library, the generated image may have an unnatural pose and be inconsistent with the original image content, resulting in the problem of low accuracy in generating images containing AI models.
[0065] In the above technical background, the present application provides Figure 2 The method for generating the image shown. Figure 2 The process of the image generation method provided in the first embodiment of the present application is as follows Figure 1 The method comprises:
[0066] Step S201 , obtaining an initial image, a posture image of a virtual model to be drawn in the initial image, and attribute information of the virtual model, wherein the virtual model is used to assist in displaying a display object in the initial image.
[0067] Optionally, electronic devices, application systems, servers and other devices may be used as the execution subject of the present application. In this embodiment, the target processing system is used as the execution subject to execute the above-mentioned image generation method.
[0068] Optionally, the initial image may refer to an initial scene image, and the image content of the initial image may include at least one display object and a related background (or application scene of the display object) for displaying the corresponding display object. For example, if the display object is a sofa, the related background may be a room, which may also include related furnishings commonly found in the room (e.g., carpets, coffee tables, bookshelves, and vases, etc.). The initial image does not include an AI model (i.e., a virtual model) used to assist in displaying the display object. If the application scenario targeted by the image generation method is an e-commerce scenario, the display object is a commodity (also referred to as a product) to be promoted or sold in the scenario, and the initial image containing the commodity may be a promotion scene map or a sales display map, which may be used to better display the effect of the commodity to be promoted. If the application scenario targeted by the image generation method is a product usage method display scenario, the display object is a product whose usage method is to be explained in the scenario, and the initial image containing the product may be a reference map in the product instruction manual, which may be used to better display the usage of the commodity.
[0069] Optionally, taking the e-commerce scenario as an example of the application scenario targeted by the image generation method, an exemplary description is given of the image content of the initial image. For example, if the target is interior design, the initial image can be an indoor space, and the display objects can be indoor furniture, lighting fixtures, and decorations, etc., to display the interior decoration effects of different styles and design concepts. If the target is a car advertisement, the initial image can be a road or a parking lot, and the display objects can be cars of different brands and models, to display the appearance, interior, and performance of the cars. If the target is clothing matching, the initial image can be a street or a shopping mall, and the display objects can be clothes of different styles and colors, to display the matching effects and styles of the clothes. If the target is food and beverage, the initial image can be a restaurant or a coffee shop, and the display objects can be different types of food and drinks, to display the appearance and taste of different foods.
[0070] It should be noted that the application scenarios of the above image generation method, the different situations targeted under different application scenarios, and the different initial images and display objects targeted under different situations are only for illustrative purposes and are not specifically limited here. In the embodiments of the present application, in different application scenarios, the types of the initial image and the display object can be selected according to specific needs and purposes to display different products, services, or concepts.
[0071] Optionally, the virtual model is an AI model. The model type of the virtual model can be a human, an animal, etc. The pose image of the virtual model is used to represent the pose information of the virtual model to be drawn. For example, the pose image of the virtual model can include the skeletal pose of the virtual model. The skeletal pose includes multiple skeletal joint points and the skeletal connections between the skeletal joint points, so that the pose of the virtual model can be represented based on the positional relationship between the multiple skeletal joint points.
[0072] Optionally, the attribute information of the virtual model can be information used to describe the appearance characteristics of the virtual model. For example, the attribute information of the virtual model can include the gender, age, and region information of the virtual model. For example, an optional attribute information is "female, young, Asian".
[0073] In an optional embodiment, the above-mentioned initial image, pose image, and attribute information can be determined depending on the operations of a user (such as a merchant, etc.) on the human-computer interaction interface at the front end of the target processing system. For example, if the target processing system is a home decoration application, at the front end of this application, a merchant user can upload a scene graph (i.e., the initial image) and a pose image designed by themselves, and input (or select from a material library) model attribute information, etc. Another example is that the user uploads a scene graph and a pose reference graph designed by themselves, and the pose reference graph includes a reference virtual model presenting an initial pose. The target processing system determines the pose image according to the pose reference graph. In the process of determining the pose image according to the pose reference graph, the pose image can be directly determined based on the initial pose in the pose reference graph, or the initial pose in the pose reference graph can be used to render adjustable bone joint points in the human-computer interaction interface, so as to respond to the position adjustment operation of the user on the human-computer interaction interface, determine the model pose that the user expects the virtual model to present according to the user's position adjustment operation, and then determine the pose image.
[0074] Step S202: Process the initial image, pose image, and attribute information through a vision-language model to obtain text information, where the text information is used to describe the pose of the virtual model.
[0075] In an optional embodiment, the target processing system can directly use the initial image, pose image, and attribute information as input data and input them into the vision-language model. The vision-language model processes the input data to infer the action pose that the virtual model should present in a specific scene (i.e., the scene shown in the initial image), and obtains text information. For example, an optional piece of text information is "sitting in front of a dressing table doing makeup". Among them, by using the model attribute as one of the input conditions of the vision-language model, it can effectively guide the model to generate results that better meet the actual needs and scene logic. If the model attribute is not specified, the model may generate some results that do not conform to the real scene, such as "a child playing with a mobile phone" or "a woman debugging equipment" and other situations that do not conform to common sense.
[0076] In an optional embodiment, in order to improve the accuracy of the text information output by the visual language model, the target processing system may generate a corresponding mask image according to the pose image, where the mask image is used to determine the regional image of the virtual model in the initial image, that is, the mask image is used to determine the redrawing area of the virtual model in the initial scene graph image. After determining the mask image, the target processing system may further determine the position information of the mask image in the initial image, and then use the initial image, the mask image, the position information of the mask image in the initial image, and the attribute information as input data and input them into the visual language model, so as to more finely guide the visual language model to infer the action pose of the virtual model based on the mask image and its position information in the initial image, and obtain text information that can accurately describe the action intention of the virtual model and its interaction relationship with the scene. Among them, the position information may be determined based on the model movement operation of the user on the human-computer interaction interface. The foregoing model movement operation may refer to the operation of the user inputting the coordinate information of the virtual model in the interface, or the dragging operation of the user on the overall skeletal joint points in the initial image in the interface.
[0077] Optionally, the position information of the foregoing mask image in the initial image may be input into the visual language model in text form (such as coordinate information). Optionally, the target processing system may also paste the mask image at the corresponding position in the initial image to obtain the pasted image, and then input the pasted image into the visual language model to input the position information, so that the visual language model determines the position information of the mask image in the initial image according to the pasted image.
[0078] In an optional embodiment, in order to improve the accuracy of the text information output by the visual language model and the richness of its content, the target processing system may also supplement the clothing information of the virtual model as input data of the visual language model. The clothing information of the virtual model may be in text form or in image form. For example, the user may select clothing materials (the materials may be text or images) in the material library displayed on the human-computer interaction interface for the target processing system to determine the clothing information of the virtual model, so that the generated text information may also include clothing description information of the virtual model.
[0079] Step S203, process the text information, attribute information, pose image, and initial image through an image generation model to obtain a target image, where the target image includes a virtual model.
[0080] Optionally, the image generation model may also be referred to as a redrawing model. The image generation model may be obtained by training (or referred to as fine-tuning) a diffusion model, and its training method may be reinforcement learning, or supervised training, or alternately performing reinforcement learning and supervised training.
[0081] In an alternative embodiment, the target processing system can directly input text information, attribute information, pose images, and initial images into an image generation model to obtain a target image.
[0082] In an alternative embodiment, the target processing system can generate a corresponding mask image based on the pose image, and then input the text information, attribute information, pose image, initial image, and the corresponding mask image into the image generation model to obtain a target image.
[0083] In an alternative embodiment, the target processing system can input the text information, attribute information, pose image, initial image, and the corresponding mask image into the image generation model. First, the image generation model generates an initial image with a virtual model, and then the human segmentation model processes the image with the virtual model to extract an accurate mask of the model part. Thus, the accurate mask, text information, attribute information, pose image, and initial image are input into the image generation model to obtain a target image.
[0084] Optionally, the target image may refer to a target scene image. The target image can be understood as an image obtained by drawing a virtual model in the initial image. The target image includes a display object and a virtual model. The target image can be called a "model display image", a "scene model image", etc.
[0085] In this solution, a vision-language model is used to process the initial image, the pose image of the virtual model, and the attribute information of the virtual model to obtain text information, effectively determining the text conditions for intelligently inferring the action poses that the virtual model should present in a specific scene. These text conditions can accurately describe the action intentions of the virtual model and its interaction relationship with the content of the initial image, providing an accurate guiding basis for generating natural and reasonable model poses. Therefore, when using the image generation model to process the text information, attribute information, pose image, and initial image to obtain the target image, the degree of fit between the model pose and the content of the initial image can be improved, avoiding the problem of generating unreasonable poses due to lack of context understanding in traditional methods, and thus effectively improving the accuracy of image generation. This application achieves the purpose of generating an image containing a virtual model by combining text conditions related to the model pose, thus realizing the technical effect of improving the accuracy of generating an image containing a virtual model, and further solving the technical problem of low accuracy in generating an image containing a virtual model in the related art.
[0086] How to determine accurate text information is crucial. Therefore, in the image generation method provided in Embodiment 1 of this application, the visual language model processes the initial image, the pose image, and the attribute information to obtain text information, including: generating a first mask image according to the pose image, where the first mask image is used to represent the regional image of the virtual model in the initial image; performing a fitting process on the first mask image and the initial image to obtain a first image; and inputting the first image and the attribute information into the visual language model to obtain text information.
[0087] In an alternative embodiment, the target processing system can input the pose image into the mask image generation model, and the mask image generation model generates a mask image corresponding to the pose image, that is, generates the first mask image. The training sample used by the mask image generation model during training can be a sample pose image, and the corresponding true label of the training sample is a reference mask image. Thus, the initial mask image generation model is trained based on the supervised training method to obtain the mask image generation model.
[0088] In an alternative embodiment, the target processing system can also supplement the morphological features of the virtual model in the pose image based on a preset morphological processing rule to obtain an object morphology map, and then perform a dilation process on the image content in the object morphology map to obtain the first mask image.
[0089] Optionally, Figure 3 is a schematic diagram of generating text information provided in Embodiment 1 of this application. As Figure 3 shown, after obtaining the first mask image, the target processing system can perform a fitting process on the first mask image and the initial image according to the position information of the first mask image in the initial image to obtain the first image.
[0090] Optionally, the position information of the first mask image in the initial image can be determined based on the model movement operation of the user on the human-computer interaction interface. For example, the aforementioned model movement operation can refer to the operation of the user inputting the coordinate information of the virtual model in the interface, or for another example, the model movement operation can refer to the operation of the user dragging the overall bone joint points in the initial image on the interface.
[0091] Optionally, the skeletal joint points in the foregoing initial image can be rendered in the initial image based on preset skeletal joint point information, or can also be rendered in the initial image according to a pose reference diagram provided by the user. The pose reference diagram includes a reference virtual model presenting an initial pose. The skeletal joint points in the initial image can not only be used by the user for overall dragging and dropping operations, but also for position adjustment operations of some skeletal joint points. Thus, when the user performs an overall dragging and dropping operation, the target processing system can determine the position information of the first mask image in the initial image based on this operation. When the user performs a position adjustment operation on some skeletal joint points, the target processing system can determine the pose information of the virtual model based on this operation, and then determine a pose image.
[0092] After obtaining the first image, as Figure 3 shown, the target processing system can input the first image and the attribute information into a vision-language model to obtain text information. In an optional embodiment, in addition to inputting the first image and the attribute information into the vision-language model, the target processing system can also input a preset prompt into the vision-language model. The prompt is used to guide the vision-language model to generate text information according to the first image and the attribute information. For example, an optional prompt is "The mask in the input image (i.e., the first image) is the area image occupied by the virtual model to be drawn in this image. Combining the following model attribute information, infer the action pose that the virtual model should present in the image and generate a corresponding text description."
[0093] Optionally, the text information derived in the above manner can be used to guide an image generation model, which can effectively improve the pose rationality of the virtual model. For example,
[0094] is a schematic diagram of a target image generated in the prior art. As Figure 4 shown, in the prior art, the generated virtual model is sitting in front of a dressing table holding a book. The shape of the pants is irregular, and the shoes are stepping on the bed body. There are many unreasonable points in this scene image. Compared with the prior art, Figure 4 shown, Figure 5Schematic diagram of the target image provided in the first embodiment of the present application Figure 1 As shown in Figure 5 Under the guidance of the text information in the present application, in the generated target image, the virtual model is sitting in front of a dressing table presenting a makeup pose, the shape of the pants is regular, the position of the feet is real, and the pose of the virtual model is relatively more reasonable.
[0095] It should be noted that by generating the first mask image according to the pose image, it can provide clear boundary guidance for subsequent image processing and fusion. By performing a fitting process on the first mask image and the initial image, the redrawn image of the virtual model can be visualized in the initial image, enabling the visual language model to more intuitively understand the relationship between the virtual model and the image content in the initial image, such as the relative position of the model in the scene, the layout of surrounding objects, and the possible interaction patterns of the model. Thus, it provides a visual basis for intelligently inferring the natural and reasonable actions of the model in a specific scene, improving the accuracy of text information generation.
[0096] In order to generate an accurate first mask image, in the image generation method provided in the first embodiment of the present application, generating the first mask image according to the pose image includes: supplementing the morphological features of the virtual model in the pose image based on a preset morphological processing rule to obtain an object morphology map; performing a dilation process on the image content in the object morphology map to obtain the first mask image.
[0097] Optionally, the preset morphological processing rule includes a method for supplementing the morphological features of the virtual model. For example, when the model type of the virtual model is a human, one optional morphological processing rule is as follows:
[0098] (1) Fitting an ellipse with the smallest area based on the joint points of the head as the representation of the head;
[0099] (2) Inferring the direction of the sole according to the parallelogram principle and inferring the length of the sole according to the length of the thigh.
[0100] Optionally, the target processing system can supplement the morphological features of the virtual model in the pose image based on the morphological processing rule to obtain an object morphology map. The object morphology map is an image including the complete morphological contour of the virtual model.
[0101] After obtaining the object morphology map, the target processing system can perform a dilation process on the image content in the object morphology map. For example, expanding the pixels on the contour line in the object morphology map by a preset magnification. Generally speaking, it can also be understood as thickening the contour line in the object morphology map by a preset magnification, thereby determining the image content area in the dilated object morphology map as the mask area to obtain the first mask image.
[0102] For example,Figure 6 It is a schematic diagram of generating the first mask image provided in the first embodiment of the present application. As Figure 6 shown, Figure 6 after the pose image on the far left in Figure 6 is processed by rules, an object morphology map as shown in the middle image in Figure 6 is obtained. After the object morphology map is dilated, the first mask image as shown in the right image in
[0103] is obtained. It should be noted that by combining morphological features to supplement and determine the first mask image, the visual language model can more comprehensively understand the morphology of the virtual model after receiving the first mask image. The dilation process effectively expands the redrawing area, avoids the unreality of the virtual model caused by improper redrawing range, thereby improving the accuracy of the generated first mask image and effectively reducing the variation rate between the virtual model and the scene.
[0104] In order to obtain a more accurate pose image, in the image generation method provided in the first embodiment of the present application, the pose image of the virtual model is obtained in the following way: obtaining a pose reference map, where the pose reference map includes a reference virtual model presenting an initial pose; extracting bone joint points from the reference virtual model in the pose reference map to obtain bone joint point information; determining the pose image of the virtual model according to the bone joint point information.
[0105] Optionally, the user can upload the pose reference map through the human-computer interaction interface for the target processing system to obtain. The pose reference map includes the image content of the initial image and the reference virtual model presenting the initial pose. In an optional embodiment, the pose reference map has the same image size as the initial image and the same image content of the scene. In other words, the pose reference map can be regarded as an image obtained by adding a reference virtual model to the initial image. For example, Figure 7 is a schematic diagram of generating a pose image provided in the first embodiment of the present application. As Figure 7 shown, Figure 7 the leftmost image in
[0106] is an optional pose reference map, and the pose reference map includes a reference virtual model sitting in front of a dressing table, that is, the initial pose of the reference virtual model is "sitting in front of a dressing table".
[0107] After obtaining the pose reference image, the target processing system can extract the bone joint points of the reference virtual model in the pose reference image to obtain bone joint point information. Optionally, the bone joint point information may include the coordinate data of multiple bone joint points in the virtual model's body in the pose reference image and the connection relationships between the multiple bone joint points. For example, for Figure 7 the left pose reference image in Figure 7 the bone joint point information shown in the middle is obtained.
[0108] Optionally, the target processing system can use a pose estimation network to extract the bone joint points of the reference virtual model to identify and locate the bone key points of the reference virtual model, such as the head, shoulders, elbows, wrists, hips, knees, and ankles. For example, Figure 8 is a schematic diagram of the pose estimation network provided in the first embodiment of the present application. As Figure 8 shown, the pose estimation network may include an encoder and a decoder. The Encoder may include Multi-Head Self-Attention (abbreviated as MHSA), Layer Normalization (abbreviated as LN), Feedforward Neural Network (abbreviated as FFN), Deconvolution (abbreviated as Deconv), Batch Normalization (abbreviated as BN), Residual Connection (abbreviated as ReLU), Predictor, Bilinear, and may also include L Transformer Blocks. This model uses a Transformer architecture to replace the convolutional neural network, and can better capture the relationships between different key points through the self-attention mechanism to achieve efficient feature learning. Its parallel processing ability improves the efficiency of feature extraction, and at the same time can comprehensively consider global context information, making it perform better in complex scenarios.
[0109] Optionally, as Figure 8 shown, when processing an image, the pose estimation network divides it into multiple blocks and uses the self-attention mechanism to parallelly obtain the information connections between different blocks. This can improve the ability to extract global information, enabling the model to better capture the relationships between different regions, thereby improving the accuracy of pose estimation. The pose estimation network allows the user to set the number of divided local regions by themselves, so that it can be adjusted according to the specific tasks and scene requirements, making the model more flexible and adaptable to different situations.
[0110] In an optional embodiment, after obtaining the skeletal joint point information, the target processing system can directly generate a corresponding pose image based on the skeletal joint point information, that is, generate the pose image corresponding to the aforementioned initial pose. For example, a blank background with the same size as the reference pose map (that is, the same image size as the initial image) is set, and the skeletal joint points are rendered on the blank background based on the Figure 7 skeletal joint point information shown in the middle part to obtain an initial pose image, and then as shown in Figure 7 , a cropping operation is performed on the initial pose image so that the skeletal joint points are located at the center of the blank background to obtain a pose image. In an optional embodiment, the initial pose image can also be directly determined as the pose image, that is, without passing through the cropping process.
[0111] Optionally, the pose image does not include the specific appearance details of the virtual model, but only reflects the relative positions and poses of the skeletal joint points, so as to provide intuitive visual information for generating a natural and reasonable pose of the virtual model. The target processing system can determine the position information of the reference virtual model in the pose reference map as the position information of the first mask image in the initial image. The position information of the reference virtual model in the pose reference map can refer to the position information of at least two skeletal joint points of the reference virtual model in the pose reference map.
[0112] In an optional embodiment, the target processing system can also display the initial image on the human-computer interaction interface, and after obtaining the skeletal joint point information, render adjustable skeletal joint points in the initial image displayed on the human-computer interaction interface according to the skeletal joint point information, so that the user can adjust the skeletal joint points according to actual needs. The target processing system can adjust the positions of the rendered skeletal joint points according to the user's position adjustment operation on the skeletal joint points to obtain target skeletal joint points, and then determine the pose image of the virtual model based on the target skeletal joint points to flexibly adapt to the scene.
[0113] It should be noted that by obtaining the pose reference map and extracting the skeletal joint point information from it, it is convenient to accurately obtain the key point information of the model pose, providing an accurate and detailed morphological control basis for subsequent generation of the pose image. By generating the pose image based on the skeletal joint point information, the visualization processing of the skeletal joint point information is realized, providing intuitive visual information for generating a natural and reasonable pose of the virtual model, thereby facilitating the improvement of the accuracy of generating the target image.
[0114] In order to more flexibly determine the pose image, in the image generation method provided in Embodiment 1 of the present application, determining the pose image of the virtual model according to the skeletal joint point information includes: rendering the skeletal joint points in the initial image displayed on the human-computer interaction interface according to the skeletal joint point information; responding to the position adjustment operation applied to the human-computer interaction interface, adjusting the positions of the skeletal joint points to obtain target skeletal joint points; determining the pose image of the virtual model according to the target skeletal joint point information of the target skeletal joint points.
[0115] Optionally, the target processing system may display the initial image in the editing area of the human-computer interaction interface and visualize the skeletal joint point information, overlaying it in the form of joint point markers on the initial image, so that the user can intuitively see the positions of the skeletal joint points of the virtual model in the initial image through the human-computer interaction interface, thereby providing a visual reference for subsequent adjustments. For example, Figure 9 is a schematic diagram of the editing area provided in Embodiment 1 of the present application. In Figure 9 the initial image, the skeletal joint points are rendered, and the skeletal joint points are marked as circular contacts and connected by white semi-transparent rectangles to present a specific pose. When the user expects to adjust the model pose, the user can drag the Figure 9 circular contacts shown in it. The target processing system responds to the position adjustment operation (i.e., the aforementioned operation of dragging the circular contact) applied by the user to the human-computer interaction interface, and adjusts the positions of the skeletal joint points rendered in the human-computer interaction interface to obtain target skeletal joint points.
[0116] In an optional embodiment, the user can perform multiple rounds of position adjustment operations on the skeletal joint points in the human-computer interaction interface. When the user determines that the adjustment is completed, the user can click a specific button on the human-computer interaction interface, such as a button containing words like "start generating image" or "determine model pose". The target processing system can respond to the user's confirmation operation (i.e., the aforementioned clicking of the specific button) and determine the skeletal joint points currently presented on the human-computer interaction interface as the target skeletal joint points.
[0117] In an optional embodiment, the target processing system supports the user to perform an overall dragging operation on the skeletal joint points in the initial image, so as to flexibly determine the position information of the subsequently generated first mask image in the initial image. For example, in Figure 9 a rectangular frame is set around the whole of the skeletal joint points, and buttons for operating on the whole of the skeletal joint points are provided at the four corners of the rectangular frame, such as Figure 9As shown, the "Move" button is located at the upper left corner of the rectangular frame, which can be used to move the entire skeletal joint points. The "Zoom" button is located at the upper right corner of the rectangular frame, which can be used to enlarge or reduce the entire skeletal joint points. The "Mirror" button is located at the lower left corner of the rectangular frame, which can be used to perform mirror processing on the entire skeletal joint points. The "Rotate" button is located at the lower right corner of the rectangular frame, which can be used to perform rotation processing on the entire skeletal joint points. The target processing system can respond to the adjustment operations applied to the four corners of the rectangular frame, perform adjustment operations on the entire skeletal joint points, and then, in response to the user's confirmation operation (i.e., the aforementioned clicking of a specific button), determine the position information of the skeletal joint points currently presented on the human-computer interaction interface (i.e., the target skeletal joint points) in the initial image (e.g., including the position information of at least two skeletal joint points in the initial image) as the position information of the first mask image in the initial image.
[0118] Optionally, after determining the target skeletal joint points, the target processing system can determine the pose image of the virtual model according to the target skeletal joint point information of the target skeletal joint points. For example, a blank background with the same image size as the initial image is set, and the skeletal joint points are rendered on the blank background based on the target skeletal joint point information to obtain an initial pose image, and then the initial pose image is cropped so that the rendered skeletal joint points are located at the center of the blank background to obtain the pose image. In an alternative embodiment, the initial pose image can also be directly determined as the pose image.
[0119] It should be noted that by introducing a pose reference diagram, automatically detecting human skeletal key points, and providing users with a fine-tuning function for skeletal joint points, the difficulty of obtaining pose images is significantly reduced. Specifically, users only need to upload a pose reference diagram, and the system can automatically extract the skeletal key point information therein, while allowing users to manually adjust the key points to meet personalized needs. This design not only simplifies the operation process for users to customize the model pose, but also greatly improves the flexibility and accuracy of generating poses, providing users with a highly customized solution. In addition, this application can also automatically determine the area that needs to be redrawn according to the changes in the skeletal key points, that is, automatically determine the first mask image, so as to generate a natural and reasonable model display diagram with a high degree of fit to the scene requirements, and improve the generation accuracy of the target image.
[0120] In order to generate a more accurate target image, in the image generation method provided in the first embodiment of the present application, the text information, attribute information, pose image, and initial image are processed by an image generation model to obtain a target image, including: generating a first mask image according to the pose image, where the first mask image is used to represent the regional image of the virtual model in the initial image; inputting the text information, attribute information, pose image, first mask image, and initial image into the image generation model to obtain a second image, where the second image includes the initial virtual model; determining the target image according to the second image.
[0121] Optionally, the method of generating the first mask image according to the pose image has been explained in the foregoing, so it will not be elaborated here.
[0122] Optionally, Figure 10 is a schematic diagram of determining the second image provided in the first embodiment of the present application. As Figure 10 shown, the target processing system can use the text information and attribute information as the target text conditions, and then input the target text conditions, initial image, first mask image, and pose image into the image generation model to obtain the second image.
[0123] In an optional embodiment, the image generation model can also be called a redrawing model. The image generation model can be obtained by training (or called fine-tuning) a diffusion model, and its training method can be reinforcement learning, or supervised training, or alternately performing reinforcement learning and supervised training. Optionally, the training samples in the training sample set used in the training process of the diffusion model can include sample text conditions, sample pose images of the sample virtual model, mask images corresponding to the sample virtual model, and sample initial images. The true label corresponding to the training sample can be a reference target image including the foregoing sample virtual model. The sample text conditions include sample text information for describing the pose of the sample virtual model and sample attribute information of the sample virtual model. In an optional embodiment, the fineness of the mask image corresponding to the sample virtual model in the training sample can be divided into multiple levels to improve the generalization ability of model training. For example, a mask image with a low fineness level can only show the model pose of the sample virtual model, while a mask image with a high fineness level can show not only the model pose of the sample virtual model but also information such as the shape of the clothing worn by the model.
[0124] Optionally, as Figure 10As shown, during the process of generating the second generated image, image generation can be combined with a pose control network. For example, the pose control network can be used to extract pose features at different scales from the pose image. The pose control network can be trained using the poses of sample virtual models to control the network structure. The aforementioned different scales can also be referred to as different dimensions. For example, scales such as 512*512, 256*256, 128*128, etc. are used here only for illustration and are not specifically limited. After obtaining the pose features, the attention mechanism in the image generation model is used to convert the pose features into control information for the image generation model. Using this control information, the image generation model analyzes the target text condition, initial image, first mask image, and pose image under its model parameters to obtain the second generated image. For example, the first mask image and the initial image are combined to obtain a combined image. The encoder is used to generate an initial latent variable in the latent variable space based on the combined image, the first mask image, the initial image, and the target text condition. Using the control information, the image generation model analyzes the initial latent variable under the model parameters to obtain the second generated image.
[0125] Optionally, the ControlNet pose control method can be used to implement the above-mentioned pose control network. This method combines deep learning techniques for extracting and processing pose features. During training, the outputs of each layer of the pose control network are injected into the decoding part of the diffusion model in a cross-attention manner. Thus, during the process of generating an image, the pose information can be combined to adjust and generate the image, achieving a more accurate and expected pose performance.
[0126] In an optional embodiment, after obtaining the second image, the initial virtual model in the second image can be directly determined as the virtual model to be drawn in the initial image, that is, the second image is directly determined as the target image.
[0127] In an optional embodiment, after obtaining the second image, in order to further improve the rationality and authenticity of the generated target image, a fine mask of the virtual model to be drawn can be determined based on the second image, and then the target image can be generated in combination with this fine mask.
[0128] It should be noted that by generating the first mask image according to the pose image and inputting the first mask image together with the text information, attribute information, pose image, and initial image into the image generation model, the image generation model can achieve more precise control over the redrawing area of the virtual model, reduce background variation, improve the naturalness of the generated image, and thus improve the generation accuracy of the target image.
[0129] In order to generate a more accurate target image, in the image generation method provided in the first embodiment of this application, determining the target image according to the second image includes: segmenting the regional image of the initial virtual model in the second image to obtain a second mask image, where the second mask image is used to represent the target regional image of the virtual model in the initial image; processing the text information, attribute information, pose image, second mask image, and initial image through an image generation model to obtain the target image.
[0130] Since the first mask image is determined according to the pose image of the virtual model, the information contained in the pose image of the virtual model is less, and the fineness of the first mask image is relatively low. Therefore, in order to improve the accuracy of the target image generated by the image generation model, the mask image of the virtual model can be optimized, that is, a relatively finer mask image is determined to replace the first mask image and input it into the image generation model, so as to facilitate obtaining a target image with a better image effect.
[0131] Optionally, after obtaining the second image, the target processing system can segment the regional image of the initial virtual model in the second image to obtain a second mask image, so as to realize the acquisition of a finer mask image of the virtual model. The second mask image is used to represent the target regional image of the virtual model in the initial image, and the fineness of the target regional image represented by the second mask image is higher than that of the regional image represented by the first mask image. Specifically, the second mask image can more accurately define the regional image occupied by the virtual model in the initial image, and the defined boundary is more detailed and more conforms to the body contour of the virtual model. For example, Figure 11 is a schematic diagram of determining the second mask image provided in the first embodiment of this application, Figure 11 showing Figure 10 the second mask image obtained after image segmentation of the second image in Figure 10 The fineness of this second mask image has been significantly improved compared with the first mask image shown in
[0132] Optionally, the second image can be segmented through a human segmentation model to obtain a second mask image. The human segmentation model can be trained based on a supervised training method. The training samples used by the human segmentation model during training can be images containing sample virtual models (or scene images containing sample virtual models), and the true labels of the training samples can be mask images corresponding to the sample virtual models.
[0133] After determining the second mask image, the text information, attribute information, pose image, second mask image, and initial image are processed through an image generation model to obtain the target image. For example, Figure 12It is a schematic diagram of determining a third image provided in the first embodiment of the present application. As Figure 12 shown, the target processing system can use text information and attribute information as target text conditions, and then input the target text conditions, the initial image, the second mask image, and the pose image into the image generation model, and determine the image output by the image generation model as the third image, so as to determine the target image according to the third image. For example, directly use the third image as the target image, or further process the third image to obtain the target image.
[0134] Optionally, guiding the image generation model based on the fine mask generated in the above manner can effectively improve the quality and naturalness of the generated image. For example, Figure 13 It is a schematic diagram of the target image provided in the first embodiment of the present application Figure 2 , Figure 13 The target image shown in Figure 14 is a model image generated according to the first mask image. The generated virtual model wears a hat on the head indoors. Figure 3 , Figure 14 The target image shown in
[0135] It should be noted that by first generating a rough redrawing area based on a preset rule, then generating a preliminary image (i.e., the initial image) through primary redrawing, and extracting an accurate model mask from the preliminary image, this method not only greatly reduces the complexity of user operations, but also effectively solves the variation problem between the model and the scene caused by inaccurate redrawing areas, thereby improving the realism and naturalness of the generated image, that is, improving the accuracy of the target image.
[0136] In order to generate a more accurate target image, in the image generation method provided in the first embodiment of the present application, processing the text information, attribute information, pose image, second mask image, and initial image through the image generation model to obtain the target image includes: inputting the text information, attribute information, pose image, second mask image, and initial image into the image generation model to obtain the third image; repairing and processing the target area image in the third image through the image repair model to obtain the updated target area image, where the target area image refers to the area image including the virtual model; superimposing the updated target area image on the initial image to obtain the target image.
[0137] Optionally, the target processing system may use the text information and attribute information as target text conditions, and then input the target text conditions, the initial image, the second mask image, and the pose image into the image generation model, and determine the image output by the image generation model as the third image.
[0138] After obtaining the third image, the target processing system may identify the target region image from the third image, and then use the redrawing and restoration model to repair the model details in the target region image with a relatively low redrawing amplitude. For example, repair the details of hands, feet, clothing, and face to obtain the updated target region image, and superimpose the updated target region image on the initial image to obtain the target image. For example, Figure 15 FIG. is a schematic diagram of determining a target image according to Embodiment 1 of the present application. For example, Figure 15 as shown, for Figure 15 the leftmost target region images in are respectively subjected to clothing detail repair and face detail repair, and the repaired images are superimposed on the initial image to obtain the target image.
[0139] Optionally, the redrawing and restoration model is a model with image generation function, which can be trained based on supervised training. The training samples used by the redrawing and restoration model during training may be the model images before detail repair, and the true labels of the training samples may be the model images after manual detail repair.
[0140] It should be noted that through the above method, the detail presentation of the virtual model in the generated target image and the naturalness between its fusion with the scene can be effectively improved, thereby improving the accuracy of the target image.
[0141] In an alternative embodiment, the schematic diagram shown in Figure 16 may be adopted to implement the generation of the target image. Figure 16 FIG. is a flowchart of the image generation method according to Embodiment 1 of the present application Figure 2 as Figure 16As shown, the user can upload an initial image, custom skeletal joint point information, and model attributes through the human-computer interaction interface. The target processing system can, through the vision-language large model, based on the initial image, the pose image corresponding to the custom skeletal joint point information, and the model attributes, intelligently infer the action pose that the virtual model should present in a specific scenario, generate corresponding text information, and then determine the text information and attribute information as the target text conditions. This process provides clear action guidance for subsequent generation. In another branch, the pose image is transformed based on preset shape processing rules to preliminarily determine a rough redrawing area (i.e., the first mask image). Subsequently, the redrawing model (i.e., the image generation model) generates an initial model image (i.e., the second image) according to the redrawing area, the target text conditions, the pose image, and the initial image. Although there may be a certain mutation rate in the image generated at this stage, it provides a basic reference for the subsequent accurate human body mask. Then, the generated initial model image is processed by the human body segmentation model to extract an accurate mask of the human body part (i.e., the second mask image). Finally, the above target text conditions, accurate mask, pose image, and initial image are input into the redrawing model together to complete the generation of the final model image (i.e., the third image). After post-processing steps by the image repair model, a model scene image (i.e., the target image) highly consistent with the user's needs can be obtained. Among them, Figure 16 Some data that need to be input into the redrawing model are omitted, such as the initial image and the pose image.
[0142] In the embodiment of the present application, a vision-language model is used to process the initial image, the pose image of the virtual model, and the attribute information of the virtual model to obtain text information, effectively determining the text conditions for intelligently inferring the action pose that the virtual model should present in a specific scenario. These text conditions can accurately describe the action intention of the virtual model and its interaction relationship with the content of the initial image, providing an accurate guiding basis for generating natural and reasonable model poses. Thus, when using the image generation model to process the text information, attribute information, pose image, and initial image to obtain the target image, the fit degree between the model pose and the content of the initial image can be improved, avoiding the problem of generating unreasonable poses due to lack of context understanding in traditional methods, and thus effectively improving the accuracy of image generation. The present application achieves the purpose of generating an image containing a virtual model by combining text conditions related to the model pose, thereby realizing the technical effect of improving the accuracy of generating an image containing a virtual model, and further solving the technical problem of low accuracy in generating an image containing a virtual model in the related art.
[0143] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0144] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of the various embodiments of this application.
[0145] Embodiment 2
[0146] According to an embodiment of this application, there is also provided a method for generating an image, as Figure 17 shown, the method includes:
[0147] Step S1701, obtaining an initial image uploaded by a client, a pose image of a virtual model to be drawn on the initial image, and attribute information of the virtual model, wherein the virtual model is used to assist in displaying a display object in the initial image.
[0148] Step S1702, processing the initial image, the pose image, and the attribute information in a cloud server through a vision-language model to obtain text information, wherein the text information is used to describe the pose of the virtual model; processing the text information, the attribute information, the pose image, and the initial image through an image generation model to obtain a target image, wherein the target image includes the virtual model.
[0149] Step S1703, feeding back the target image to the client.
[0150] Through the above solution, a visual language model is used to process the initial image, the pose image of the virtual model, and the attribute information of the virtual model to obtain text information, effectively determining the text conditions for intelligently inferring the action poses that the virtual model should present in a specific scenario. These text conditions can accurately describe the action intentions of the virtual model and its interaction relationship with the content of the initial image, providing an accurate guiding basis for generating natural and reasonable model poses. Therefore, when using an image generation model to process the text information, attribute information, pose image, and initial image to obtain the target image, the degree of fit between the model pose and the content of the initial image can be improved, avoiding the problem of generating unreasonable poses due to lack of context understanding in traditional methods, thereby effectively improving the accuracy of image generation. This application achieves the purpose of generating an image containing a virtual model by combining text conditions related to the model pose, thus realizing the technical effect of improving the accuracy of generating an image containing a virtual model, and further solving the technical problem of low accuracy in generating an image containing a virtual model in the related art.
[0151] In the cloud server, the specific method for generating the image is the same as the method in Embodiment 1 and will not be elaborated here.
[0152] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0153] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0154] Embodiment 3
[0155] According to an embodiment of the present application, there is also provided an image generation device for implementing the above image generation method, as Figure 18As shown in the figure, the device includes: a first acquisition unit 1801, a first processing unit 1802, and a second processing unit 1803.
[0156] The first acquisition unit 1801 is configured to acquire an initial image, a pose image of a virtual model to be drawn in the initial image, and attribute information of the virtual model, where the virtual model is used to assist in displaying a display object in the initial image;
[0157] The first processing unit 1802 is configured to process the initial image, the pose image, and the attribute information through a vision-language model to obtain text information, where the text information is used to describe the pose of the virtual model;
[0158] The second processing unit 1803 is configured to process the text information, the attribute information, the pose image, and the initial image through an image generation model to obtain a target image, where the target image includes the virtual model.
[0159] In the image generation device provided in Embodiment 3 of the present application, the first acquisition unit 1801 acquires an initial image, a pose image of a virtual model to be drawn in the initial image, and attribute information of the virtual model, where the virtual model is used to assist in displaying a display object in the initial image; the first processing unit 1802 processes the initial image, the pose image, and the attribute information through a vision-language model to obtain text information, where the text information is used to describe the pose of the virtual model; the second processing unit 1803 processes the text information, the attribute information, the pose image, and the initial image through an image generation model to obtain a target image, where the target image includes the virtual model. In this solution, the vision-language model is used to process the initial image, the pose image of the virtual model, and the attribute information of the virtual model to obtain text information, effectively determining the text conditions for intelligently inferring the action poses that the virtual model should present in a specific scene. These text conditions can accurately describe the action intentions of the virtual model and its interaction relationship with the content of the initial image, providing an accurate guiding basis for generating natural and reasonable model poses. Therefore, when using the image generation model to process the text information, the attribute information, the pose image, and the initial image to obtain the target image, the fit degree between the model pose and the content of the initial image can be improved, avoiding the problem of generating unreasonable poses due to lack of context understanding in traditional methods, thereby effectively improving the accuracy of image generation. The present application achieves the purpose of generating an image containing a virtual model by combining text conditions related to the model pose, thereby realizing the technical effect of improving the accuracy of generating an image containing a virtual model, and further solving the technical problem of low accuracy in generating an image containing a virtual model in the related art.
[0160] Optionally, in the image generation device provided in Embodiment 3 of the present application, the first processing unit includes: a first generation subunit, configured to generate a first mask image according to the pose image, where the first mask image is used to represent the regional image of the virtual model in the initial image; a first processing subunit, configured to perform a fitting process on the first mask image and the initial image to obtain a first image; a second processing subunit, configured to input the first image and the attribute information into a vision-language model to obtain text information.
[0161] Optionally, in the image generation device provided in Embodiment 3 of the present application, the first generation subunit includes: a first processing module, configured to supplement the morphological features of the virtual model in the pose image based on a preset morphological processing rule to obtain an object morphology map; a second processing module, configured to perform a dilation process on the image content in the object morphology map to obtain a first mask image.
[0162] Optionally, in the image generation device provided in Embodiment 3 of the present application, the image generation device further includes: a second acquisition unit, configured to acquire a pose reference image, where the pose reference image includes a reference virtual model presenting an initial pose; an extraction unit, configured to extract skeletal joint points of the reference virtual model in the pose reference image to obtain skeletal joint point information; a determination unit, configured to determine the pose image of the virtual model according to the skeletal joint point information.
[0163] Optionally, in the image generation device provided in Embodiment 3 of the present application, the determination unit includes: a third processing subunit, configured to render skeletal joint points in the initial image displayed on the human-computer interaction interface according to the skeletal joint point information; an adjustment subunit, configured to respond to a position adjustment operation applied to the human-computer interaction interface to adjust the positions of the skeletal joint points to obtain target skeletal joint points; a first determination subunit, configured to determine the pose image of the virtual model according to the target skeletal joint point information of the target skeletal joint points.
[0164] Optionally, in the image generation device provided in Embodiment 3 of the present application, the second processing unit includes: a second generation subunit, configured to generate a first mask image according to the pose image, where the first mask image is used to represent the regional image of the virtual model in the initial image; a fourth processing subunit, configured to input the text information, the attribute information, the pose image, the first mask image, and the initial image into an image generation model to obtain a second image, where the second image includes an initial virtual model; a second determination subunit, configured to determine a target image according to the second image.
[0165] Optionally, in the image generation device provided in Embodiment 3 of the present application, the second determination subunit includes: a third processing module, configured to perform segmentation processing on the regional image of the initial virtual model in the second image to obtain a second mask image, where the second mask image is used to represent the target regional image of the virtual model in the initial image; a fourth processing module, configured to process the text information, attribute information, pose image, second mask image, and initial image through an image generation model to obtain a target image.
[0166] Optionally, in the image generation device provided in Embodiment 3 of the present application, the fourth processing module includes: a first processing sub-module, configured to input the text information, attribute information, pose image, second mask image, and initial image into an image generation model to obtain a third image; a second processing sub-module, configured to perform restoration processing on the target regional image in the third image through an image restoration model to obtain an updated target regional image, where the target regional image refers to the regional image including the virtual model; a third processing sub-module, configured to superimpose the updated target regional image on the initial image to obtain a target image.
[0167] It should be noted here that the above-mentioned first acquisition unit 1801, first processing unit 1802, and second processing unit 1803 correspond to steps S201 to S203 in Embodiment 1. The above units have the same examples and application scenarios as the corresponding steps, but are not limited to the content disclosed in the above-mentioned Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.
[0168] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0169] Embodiment 4
[0170] An embodiment of the present application may provide an electronic device, and the electronic device may be any one of a group of electronic devices. Optionally, in this embodiment, the above-mentioned electronic device may also be replaced with a terminal device such as a mobile terminal.
[0171] Optionally, in this embodiment, the above-mentioned electronic device may be located in at least one of multiple network devices in a computer network.
[0172] In this embodiment, the above electronic device may execute the program code of the following steps in the method for generating an image: obtaining an initial image, a pose image of a virtual model to be drawn in the initial image, and attribute information of the virtual model, where the virtual model is used to assist in displaying a display object in the initial image; processing the initial image, the pose image, and the attribute information through a vision-language model to obtain text information, where the text information is used to describe the pose of the virtual model; processing the text information, the attribute information, the pose image, and the initial image through an image generation model to obtain a target image, where the target image includes the virtual model.
[0173] The above electronic device may also execute the program code of the following steps in the method for generating an image: generating a first mask image according to the pose image, where the first mask image is used to represent the regional image of the virtual model in the initial image; performing a fitting process on the first mask image and the initial image to obtain a first image; inputting the first image and the attribute information into the vision-language model to obtain text information.
[0174] The above electronic device may also execute the program code of the following steps in the method for generating an image: supplementing the morphological features of the virtual model in the pose image based on a preset morphological processing rule to obtain an object morphology map; performing a dilation process on the image content in the object morphology map to obtain a first mask image.
[0175] The above electronic device may also execute the program code of the following steps in the method for generating an image: obtaining a pose reference image, where the pose reference image includes a reference virtual model presenting an initial pose; extracting bone joint points from the reference virtual model in the pose reference image to obtain bone joint point information; determining the pose image of the virtual model according to the bone joint point information.
[0176] The above electronic device may also execute the program code of the following steps in the method for generating an image: rendering bone joint points in the initial image displayed on the human-computer interaction interface according to the bone joint point information; responding to a position adjustment operation applied to the human-computer interaction interface to adjust the positions of the bone joint points to obtain target bone joint points; determining the pose image of the virtual model according to the target bone joint point information of the target bone joint points.
[0177] The above electronic device may also execute the program code of the following steps in the method for generating an image: generating a first mask image according to the pose image, where the first mask image is used to represent the regional image of the virtual model in the initial image; inputting the text information, the attribute information, the pose image, the first mask image, and the initial image into the image generation model to obtain a second image, where the second image includes an initial virtual model; determining a target image according to the second image.
[0178] The above-mentioned electronic device can also execute the program code of the following steps in the method for generating an image: performing segmentation processing on the regional image of the initial virtual model in the second image to obtain a second mask image, where the second mask image is used to represent the target regional image of the virtual model in the initial image; processing the text information, attribute information, pose image, second mask image, and initial image through an image generation model to obtain a target image.
[0179] The above-mentioned electronic device can also execute the program code of the following steps in the method for generating an image: inputting the text information, attribute information, pose image, second mask image, and initial image into an image generation model to obtain a third image; performing restoration processing on the target regional image in the third image through an image restoration model to obtain an updated target regional image, where the target regional image refers to the regional image including the virtual model; superimposing the updated target regional image on the initial image to obtain a target image.
[0180] Optionally, Figure 19 is a structural block diagram of an electronic device according to an embodiment of the present application. As Figure 19 shown, the electronic device 190 may include: one or more ( Figure 19 only one is shown in the figure) processors 1902, a memory 1904. The electronic device 190 may further include a storage controller for controlling and managing the memory 1904; the electronic device 190 may further include a peripheral interface for connecting a radio frequency module, an audio module, a display screen, etc.
[0181] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the method and device for generating an image in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above-mentioned method for generating an image. The memory may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely set relative to the processor, and these remote memories may be connected to the terminal 10 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.
[0182] The processor can call the information and application programs stored in the memory through the transmission device to execute the following steps: obtain an initial image, a pose image of a virtual model to be drawn in the initial image, and attribute information of the virtual model, where the virtual model is used to assist in displaying the display object in the initial image; process the initial image, the pose image, and the attribute information through a vision-language model to obtain text information, where the text information is used to describe the pose of the virtual model; process the text information, the attribute information, the pose image, and the initial image through an image generation model to obtain a target image, where the target image includes the virtual model.
[0183] Optionally, the above-mentioned processor can also execute the program code of the following steps: generate a first mask image according to the pose image, where the first mask image is used to represent the regional image of the virtual model in the initial image; perform a fitting process on the first mask image and the initial image to obtain a first image; input the first image and the attribute information into the vision-language model to obtain text information.
[0184] Optionally, the above-mentioned processor can also execute the program code of the following steps: supplement the morphological features of the virtual model in the pose image based on a preset morphological processing rule to obtain an object morphology map; perform a dilation process on the image content in the object morphology map to obtain a first mask image.
[0185] Optionally, the above-mentioned processor can also execute the program code of the following steps: obtain a pose reference map, where the pose reference map includes a reference virtual model presenting an initial pose; extract the skeletal joint points of the reference virtual model in the pose reference map to obtain skeletal joint point information; determine the pose image of the virtual model according to the skeletal joint point information.
[0186] Optionally, the above-mentioned processor can also execute the program code of the following steps: render the skeletal joint points in the initial image displayed on the human-computer interaction interface according to the skeletal joint point information; respond to the position adjustment operation acting on the human-computer interaction interface to adjust the position of the skeletal joint points to obtain target skeletal joint points; determine the pose image of the virtual model according to the target skeletal joint point information of the target skeletal joint points.
[0187] Optionally, the above-mentioned processor can also execute the program code of the following steps: generate a first mask image according to the pose image, where the first mask image is used to represent the regional image of the virtual model in the initial image; input the text information, the attribute information, the pose image, the first mask image, and the initial image into the image generation model to obtain a second image, where the second image includes the initial virtual model; determine the target image according to the second image.
[0188] Optionally, the above processor may also execute program code for the following steps: performing segmentation processing on the regional image of the initial virtual model in the second image to obtain a second mask image, where the second mask image is used to represent the target regional image of the virtual model in the initial image; processing the text information, attribute information, pose image, second mask image, and initial image through an image generation model to obtain a target image.
[0189] Optionally, the above processor may also execute program code for the following steps: inputting the text information, attribute information, pose image, second mask image, and initial image into an image generation model to obtain a third image; performing restoration processing on the target regional image in the third image through an image restoration model to obtain an updated target regional image, where the target regional image refers to the regional image including the virtual model; superimposing the updated target regional image on the initial image to obtain a target image.
[0190] Those of ordinary skill in the art can understand that Figure 19 The structure shown is only schematic, and the electronic device can also be a terminal device such as a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, and a Mobile Internet Device (MID), a PAD, etc. Figure 19 It does not limit the structure of the above electronic device. For example, the electronic device 190 may further include more or fewer components (such as a network interface, a display device, etc.) than those shown Figure 19 in the figure, or have a different configuration from that shown Figure 19 in the figure.
[0191] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk, or an optical disc, etc.
[0192] Embodiment 5
[0193] An embodiment of the present application further provides a computer-readable storage medium. Optionally, in this embodiment, the above storage medium may be used to store the program code executed by the image generation method provided in the first embodiment above.
[0194] Optionally, in this embodiment, the above storage medium may be located in any one of the electronic devices in the electronic device group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0195] Example 6
[0196] An embodiment of the present application further provides a computer program product. Optionally, in this embodiment, the above computer program product may include a computer program, and when the computer program is executed by a processor, it implements the method for generating an image provided in the first embodiment above.
[0197] The serial numbers of the embodiments of the present application are only for description and do not represent the superiority or inferiority of the embodiments.
[0198] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0199] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of units or modules can be in electrical or other forms.
[0200] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0201] In addition, the functional units in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0202] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0203] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A method for generating an image, characterized in that, Including: Obtain an initial image, a pose image of a virtual model to be drawn in the initial image, and attribute information of the virtual model, where the virtual model is used to assist in displaying a display object in the initial image; Process the initial image, the pose image, and the attribute information through a vision - language model to obtain text information, where the text information is used to describe the pose of the virtual model; Process the text information, the attribute information, the pose image, and the initial image through an image generation model to obtain a target image, where the target image includes the virtual model.
2. The method according to claim 1, wherein Processing the initial image, the pose image, and the attribute information through a vision - language model to obtain text information includes: Generate a first mask image according to the pose image, where the first mask image is used to represent the regional image of the virtual model in the initial image; Perform a fitting process on the first mask image and the initial image to obtain a first image; Input the first image and the attribute information into the vision - language model to obtain the text information.
3. The method according to claim 2, wherein Generating a first mask image according to the pose image includes: Based on a preset morphological processing rule, supplement the morphological features of the virtual model in the pose image to obtain an object morphology map; Perform a dilation process on the image content in the object morphology map to obtain the first mask image.
4. The method according to claim 1, wherein The pose image of the virtual model is obtained by the following method: Obtain a pose reference image, where the pose reference image includes a reference virtual model presenting an initial pose; Extract bone joint points of the reference virtual model in the pose reference image to obtain bone joint point information; Determine the pose image of the virtual model according to the bone joint point information.
5. The method according to claim 4, characterized in that Determining the pose image of the virtual model according to the bone joint point information includes: Render bone joint points in the initial image displayed on the human - computer interaction interface according to the bone joint point information; Respond to a position adjustment operation acting on the human - computer interaction interface, adjust the positions of the bone joint points to obtain target bone joint points; Determine the pose image of the virtual model according to the target bone joint point information of the target bone joint points.
6. The method according to claim 1, wherein Processing the text information, the attribute information, the pose image, and the initial image through an image generation model to obtain a target image includes: Generate a first mask image according to the pose image, where the first mask image is used to represent the regional image of the virtual model in the initial image; Input the text information, the attribute information, the pose image, the first mask image, and the initial image into the image generation model to obtain a second image, where the second image includes an initial virtual model; Determine the target image according to the second image.
7. The method according to claim 6, characterized in that, Determining the target image according to the second image includes: Segment the regional image of the initial virtual model in the second image to obtain a second mask image, where the second mask image is used to represent the target regional image of the virtual model in the initial image; Process the text information, the attribute information, the pose image, the second mask image, and the initial image through the image generation model to obtain the target image.
8. The method according to claim 7, wherein Processing the text information, the attribute information, the pose image, the second mask image, and the initial image through the image generation model to obtain the target image includes: Input the text information, the attribute information, the pose image, the second mask image, and the initial image into the image generation model to obtain a third image; Process the target regional image in the third image through an image inpainting model to obtain an updated target regional image, where the target regional image refers to the regional image including the virtual model; Overlay the updated target regional image on the initial image to obtain the target image.
9. A method for generating an image, characterized in that, Includes: Obtain the initial image uploaded by the client, the pose image of the virtual model to be drawn in the initial image, and the attribute information of the virtual model, where the virtual model is used to assist in displaying the display object in the initial image; Process the initial image, the pose image, and the attribute information through a vision-language model in a cloud server to obtain text information, where the text information is used to describe the pose of the virtual model; process the text information, the attribute information, the pose image, and the initial image through an image generation model to obtain a target image, where the target image includes the virtual model; Feedback the target image to the client.
10. An image generation device, characterized in that, Includes: A first acquisition unit for acquiring an initial image, a pose image of a virtual model to be drawn in the initial image, and attribute information of the virtual model, where the virtual model is used to assist in displaying the display object in the initial image; A first processing unit for processing the initial image, the pose image, and the attribute information through a vision-language model to obtain text information, where the text information is used to describe the pose of the virtual model; A second processing unit for processing the text information, the attribute information, the pose image, and the initial image through an image generation model to obtain a target image, where the target image includes the virtual model.
11. An electronic device, characterized in that, Includes: A memory storing an executable program; A processor for running the program, where when the program runs, it executes the image generation method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, where when the executable program runs, it controls the device where the storage medium is located to execute the image generation method according to any one of claims 1 to 9.
13. A computer program product, characterized in that, Comprising a computer program or instructions which, when executed by a processor, implement the method for generating an image according to any one of claims 1 to 9.
Citation Information
Cited By
AI model chart generation method, system and device and medium
CN121685752A
A method, system, device and medium for generating AI model images
CN121685752B
Garment display posture transformation method based on diffusion model
CN122090512A