Image generation method, training method of image generation model and computer terminal

By extracting the structural features of virtual images and the stylistic features of real images, and combining them with target text data to generate target images, the problems of low granularity, diversity, and realism of virtual images are solved, and high-quality virtual reality image generation is achieved.

CN121921391APending Publication Date: 2026-04-24ALIBABA CLOUD COMPUTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ALIBABA CLOUD COMPUTING CO LTD
Filing Date
2024-10-17
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

The virtual images generated by existing technologies have low granularity, diversity, and realism, which cannot meet the requirements of high-quality virtual reality.

Method used

By extracting the structural features of virtual images and the stylistic features of real images, and combining them with target text data to generate target images, the structure of virtual objects is consistent with the style of real objects. This is achieved using techniques such as generative adversarial networks.

Benefits of technology

The generated target image has the same structure as the virtual image and the same style as the real image, which improves the fineness, diversity and realism of the image and meets the user's personalized needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921391A_ABST
    Figure CN121921391A_ABST
Patent Text Reader

Abstract

The invention discloses an image generation method, a training method of an image generation model and a computer terminal. The method comprises the steps that a virtual image, a real image and target text data are obtained, the virtual image is an image obtained by rendering a first virtual object in a virtual world through a virtual engine, the real image is an image obtained by shooting the real world, the virtual world is constructed based on the real world, and the target text data are obtained; the target text data is used for describing a second virtual object contained in the target image; extracting the structure of the first virtual object contained in the virtual image to obtain the structural features of the virtual image; extracting the style of the real image to obtain style features of the real image; and generating a target image based on the structural features, the style features and the target text data. According to the method and the device, the technical problems of relatively low fine granularity, diversity and authenticity of the generated virtual image in related technologies are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of digital twins and image processing, and more specifically, to an image generation method, a training method for an image generation model, and a computer terminal. Background Technology

[0002] Digital twin technology involves creating a virtual entity or system in the digital world that corresponds to an entity or system in the real world to simulate and predict behavior, performance, and conditions in the real world. This is known as generating virtual images that simulate real-world scenarios using digital twin technology. These virtual images can be various themes such as virtual characters, virtual scenes, and virtual objects, and can be used in various fields such as movies, games, and advertising. For example, they can be used to create realistic virtual worlds and virtual characters, providing a more immersive experience, and can be used for special effects in movies and games to make scenes and characters more realistic.

[0003] Currently, the virtual images generated by related technologies are often of relatively coarse quality, have a monotonous style, or contain content that does not conform to the real world, resulting in low granularity, diversity, and realism of the generated virtual images.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] This application provides an image generation method, an image generation model training method, and a computer terminal to at least solve the technical problems of low granularity, low diversity, and low realism of generated virtual images in related technologies.

[0006] According to one aspect of the embodiments of this application, an image generation method is provided. The method includes: acquiring a virtual image, a real image, and target text data, wherein the virtual image is an image obtained by rendering a first virtual object in a virtual world using a virtual engine, the real image is an image obtained by photographing the real world, the virtual world is constructed based on the real world, and the target text data is used to describe a second virtual object contained in the target image; extracting the structure of the first virtual object contained in the virtual image to obtain structural features of the virtual image; extracting the style of the real image to obtain style features of the real image; and generating a target image based on the structural features, style features, and target text data, wherein the second virtual object in the target image has the same structure as the first virtual object and the same style as the real image.

[0007] According to another aspect of the embodiments of this application, a training method for an image generation model is also provided. The image generation model is used to perform the above-described image generation method. The method includes: acquiring first training data and second training data, wherein the first training data includes a first training image and first text data corresponding to the first training image, the first training image includes a first virtual image or a first real image, and the second training data includes a second virtual image, a second real image corresponding to the second virtual image, and second text data; training a structural feature extractor and a style feature extractor in the image generation model using the first training data; and retraining the style feature extractor using the second training data to obtain the image generation model.

[0008] According to another aspect of the embodiments of this application, an image generation method is also provided. The method includes: acquiring a virtual city image, a real city image, and target text data, wherein the virtual city image is an image obtained by rendering a first virtual city in a virtual world using a virtual engine, the real city image is an image obtained by photographing a real city in the real world, the virtual world is constructed based on the real world, and the target text data is used to describe a second virtual city in the target image; extracting the structure of the first virtual city contained in the virtual city image to obtain structural features of the virtual city image; extracting the style of the real city image to obtain style features of the real city image; and generating a target city image based on the structural features, style features, and target text data.

[0009] According to another aspect of the embodiments of this application, an image generation method is also provided. The method includes: acquiring a virtual avatar image, a real avatar image, and target text data, wherein the virtual avatar image is an image obtained by rendering a first virtual avatar in a virtual world using a virtual engine, the real avatar image is an image obtained by photographing a real avatar in the real world, the virtual world is constructed based on the real world, and the target text data is used to describe a second virtual avatar in the target avatar image; extracting the structure of the first virtual avatar contained in the virtual avatar image to obtain the structural features of the virtual avatar image; extracting the style of the real avatar image to obtain the style features of the real avatar image; and generating the target avatar image based on the structural features, style features, and target text data.

[0010] According to another aspect of the embodiments of this application, an image generation method is also provided. The method includes: responding to an input command applied to an operation interface, displaying a virtual image, a real image, and target text data on the operation interface, wherein the virtual image is an image obtained by rendering a virtual model in a virtual world using a virtual engine, the real image is an image obtained by photographing the real world, the virtual world is constructed based on the real world, and the target text data is used to describe the image content of the target image; responding to a generation command applied to the operation interface, displaying a target image on the operation interface, wherein the target image is an image generated based on the structural features of the virtual image, the style features of the real image, and the target text data, wherein the structural features are features extracted from the structure of virtual objects contained in the virtual image, and the style features are features extracted from the style of the real image.

[0011] According to another aspect of the embodiments of this application, an image generation method is also provided. The method includes: acquiring a virtual image, a real image, and target text data by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter including the virtual image, the real image, and the target text data; the virtual image is an image obtained by rendering a virtual model in a virtual world using a virtual engine; the real image is an image obtained by photographing the real world; the virtual world is constructed based on the real world; and the target text data is used to describe the image content of the target image; extracting the structure of virtual objects contained in the virtual image to obtain the structural features of the virtual image; extracting the style of the real image to obtain the style features of the real image; generating a target image based on the structural features, style features, and target text data; and outputting the target image by calling a second interface, wherein the second interface includes a second parameter, the parameter value of the second parameter including the target image.

[0012] According to another aspect of the embodiments of this application, a computer terminal is also provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of this application when it runs.

[0013] According to another aspect of the embodiments of this application, an image processing system is also provided, the system comprising: a virtual engine for rendering a first virtual object in a virtual world to obtain a virtual image, wherein the virtual world is constructed based on the real world; a client for sending a real image and target text data, wherein the real image is an image obtained by taking a picture of the real world, and the target text data is used to describe a second virtual object contained in the target image; a server connected to the virtual engine and the client for extracting the structure of the first virtual object contained in the virtual image to obtain structural features of the virtual image, extracting the style of the real image to obtain style features of the real image, and generating a target image based on the structural features, style features, and target text data, wherein the second virtual object in the target image has the same structure as the first virtual object and the same style as the real image; the client is also used to output the target image.

[0014] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.

[0015] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.

[0016] According to another aspect of the embodiments of this application, a computer program product is also provided, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods in various embodiments of this application.

[0017] According to another aspect of the embodiments of this application, a computer program is also provided, which, when executed by a processor, implements the methods of the various embodiments of this application.

[0018] In this application embodiment, an image generation method is provided. The method includes: acquiring a virtual image, a real image, and target text data, wherein the virtual image is an image obtained by rendering a first virtual object in a virtual world using a virtual engine, the real image is an image obtained by photographing the real world, the virtual world is constructed based on the real world, and the target text data is used to describe a second virtual object contained in the target image; extracting the structure of the first virtual object contained in the virtual image to obtain the structural features of the virtual image; extracting the style of the real image to obtain the style features of the real image; and generating a target image based on the structural features, style features, and target text data, wherein the second virtual object in the target image has the same structure as the first virtual object and the same style as the real image. It is noteworthy that the target image is generated based on structural features, style features, and target text data. The structural features reflect the structure of the first virtual object, the style features reflect the style of the real image, and the target text data reflects the user's requirements for image generation. Therefore, the target image has the same structure and shape as the virtual image, the same style as the real image, and conforms to the user's intent. By combining virtual images, real images, and target text data, more realistic, diverse, and user-compliant target images can be generated. The content of the generated target image can be customized to simulate various scenarios according to application requirements, broadening the application scenarios of digital twins and achieving the technical effect of improving the fineness, diversity, and realism of the generated images. This solves the technical problem of low fineness, diversity, and realism of generated virtual images in related technologies.

[0019] It is worth noting that the general description above and the detailed description that follow are merely for illustrative purposes and do not constitute a limitation on this application. Attached Figure Description

[0020] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0021] Figure 1 This is a schematic diagram illustrating an application scenario of an image generation method according to an embodiment of this application;

[0022] Figure 2 This is a schematic diagram of a computer terminal as a computing node in a computing environment according to an embodiment of this application;

[0023] Figure 3 This is a flowchart of an image generation method according to an embodiment of this application;

[0024] Figure 4This is a schematic diagram illustrating an optional method of virtual image generation using an image generation model according to an embodiment of this application;

[0025] Figure 5 This is a flowchart of a training method for an image generation model according to an embodiment of this application;

[0026] Figure 6 This is a schematic diagram of an optional initial training of an image generation model according to an embodiment of this application;

[0027] Figure 7 This is a schematic diagram illustrating an optional retraining of an image generation model according to an embodiment of this application;

[0028] Figure 8 This is a flowchart of an optional image generation method according to an embodiment of this application;

[0029] Figure 9 This is a flowchart of an optional image generation method according to an embodiment of this application;

[0030] Figure 10 This is a flowchart of an optional image generation method according to an embodiment of this application;

[0031] Figure 11 This is a flowchart of an optional image generation method according to an embodiment of this application;

[0032] Figure 12 This is a schematic diagram of an image generation apparatus according to an embodiment of this application;

[0033] Figure 13 This is a schematic diagram of a training apparatus for an image generation model according to an embodiment of this application;

[0034] Figure 14 This is a schematic diagram of an optional image generation apparatus according to an embodiment of this application;

[0035] Figure 15 This is a schematic diagram of an optional image generation apparatus according to an embodiment of this application;

[0036] Figure 16 This is a schematic diagram of an optional image generation apparatus according to an embodiment of this application;

[0037] Figure 17 This is a schematic diagram of an optional image generation apparatus according to an embodiment of this application;

[0038] Figure 18 This is a structural block diagram of a computer terminal according to an embodiment of this application;

[0039] Figure 19This is a schematic diagram of an image processing system according to an embodiment of this application. Detailed Implementation

[0040] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0041] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0042] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:

[0043] Virtual images can refer to images generated by a basic rendering virtual engine, or virtual scenes or objects simulated by a computer program.

[0044] A real image can refer to an image obtained by optically imaging the real world through a camera, or an image obtained by optically imaging the real world through a camera or other optical devices.

[0045] Structural images can refer to various types of images that represent the outline and shape of an object, such as edge maps. They can contain information such as the object's edges, textures, and lines, and can be used for tasks such as object recognition, object segmentation, and object detection. In the field of computer vision, structural images can be used in image processing, image analysis, and other applications. Structural images can include edge maps, corner maps, line maps, etc.

[0046] Style can refer to the unique characteristics of an image in terms of its form of expression, color, composition, and emotional expression.

[0047] Text tags, which can be a text description or a few words, indicate the content or style information of an image. They can be a description of the image's content or style and can be used to help search engines or users find the images they want more quickly. Text tags can include information such as objects, scenes, emotions, colors, and styles to better describe the content characteristics of the image.

[0048] Mean squared loss can be the average of the squares of the differences between the predicted and actual values.

[0049] Stable diffusion (SD) is a type of generative model.

[0050] The structural feature extractor can be a pre-trained SD-based adapter used to inject structural information, thereby guiding the generation of image content and structure.

[0051] A style feature extractor can be an adapter based on pre-trained SD training, used to inject style information to guide style transfer.

[0052] A structure graph generator is a pre-trained generator that can return the corresponding structure image from a given image.

[0053] The initial noisy image can refer to the random noisy image that serves as input during the image generation process, which drives the generative model to generate new images.

[0054] A deep convolutional neural network (Visual Geometry Group, or VGG for short) is a type of convolutional neural network that is widely used in the field of image style transfer.

[0055] According to an embodiment of this application, an image generation method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0056] Considering that machine learning models consume significant computing resources on mobile terminals, the methods described above in this application can be applied to, for example... Figure 1 The application scenarios shown are not limited to these. In, for example... Figure 1In the application scenario shown, the machine learning model is deployed on server 10. Server 10 can connect to one or more clients 20 via a local area network (LAN), wide area network (WAN), internet connection, or other types of data network. Clients 20 may include, but are not limited to, smartphones, tablets, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices. Clients 20 can interact with users through a graphical user interface to invoke the large model, thereby implementing the method provided in this application embodiment.

[0057] It should be noted that, provided that the client device's operating resources can meet the deployment and operation conditions of the large model, the embodiments of this application can be performed on the client device.

[0058] Figure 2 This is a schematic diagram illustrating a computer terminal as a computing node in a computing environment according to an embodiment of this application, such as... Figure 2 As shown, computing environment 201 includes multiple computing nodes (such as servers) running on a distributed network (represented as 210-1, 210-2, ... in the diagram). Each computing node contains local processing and memory resources, and end user 202 can remotely run applications or store data within computing environment 201. Applications can be provided as multiple services 220-1, 220-2, 220-3, and 220-4 within computing environment 201, representing services "A", "D", "E", and "H", respectively.

[0059] End user 202 can provide and access services through a web browser or other software application on a client. In some embodiments, the provisioning and / or requests of end user 202 can be provided to ingress gateway 230. Ingress gateway 230 may include a corresponding agent to handle the provisioning and / or requests for services (one or more services provided in computing environment 201).

[0060] Services are provided or deployed based on various virtualization technologies supported by the computing environment 201. In some embodiments, services may be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar methods. VM-based virtualization can simulate a real computer by initializing a virtual machine, executing programs and applications without directly accessing any actual hardware resources. While the machine is virtualized by a virtual machine, container-based virtualization can launch containers to virtualize an entire operating system (OS), allowing multiple workloads to run on a single OS instance.

[0061] In one embodiment based on container virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). For example, such as Figure 2 As shown, service 220-2 can be equipped with one or more Pods 240-1, 240-2, ..., 240-N (collectively referred to as Pods). A Pod can include a proxy 245 and one or more containers 242-1, 242-2, ..., 242-M (collectively referred to as containers). One or more containers within a Pod handle requests related to one or more corresponding functions of the service. Proxy 245 typically controls service-related network functions such as routing and load balancing. Other services can also be equipped with similar Pods.

[0062] During operation, executing a user request from end user 202 may require calling one or more services in computing environment 201, and executing one or more functions of one service may require calling one or more functions of another service. For example... Figure 2 As shown, service "A" 220-1 receives user requests from terminal user 202 from ingress gateway 230. Service "A" 220-1 can call service "D" 220-2, and service "D" 220-2 can request service "E" 220-3 to perform one or more functions.

[0063] The aforementioned computing environment can be a cloud computing environment, where resource allocation is managed by cloud services, allowing functionality development without needing to consider implementation, adjustment, or server scaling. This computing environment allows developers to execute event-responsive code without building or maintaining complex infrastructure. Services can be partitioned into a set of functions that can automatically and independently scale, rather than scaling a single hardware device to handle potential loads.

[0064] Under the aforementioned operating environment, this application provides the following: Figure 3 The image generation method shown. Figure 3 This is a flowchart of an image generation method according to an embodiment of this application. Figure 3 As shown, the specific steps may include the following:

[0065] Step S302: Obtain virtual image, real image, and target text data.

[0066] Among them, virtual images are images rendered by a virtual engine to depict the first virtual object in the virtual world, real images are images captured in the real world, the virtual world is constructed based on the real world, and target text data is used to describe the second virtual object contained in the target image.

[0067] The aforementioned virtual engine can refer to a software system used to create and render images and scenes in a virtual world, such as the Unity Game Engine (Unity for short) and Unreal Engine (Unreal Engine). It can be used to simulate physical shapes, lighting effects, materials, etc. in the real world, thereby presenting a realistic virtual world. Virtual engines can be used in fields such as game development, virtual reality, and augmented reality, and can help developers quickly create realistic virtual worlds and provide rich interactive experiences.

[0068] The aforementioned first virtual object can refer to various elements or entities in the virtual world, representing objects created and rendered by the virtual engine. The first virtual object can be a virtual character, virtual scene, virtual item, etc. The first virtual object can also be determined according to actual needs, which is not limited here.

[0069] The aforementioned virtual image can refer to an image generated by rendering virtual objects in a virtual world through a virtual engine. The virtual image is generated by rendering a first virtual object. By rendering the first virtual object through a virtual engine, the first virtual object can present a realistic appearance and dynamic effects. Virtual images can be used in fields such as virtual reality, augmented reality, and game development. The specific content of the virtual image can be determined according to actual needs and is not limited here.

[0070] The aforementioned real images refer to images that reflect the real world, captured through photographic or video recording techniques. Real images are obtained by directly observing or recording real scenes, objects, and people, and possess objective authenticity and credibility. Compared to fictional, artistic, or digitally synthesized images, real images can more closely resemble the appearance and situation of the real world, accurately reflecting the appearance and characteristics of things. At the same time, real images may also be affected by factors such as lighting, angle, and post-processing, and thus have subjectivity and limitations, requiring users to conduct rational analysis and judgment.

[0071] The target image mentioned above can refer to a virtual image that the user needs to generate. The content of the target image can be specified by the user in advance and can be determined according to actual needs. No restrictions are imposed here.

[0072] The aforementioned second virtual object can refer to various elements or entities in the target image. The second virtual object can be specified by the user in advance. The second virtual object can be a virtual character, virtual scene, virtual item, etc. The second virtual object can also be determined according to actual needs, which is not limited here.

[0073] The aforementioned target text data can refer to text tags used to describe the second virtual object contained in the target image, keywords or phrases used to describe the image content, a text description or a few words used to indicate the content or style information of the image, and can include descriptive words, theme keywords or style information, which can accurately convey the theme or characteristics of the image or video.

[0074] In one optional embodiment, virtual images, real images, and target text data can be acquired to generate virtual images. Specifically, a virtual engine can be used to render a first virtual object in the virtual world to generate a virtual image; a real image can be obtained by taking pictures of scenes or objects in the real world using a camera or video camera, and the real image can include various scenes, objects, or people; the target text data can be a description of a second virtual object in the target image, and can include information such as the object's appearance, attributes, and location. The target text data can be acquired through manual annotation or automatic generation.

[0075] In relevant application scenarios, users can interact with the client application or webpage to upload real images and input target text data. Users can choose to use the client's camera to directly capture real images, select real images from the client's stored album, or retrieve the required real images from the cloud storage system. Users can use natural language to describe the text tags of the target image they want to generate, i.e., the target text data. The content of the target text data described by the user can include the content, style, resolution, format, storage size, etc. of the target image to be generated. Users can input target text data by entering text, selecting tags, or by voice. The input method of target text data can also be determined according to actual needs and is not limited here.

[0076] During the acquisition of virtual images, users can also select virtual objects in the virtual engine for rendering to generate virtual images. Users can choose a specific virtual engine and a first virtual object. Users can control the virtual engine to generate the first virtual object. Users can view the rendering results of the virtual image in real time and adjust the rendering results of the virtual engine in real time to directly obtain a satisfactory virtual image. Alternatively, users can select pre-rendered virtual images from the client's stored album, or obtain virtual images shared by other users based on the client's sharing function. The method of obtaining virtual images can be determined according to actual needs and is not limited here.

[0077] Through the above process, users can easily obtain virtual images, real images, and target text data to realize the function of virtual image generation, which can provide users with a richer and more diverse virtual experience.

[0078] After receiving the virtual image, real image, and target text data specified by the client, the server can integrate these three data points into a single dataset. Once the dataset is acquired, the server can preprocess the data, including but not limited to image cropping, scaling, and denoising, as well as text data encoding and standardization. Specifically, the server can use image processing algorithms to crop the original image, removing unnecessary parts and retaining the region of interest; it can scale the image to change its size; it can denoise the image to eliminate noise and improve image quality; and it can encode and standardize the text data, converting it into a format that the model can process. This preprocessing improves the data quality, making it suitable for subsequent virtual image generation. These preprocessing operations help the model better learn the features of the data, improving the quality and accuracy of the generated images. Optionally, if the client's resources meet the above requirements, the above process can also be performed on the client device; this is not limited here.

[0079] By following the steps above, a dataset containing virtual images, real images, and target text data can be obtained, which can provide data support for subsequent virtual image generation, thereby enabling the generation of images with virtual world characteristics.

[0080] Step S304: Extract the structure of the first virtual object contained in the virtual image to obtain the structural features of the virtual image.

[0081] The aforementioned structural features can refer to the feature information describing the structure of the first virtual object extracted from the virtual image. The structural features can be information such as the shape, size, position, and boundary of the first virtual object, information such as the relative positional relationship and hierarchical structure between objects, or appearance features such as the texture and color of the object. The structural features can be determined according to actual needs and are not limited here.

[0082] In one optional embodiment, the structure of the first virtual object contained in the virtual image can be extracted to obtain the structural features of the virtual image. Specifically, a virtual image dataset containing different types of virtual objects and their corresponding structural information can be obtained. Virtual objects can be created using 3D modeling software, and their structural information can be saved. Computer vision technology can be used to analyze and identify the virtual image to determine the first virtual object contained in the virtual image. Deep learning models such as convolutional neural networks can be used for object detection and classification. The structure of the virtual object can be extracted using image processing and computer graphics techniques. Edge detection, contour extraction, and other algorithms can be used to obtain the boundary and contour information of the virtual object. After obtaining the structural information of the virtual object, structural features, including shape features, texture features, color features, etc., can be further extracted using feature extraction algorithms (such as Scale-Invariant Feature Transform (SIFT) and Histogram of Oriented Gradients). Gradients (HOG, etc.) are used to obtain feature descriptions of images. Finally, the extracted structural features can be analyzed and processed. Machine learning techniques such as clustering and classification can be used to describe and classify virtual images. Through the above steps, the structure of the first virtual object contained in the virtual image can be extracted, and the structural features of the virtual image can be obtained, providing a foundation for subsequent virtual image generation and analysis.

[0083] Step S306: Extract the style of the real image to obtain the style features of the real image.

[0084] The aforementioned style features can refer to unique visual style features extracted from real images, including features such as tone, texture, and lines. Style features can help understand the stylistic characteristics of real images and apply them to the generation of virtual images, making the generated virtual images more realistic and artistic. By extracting the style features of real images, virtual image generation models can learn and imitate the style of real images, thereby generating more artistic and realistic virtual images.

[0085] In one optional embodiment, the style of the real image can be extracted. This helps preserve the style features of the real image when generating virtual images, making the generated virtual images more realistic and artistic. Specifically, the acquired real image can be preprocessed, including image resizing, grayscale conversion, and denoising, to facilitate subsequent style extraction. Deep learning models such as convolutional neural networks can be used to extract the style features of the real image. Pre-trained models, such as Visual Geometry Group (VGG) or Residual Network (ResNet), can be used to extract features at different levels of the image. The extracted style features can be represented as a vector or matrix, which can represent the style features of the real image. This facilitates the application of the extracted style features of the real image to the virtual image generation process. Style transfer networks and other methods can be used to transfer the style features of the real image to the generated virtual image. Through these steps, the style of the real image can be extracted, helping to preserve the style features of the real image when generating virtual images, thus facilitating the subsequent generation of virtual images with the style features of the real image.

[0086] Step S308: Generate the target image based on structural features, style features, and target text data.

[0087] The second virtual object in the target image has the same structure as the first virtual object and the same style as the real image.

[0088] In an optional embodiment, this application can generate a target image based on the structural features of a virtual image, the style features of a real image, and target text data. The structural features of the virtual image can be a model generated by 3D modeling software or a manually designed contour; the style features of the real image can be style information extracted through a style transfer algorithm; and the target text data can be text tags describing the image content. Specifically, feature extraction can be performed on the prepared structural features, style features, and target text data, extracting the structural features of the virtual image, the style features of the real image, and the features of the target text data respectively. Feature extraction can be performed using various deep learning networks, such as convolutional neural networks. The three extracted features can be fused to obtain a comprehensive feature representation, which can be accomplished through simple concatenation or a more complex feature fusion network. Finally, the comprehensive feature representation can be used to generate the target image, which can be achieved through generative adversarial networks, etc. The generative model completes the process. It can generate a new image based on the input feature representation while maintaining consistency between structural features, style features, and text data. Through the above steps, based on the structural features of the virtual image, structural information such as the shape and outline of different objects can be learned. Structural features help generate images with accurate shapes and layouts. The style features of the real image help learn image features of different styles, such as color, texture, and lighting. Style features can make the generated images more realistic and detailed. The target text data, as input, can help understand the user's needs and requirements, thereby better generating images that meet the user's intentions. By combining virtual images, real images, and text data, more realistic, diverse, and user-relevant images can be generated. The generated target images have broad application prospects in fields such as virtual reality, human-computer interaction, and design, and can provide users with a more personalized and vivid image generation experience.

[0089] For example, in a smart city scenario, users can acquire virtual city images, real city images, and target text data through the following steps, and generate a target city image. Users then use a virtual engine to render a first virtual object in the virtual world to obtain a virtual city image. This virtual city image can be a city scene constructed using virtual reality technology, such as a 3D city model created using modeling software. Users can also photograph a real-world city to obtain a real city image, which can be a cityscape captured by a camera, including streets and buildings. Simultaneously, target text data describing a second virtual object contained in the target image can be collected. This target text data can be a description of the second virtual object, such as "a park in the city" or "a school in the city." Finally, image processing techniques can be used to extract the structure of the first virtual object contained in the virtual city image to obtain a virtual city image. The structural features of city images can be extracted to obtain the style features of real city images. Finally, based on the structural features of virtual city images, the style features of real city images, and target text data, generative adversarial networks and other technologies can be used to generate target city images. The generated target image will combine the structure of the virtual city, the style of the real city, and the target text description to obtain a new city image with a sense of realism. For example, in the application scenario of smart cities, the above method can be used to generate a virtual city image that includes a virtual intelligent transportation system. By collecting target text data describing the transportation system, the structural features of the transportation system in the virtual city image are extracted, and the style features of the real city image are extracted. Then, a target city image with an intelligent transportation system is generated. The generated target image can be used to study traffic management and the design of intelligent transportation systems in smart cities.

[0090] For example, in a digital human scenario, users can generate a target human image by acquiring virtual human images, real human images, and target text data. Specifically, users can use a virtual engine to render a first virtual object in the virtual world to obtain a virtual human image. For example, a digital human can be created in the virtual world using virtual reality technology and then rendered. Users can also photograph people in the real world to obtain real human images, capturing the appearance and features of real people using cameras or other imaging devices. Target text data is used to describe a second virtual object contained in the target image, such as describing the characteristics, clothing, and posture of the virtual human in the target image. The structure of the first virtual object contained in the virtual human image can be extracted. Image processing techniques and feature extraction algorithms are used to identify the structural features of virtual characters, such as their contours and key points. Styles of real-life images can be extracted, and image processing and style transfer algorithms can capture stylistic features such as color, texture, and lighting. Finally, based on structural features, stylistic features, and target text data, techniques such as generative adversarial networks can be used to fuse the structural features of the virtual character image with the stylistic features of the real-life image to generate the target character image. For example, the contours of a virtual character can be combined with the textures of a real-life character to generate a realistic digital character image. Through these steps, the process of generating a target character image from virtual character images, real-life images, and target text data can be realized, thereby enabling the generation and application of virtual images in digital human scenarios.

[0091] In this application embodiment, an image generation method is provided. The method includes: acquiring a virtual image, a real image, and target text data, wherein the virtual image is an image obtained by rendering a first virtual object in a virtual world using a virtual engine, the real image is an image obtained by photographing the real world, the virtual world is constructed based on the real world, and the target text data is used to describe a second virtual object contained in the target image; extracting the structure of the first virtual object contained in the virtual image to obtain the structural features of the virtual image; extracting the style of the real image to obtain the style features of the real image; and generating a target image based on the structural features, style features, and target text data, wherein the second virtual object in the target image has the same structure as the first virtual object and the same style as the real image. It is noteworthy that the target image is generated based on structural features, style features, and target text data. The structural features reflect the structure of the first virtual object, the style features reflect the style of the real image, and the target text data reflects the user's requirements for image generation. Therefore, the target image has the same structure and shape as the virtual image, the same style as the real image, and conforms to the user's intent. By combining virtual images, real images, and target text data, more realistic, diverse, and user-compliant target images can be generated. The content of the generated target image can be customized to simulate various scenarios according to application requirements, broadening the application scenarios of digital twins and achieving the technical effect of improving the fineness, diversity, and realism of the generated images. This solves the technical problem of low fineness, diversity, and realism of generated virtual images in related technologies.

[0092] In the above embodiments of this application, extracting the structure of the virtual object contained in the virtual image to obtain the structural features of the virtual image includes: using a structure graph generator to identify the structure of the first virtual object in the virtual image and generating a structure image corresponding to the virtual image, wherein the structure image contains the structure of the first virtual object; and using a structure feature extractor to extract features from the structure image to obtain the structural features.

[0093] The aforementioned structure graph generator can refer to a tool or algorithm that can generate image structures based on input data or parameters. Structure graph generators can be based on deep learning techniques, such as Generative Adversarial Networks (GANs) or Variational Autoencoders (VAEs). Structure graph generators can be used to generate structure images corresponding to various types of images, including faces, animals, and landscapes. By learning the image structure features in the input dataset, new images are generated based on these features. Structure graph generators can control the content and quality of the generated structure images by adjusting input parameters or random noise.

[0094] The aforementioned structural feature extractor refers to a tool or algorithm used to extract structured features from image data. It can identify and extract key structural information in images, such as edges, corners, and textures. These structural features help computers understand the content and organization of images and generate more realistic and lifelike virtual images. Based on image processing and computer vision technologies, structural feature extractors utilize various algorithms and models to extract key structural information from images. These algorithms can include edge detection algorithms (such as Sobel and Canny), corner detection algorithms (such as the Harris corner detector), and texture feature extraction algorithms (such as Gabor filters and LBP). Structural feature extractors help to better understand and learn the structural information of images, thereby generating more realistic and accurate virtual images. Through the information extracted by the structural feature extractor, models can better capture the details and features of images, improving the quality and realism of the generated images.

[0095] In one optional embodiment, a structure graph generator can be used to identify the structure of a first virtual object in a virtual image, generating a structure image corresponding to the virtual image. Specifically, image data of the first virtual object can be input and processed by the structure graph generator to identify structural information in the image. The structure graph generator can analyze and process the input image, extracting the main structure and feature information in the image to generate a corresponding structure image. The generated structure image can contain the structure of the first virtual object, such as contours, shapes, and edges. The generated structure image can be used for subsequent feature extraction and image generation processes, providing a foundation for generating the final virtual image. Then, a structure feature extractor can be used to extract features from the structure image. Specifically, the process involves inputting the generated structural image data and processing it through a structural feature extractor. This extractor extracts structural features from the image using operations such as convolution and pooling to identify key features like texture, lines, and shapes. These extracted structural features can then be used in the subsequent virtual image generation process, helping the model better understand and generate images, thus improving the quality and realism of the generated images. Through these steps, the structure of the first virtual object can be effectively identified and relevant features extracted, providing a crucial foundation and support for subsequent virtual image generation. This structural and feature information helps the model better understand the image content and generate realistic virtual images.

[0096] In the above embodiments of this application, the structural feature extractor adopts a partial network layer structure of the image generator, and the initial network parameters of the structural feature extractor are the same as the initial network parameters of the partial network layer structure. The image generator is used to generate the target image, and the partial network layer is used to extract features from the data input to the image generator.

[0097] In an optional embodiment, the structural feature extractor in this application can adopt a portion of the network layer structure of the image generator. That is, the structural feature extractor can utilize a portion of the network layers of the image generator to extract the structural information of the image. The initial network parameters of the structural feature extractor can be the same as the initial network parameters of the portion of the network layer structure. This ensures that the structural feature extractor has an initial state similar to that of the image generator when it starts training, thereby helping to improve the convergence speed and performance of the model. The image generator can be a model used to generate target images. It can receive some input signals, such as random noise or conditional information, and then learn to generate virtual images corresponding to the input signals through model learning. In the above setting, the structural feature extractor can assist the image generator in generating more realistic virtual images. By extracting image structural features, it can help the image generator learn better and generate more realistic images. The above setting combines the structural feature extractor and the image generator. By sharing a portion of the network layer structure and initial network parameters, it can improve the performance and effect of the image generation model, thereby generating more realistic virtual images.

[0098] In the above embodiments of this application, the style of a real image is extracted to obtain style features of the real image, including: inputting the real image into a style feature extractor in the image generation module, and obtaining sub-style features output by at least one network layer in the style feature extractor; and concatenating at least one sub-style feature to obtain style features.

[0099] In one optional embodiment, a real image can be input into a style feature extractor in the image generation module, and sub-style features output from at least one network layer in the style feature extractor can be obtained. Specifically, a pre-trained deep learning model (such as VGG, ResNet, etc.) can be selected as the style feature extractor. The real image can be input into the style feature extractor, and the outputs of each network layer in the extractor can be obtained through forward propagation. At least one network layer output can be selected as the sub-style feature. The outputs of network layers closer to the bottom can be selected, as these layers capture more of the image's texture and local features. At least one sub-style feature can be concatenated to obtain a style feature. Specifically, the selected sub-style features can be processed by concatenating, adding, or performing other operations to obtain richer style information. Concatenation can be achieved by simply connecting the selected sub-style features along a certain dimension to form a larger style feature vector. The final style feature can be used for subsequent image generation tasks, such as style transfer and image synthesis. Through the above steps, features with rich style information can be extracted from real images, providing useful reference and guidance for subsequent image generation tasks.

[0100] In the above embodiments of this application, generating a target image based on structural features, style features, and target text data includes: using structural features, style features, and target text data to guide an image generator to generate a target image.

[0101] In one alternative embodiment, structural features, style features, and target text data can be used to guide an image generator to generate a target image. Specifically, structural features, style features, and target text data can be used as inputs, and the image generator model can generate the target image. Generative Adversarial Networks (GANs) or other image generation algorithms can be used to train the image generator. The image generator can be optimized and adjusted based on the differences between the generated image and the target image to make the generated image closer to the target image. Through the above steps, by using structural features, style features, and target text data to guide the image generator to generate a target image, more accurate and refined virtual image generation can be achieved.

[0102] Figure 4 This is a schematic diagram illustrating an optional method for generating virtual images using an image generation model according to an embodiment of this application, such as... Figure 4As shown, a virtual image can be input into a structure graph generator to obtain a structure image, and the structure image can be input into a structure feature extractor to obtain the structure features of the virtual image; a real image can be input into a style feature extractor to obtain the style features of the real image; and an initial noisy image, text labels (which can be the target text data mentioned above), the structure features of the virtual image obtained above, and the style features of the real image can be input into a pre-trained stable diffusion model (which can be the image generation model mentioned above) to obtain an output image (which can be the target image mentioned above).

[0103] According to an embodiment of this application, a training method for an image generation model is also provided. The image generation model is used to execute the image generation method described above. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than that shown here.

[0104] Figure 5 This is a flowchart of a training method for an image generation model according to an embodiment of this application, such as... Figure 5 As shown, the specific steps may include the following:

[0105] Step S502: Obtain the first training data and the second training data.

[0106] The first training data includes a first training image and the first text data corresponding to the first training image. The first training image includes a first virtual image or a first real image. The second training data consists of a second virtual image, a second real image corresponding to the second virtual image, and the second text data.

[0107] In one optional embodiment, first training data and second training data can be obtained. Specifically, a set of first training images can be collected. The first training images can be virtual images or real images. First text data corresponding to the first training images can be obtained. The first text data can be text tags describing the content of the first training images. The first training images and the corresponding first text data can be matched and combined to construct a first training dataset. Another set of second virtual images and corresponding second real images can also be collected. The second virtual images can be generated by a generative model. The second real images can be collected by means of taking pictures in the real world, etc. The second virtual images need to match the second real images, that is, they need to correspond or be consistent in content. Second text data corresponding to the second virtual images can be collected. The second text data can be text tags describing the content of the second virtual images and the second real images. The second virtual images, the second real images, and the corresponding second text data can be matched and combined to construct a second training dataset. Through the above steps, the first training data and the second training data can be obtained and used to train the virtual image generation model. These datasets will help the model learn the relationship between virtual images and text data, thereby improving the quality and accuracy of the generated images.

[0108] Step S504: Train the structural feature extractor and style feature extractor in the image generation model using the first training data.

[0109] In an optional embodiment, during the training of the virtual image generation model, the present application can utilize first training data to train the structural feature extractor and style feature extractor in the image generation model. This helps the model learn the structural and style features of the image, thereby improving the quality of the generated image. The image generation model includes a structural feature extractor and a style feature extractor. The structural feature extractor is used to extract structural information of the image, such as edges and contours; the style feature extractor is used to extract style information of the image, such as color and texture. The first training data can be input into the image generation model, and the structural feature extractor and style feature extractor can be updated through a backpropagation algorithm. The parameters of the structural and stylistic feature extractors are adjusted to enable them to better extract structural and stylistic features from images. Optimization algorithms, such as Adam, can be used to accelerate the training process. Furthermore, parameters can be tuned based on the loss observed during training to achieve better results. After training, evaluation metrics can be used to assess the performance of the structural and stylistic feature extractors, such as the sharpness and realism of the generated images. By using the first training data to train the structural and stylistic feature extractors in the image generation model, the model can learn better feature representations, thereby improving the quality and diversity of the generated images.

[0110] Step S506: The style feature extractor is retrained using the second training data to obtain the image generation model.

[0111] In one optional embodiment, during the training of the virtual image generation model, the style feature extractor can be retrained using second training data. The second training data includes a second virtual image, a second real image, and text labels for the images. The second virtual image is a virtual image generated by the previously trained generation model, and the second real image is a real image corresponding to the second virtual image. The text labels for the images are used to describe the image content. Specifically, the second virtual image and the second real image can be input into the style feature extractor. By comparing their feature representations, the loss function of the style feature extractor can be obtained. The loss function can include content loss and style loss, used to guide the style feature extractor to learn more... Good feature representation is achieved by using the backpropagation algorithm to update the parameters of the style feature extractor according to the loss function, enabling it to better distinguish the differences between virtual and real images. After multiple iterations of training, a retrained style feature extractor is obtained. The trained style feature extractor can be used to generate more realistic and stylized virtual images, thus obtaining a better image generation model. By using a second training data to retrain the style feature extractor, the performance of the generation model and the quality of the generated images can be improved, making the generated virtual images more realistic and stylized. The above process can be iteratively trained repeatedly, and the parameters can be updated according to the guidance of the loss function to continuously optimize the performance of the generation model.

[0112] In the above embodiments of this application, training the structural feature extractor and style feature extractor in the image generation model using first training data includes: using the structure graph generator in the image generation model to identify the structure of the first training object in the first virtual image, generating a first training structure image corresponding to the first training image, wherein the training structure image contains the structure of the first training object; using the structural feature extractor to extract features from the first training image to obtain the first structural feature of the first training image; using the style feature extractor to extract style features from the first training image to obtain the first style feature of the first training image; inputting the first structural feature, the first style feature, and the first text data into the image generator to obtain the first predicted image generated by the image generator; constructing a first loss function based on the first training image and the first predicted image; and adjusting the parameters of the structural feature extractor and the style feature extractor based on the first loss function.

[0113] In one optional embodiment, during the training process of the virtual image generation model, a structure graph generator can be used to identify the structure of the first training object in the first virtual image, generating a first training structure image corresponding to the first training image. This can be achieved by training a specific neural network model. The model can accept a real image as input and output the corresponding structure image. The generated structure image and the real image can be used for training to update the parameters of the structure feature extractor. The structure feature extractor can be a convolutional neural network, which learns how to extract structural information from the image during training. The style information corresponding to the real image and the structure image can also be used for training to update the parameters of the style feature extractor. The style feature extractor can be a pre-trained convolutional neural network, which learns the style features of the image through transfer learning. Then, the first training image can be input into the structure feature extractor, and the structural features of the image can be extracted using techniques such as convolutional neural networks to obtain the first structure feature. Simultaneously, the first... The training image is input into a style feature extractor to extract style features, resulting in the first style feature. The first structural feature, the first style feature, and the corresponding text data are then input into an image generator, which can be a generative adversarial network or a variational autoencoder. The image generator learns the feature distribution of the input data to generate new images, thus generating the first predicted image. The first training image and the first predicted image can be used as input to construct a first loss function. This loss function can include metrics such as pixel-level difference and structural similarity to measure the difference between the generated image and the real image. Based on the first loss function, the parameters of the structural feature extractor and the style feature extractor are adjusted using a backpropagation algorithm. By minimizing the loss function, the parameters of the extractors are optimized, making the generated image closer to the real image. These steps can be repeated iteratively to train the model until the loss function converges or the set number of training epochs is reached. Finally, a trained virtual image generation model is obtained, which can be used to generate new images.

[0114] In the above embodiments of this application, the style feature extractor is retrained using the second training data to obtain an image generation model, including: using the structure graph generator in the image generation model to identify the second training structure image and generate a second training structure image corresponding to the second virtual image, wherein the training structure image contains the structure of the first training object; using the structure feature extractor to extract features from the second training structure image to obtain the second structural features of the second virtual image; using the style feature extractor to extract style features from the second real image to obtain the second style features of the second real image; inputting the second structural features, the second style features, and the second text data into the image generator to obtain a second predicted image generated by the image generator; constructing a second loss function based on the second virtual image, the second real image, and the second predicted image; and adjusting the parameters of the style feature extractor based on the second loss function.

[0115] In one optional embodiment, during the training of the virtual image generation model, the style feature extractor can be retrained using second training data. Specifically, the style feature extractor can be retrained using the second training data, and a second dataset, which can be images of different objects or scenes, can be used to retrain the original style feature extractor to adapt to different style features. The trained style feature extractor can be used to construct an image generation model, and the retrained style feature extractor can be used to construct a new image generation model to generate virtual images that match the second set of training data. Then, the structural feature extractor can be used to extract features from the second training structural images to obtain the second structural features of the second virtual image. The style feature extractor can be used to extract the style features from the second real image. Feature extraction is performed on the lattice to obtain the second style features of the second real image. The second structural features, the second style features, and the second text data can be input into the image generator to obtain the second predicted image generated by the image generator. A second loss function can be constructed based on the second virtual image, the second real image, and the second predicted image to measure the difference between the generated image and the real image. The parameters of the style feature extractor can be adjusted based on the second loss function to optimize the training effect of the model and improve the quality of the generated image. Through the above steps, the virtual image generation model can be continuously optimized to improve the accuracy and realism of the generated image, making it more consistent with the structure and style features of the real image. This training process can help the model better learn and understand the structural and style information of the image, thereby generating more realistic and compliant virtual images.

[0116] In the above embodiments of this application, constructing a second loss function based on a second virtual image, a second real image, and a second predicted image includes: constructing a first sub-loss function based on the error between the second real image and the second predicted image; extracting features of the structure of a second object in the second predicted image using a structural feature extractor to obtain predicted structural features of the second predicted image; constructing a second sub-loss function based on the error between the second structural features and the predicted structural features; and weighting the first sub-loss function and the second sub-loss function to obtain the second loss function.

[0117] In one optional embodiment, during the training of the virtual image generation model, a first sub-loss function can be constructed based on the error between the second real image and the second predicted image. The difference between the second real image and the second predicted image can be calculated using methods such as pixel-level difference or structural similarity indices. This difference is then used as input to the first sub-loss function to optimize model training. A structural feature extractor can be used to extract features of the structure of the second object in the second predicted image, obtaining the predicted structural features of the second predicted image. The structural feature extractor can extract structural information of the second object in the second predicted image, including features such as contours and edges. Feature extraction facilitates structural feature comparison in subsequent steps. The predicted structural features can be compared with the predicted results. The error between structural features is used to construct a second sub-loss function, calculating the difference between the second structural feature and the predicted structural feature. Various structural similarity indices or feature distance metrics can be used. The difference between the second structural feature and the predicted structural feature can be used as the input to the second sub-loss function for model training. The first and second sub-loss functions can be weighted to obtain the second loss function. The first and second sub-loss functions can be weighted separately and then added together to obtain the second loss function. The second loss function can be used as the total loss function for model training to optimize the parameters of the entire generative model. Through the above steps, the error between the real image and the predicted image, as well as the difference in structural features, can be effectively utilized to optimize the virtual image generation model, improving the quality and accuracy of the generated image.

[0118] Based on the digital twin world, this application proposes a virtual image simulation model. This model uses a basic rendering physics engine to generate virtual images and further combines a small number of paired real-style images and generative model technology to controllably transform the basic virtual images into realistic images with a given real-world image style. The content of the generated images can be customized to simulate various scenarios according to application requirements. The proposed virtual image simulation model can simulate high-quality realistic images that meet the application requirements of digital twins, providing basic data support for subsequent training of 3D scene generation and realistic video stream generation models. At the same time, it broadens the application scenarios of digital twins and reduces manpower and time costs.

[0119] This application considers that images generated based on basic rendering virtual engines are often not realistic enough in terms of content and fine granularity. For example, trees in virtual images are often jumbled together. Furthermore, many details violate real-world physics, such as distorted lampposts in images generated by virtual engines. This fails to meet the requirements for the practical application of digital twins. At the same time, although images of real scenes can be continuously obtained, it is difficult to obtain a large number of real images that are pixel-level matched with virtual images. This poses a challenge to transferring the style of real images to virtual images. Traditional style transfer algorithms, such as Whitening and Coloring Transform (WCT), achieve the effect of style transfer by extracting style information from style images and injecting it into content images, but cannot achieve the generation process.

[0120] This application proposes a virtual image simulation model based on the digital twin world. This model utilizes a basic rendering physics engine to generate virtual images and further combines a small number of paired real-style images with generative modeling techniques to controllably transform the basic virtual images into realistic images with a given real-world image style. The content of these generated images can be customized to simulate various scenarios according to application requirements. The virtual image simulation model proposed in this application can simulate high-quality, realistic images that meet the application requirements of digital twins, providing basic data support for subsequent training of 3D scene generation and realistic video stream generation models. It also broadens the application scenarios of digital twins and reduces manpower and time costs. From an application perspective, the simulator's rendering function is comparable to traditional WCT. Style transfer algorithms perform well when performing transfer or rendering tasks for specific styles. However, the virtual image simulation model proposed in this application can be applied to general digital twin systems. The simulator relies on the capabilities of the renderer, which generates realistic images by leveraging the powerful renderer's capabilities. WCT, on the other hand, relies on neural networks to extract and fuse the style features of the image. However, it cannot effectively improve the shortcomings of virtual images, such as poor granularity and generation that does not conform to the physical laws of the real world. The virtual simulation model proposed in this application uses basic virtual images. It can controllably generate images that conform to our given real style using only images generated by the most basic renderer and any given real style image, and can also perform further content editing.

[0121] The basic virtual images used in this application can be obtained from a digital twin engine. This process can use some virtual engines (such as Unreal Engine, Cesium, etc.) for basic rendering and imaging, but the imaging effect is still far from that of real images. The image quality of basic virtual images is relatively rough, the style is monotonous, and the content in basic virtual images may contain content that does not conform to the real world due to the limitations of the model in the digital twin. The above defects make it impossible for users to feel the realism and diversity of basic virtual images. This application proposes a virtual image simulation generation model that can transform the content of basic virtual images into images with a real style specified by the user. Specifically, this application first uses a basic rendering engine to generate basic virtual image data. Then, using a self-developed image generation algorithm model, the obtained basic virtual images are transformed into more realistic images based on the given real style images. This process automatically generates realistic images that users need or for digital twin applications, providing basic data support for subsequent training of 3D scene generation and realistic video stream generation models. At the same time, it broadens the application scenarios of digital twins and reduces manpower and time costs.

[0122] The training of the image generation model proposed in this application can be mainly divided into two steps: obtaining structural generation and basic style transfer capabilities based on self-reconstruction, and further fine-tuning based on a small number of paired (virtual image-real image) image pairs to improve style transfer capabilities and generalization.

[0123] This method, based on its own reconstruction capabilities to generate structure and transfer basic style, can utilize a mixed set of unpaired real and virtual images as a training set. During each training iteration, training images can be extracted from this set. These images are then fed into a structure map generator (which can be pre-trained and have its parameters frozen) to extract structural information. Alternatively, other models can be used directly. Structural information can be expressed in various ways, such as image edges, image depth, and image segmentation results. Image edge extractors (e.g., the Canny detector), image depth detectors, and image segmentation models can be used to acquire structure. The training images themselves are then fed into a style extractor to obtain style information. If the training images have corresponding text labels, these are also injected. All three types of information are simultaneously injected into a pre-trained stable diffusion model to guide generation. Finally, the mean squared loss between the output image and the training image is calculated, along with the mean squared loss between the structure maps of the output and training images. During this process, the parameters of the pre-trained stable diffusion model and the structure map generator remain frozen and do not participate in training. The trainable parts are the structure extractor and the style extractor. Through this first stage of training, a structure extractor and a style extractor can be obtained.

[0124] Specifically, the style feature extractor can use a pre-trained VGG network. When inputting a style image, the outputs of three layers of the VGG network are selected and concatenated as the style features of the image. The three layers can be determined experimentally. Considering that the shallow layers of the VGG network characterize the structure more and the deep layers characterize the features more, this application can use the last three layers of the VGG network for feature concatenation. It can also be determined according to actual needs, which is not limited here. After normalization and alignment, it is fed into the stable diffusion model to guide the style of the output image. The structural feature extractor can use the first half of the network layer structure of the stable diffusion model and initialize the structural feature extractor with the corresponding parameters of the stable diffusion model. The structural features obtained after the structural image passes through the structural feature extractor are fed into the stable diffusion model to guide the structure and content of the output image. The mean square loss function can represent the pixel tensor of the image under RGB.

[0125] The process of further fine-tuning image pairs based on a small number of virtual and real image pairs to improve style transfer capability and generalization can include using a small number of virtual and real image pairs as the training set, freezing the parameters of the trained structure extractor, the pre-trained stable diffusion model, and the structure map generator, and training only the style extractor to further enhance the style transfer capability. During training, the structure image of the virtual image in the virtual image and real image pair can be extracted and fed into the structure extractor to provide structural information, while the real image is used as the input of the style extractor to provide style information. If the virtual image has a corresponding text label, it is also injected at the same time. The mean squared loss is calculated between the output image generated by the model and the real image to ensure the transfer of the real style. At the same time, the mean squared loss is calculated between the structure map extracted from the model image and the structure map of the virtual image to ensure that the structure and content still follow the virtual image. Through fine-tuning training, the style transfer capability of the style extractor can be further improved. The loss function can represent the pixel tensor under the RGB of the image.

[0126] The working principle of the Stable Diffusion model includes the following: the forward process of the Stable Diffusion model is to continuously add noise to an image, while the reverse process is to train a neural network to estimate the noise and then remove the noise to restore the image. In the reverse denoising process, in order to make the noise removal process continuously generate the image required by the user, the algorithm can model structural information and style information as conditional parts of the probability function and integrate them into the neural network.

[0127] The automated generation process of realistic images using the image generation model proposed in this application can be as follows: Given a basic virtual image and its corresponding text labels, and a real image with a specific style specified by the user (which does not need to be paired with the virtual image), the application requires the use of text labels during the training and use of the model. The text labels are descriptions of the image content. The structure image obtained by the virtual image through the structure graph generator can be used as the input of the structure extractor, the real image with a specific style can be used as the input of the style extractor, and the text labels can be used as text input. The model trained by the above training stage generates a realistic output image that conforms to the given text labels and virtual image in terms of content structure, and follows the style of the given real image.

[0128] In related technologies, the generation of realistic images relies solely on powerful renderers. These renderers utilize their capabilities to render the generated images as detailed and realistic as possible. However, renderers cannot change the unrealistic content of the video stream, nor can they generate images according to the style type specified by the user. Furthermore, traditional style transfer algorithms are not strong enough in terms of style transfer capabilities and have weak generalization. The solution proposed in this application can generate basic virtual images using only the most basic renderer and generate realistic images according to the image style given by the user, thus achieving more realistic and reliable content.

[0129] Based on its self-developed image generation technology, this application can transform basic virtual images into real images with a specified style. It proposes to retrain two adapters that fuse virtual images and real-style images respectively based on a pre-trained diffusion model, thereby achieving a controllable generation effect. The generated simulated images conform to the given virtual images and text in terms of content, allowing for more controllable image editing, including simulating various accidents and situations that do not exist in the real world. At the same time, the style images are based on the real world, which can, to some extent, correct the defects of virtual images generated by the physics engine that do not conform to the physical laws of the real world.

[0130] Figure 6 This is a schematic diagram illustrating an optional initial training of an image generation model according to an embodiment of this application, such as... Figure 6 As shown, a training image A (which may be the first training image mentioned above) can be input into a structure map generator to obtain a structure image, and the obtained structure image can be input into a structure feature extractor with trainable parameters to obtain the structure features of the training image A; the training image A can be input into a style feature extractor with trainable parameters; and the initial noise image, text labels, the structure features of the training image A obtained above, and the style features of the training image A can be input into a pre-trained stable diffusion model to obtain an output image B (which may be the first predicted image mentioned above), and the mean square loss L (which may be the first loss function mentioned above) between the output image B and the training image A can be determined to adjust the parameters of the structure feature extractor and the style feature extractor with trainable parameters.

[0131] Figure 7 This is a schematic diagram illustrating an optional retraining of an image generation model according to an embodiment of this application, such as... Figure 7As shown, a virtual image A (which can be the second virtual image mentioned above) can be input into a structure graph generator to obtain a structure image B. The obtained structure image B is then input into a structure feature extractor with frozen parameters to obtain the structure features of the virtual image A (which can be the second structure feature mentioned above). A real image (which can be the second real image mentioned above) can be input into a style feature extractor with trainable parameters to obtain the style features of the real image (which can be the second style feature mentioned above). An initial noisy image, text labels, the structure features of the virtual image A obtained above, and the style features of the real image A can be input into a pre-trained stable diffusion model to obtain an output image D (which can be the second predicted image mentioned above). The mean square loss L1 (which can be the first sub-loss function mentioned above) between the output image D and the real image C is determined. The same structure feature extractor is used to determine the structure graph corresponding to the output image D, and the mean square error L2 (which can be the second sub-loss function mentioned above) between the structure graph corresponding to the output image D and the structure image B is determined. The parameters of the style feature extractor with trainable parameters can then be adjusted based on the mean square loss L1 and the mean square error L2.

[0132] According to an embodiment of this application, an image generation method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0133] Figure 8 This is a flowchart of an optional image generation method according to an embodiment of this application, such as... Figure 8 As shown, the specific steps may include the following:

[0134] Step S802: Obtain virtual city image, real city image and target text data.

[0135] In one optional embodiment, virtual city images, real city images, and target text data can be acquired to generate a real city image with a virtual city style. Specifically, virtual city images can be collected. These virtual city images can be rendered using a virtual engine to depict a first virtual city in the virtual world. The virtual city images can include different angles, lighting conditions, and scenes to train the model to learn the characteristics of the virtual city. Real city images can also be collected. These real city images are photographs of real cities in the real world. They can also include different angles, lighting conditions, and scenes to train the model to learn the characteristics of the real city. Target text data can be collected. This target text data can be used to describe a second virtual city in the target image. The target text data can include information such as the city's name, geographical location, and architectural style. This information helps the model generate a virtual city image that matches the target text description. By collecting and integrating the above data, the characteristics of both virtual and real cities can be learned, and a virtual city image that matches the target text description can be generated. This achieves the goal of generating a real city image with a virtual city style in the virtual world.

[0136] Step S804: Extract the structure of the first virtual city contained in the virtual city image to obtain the structural features of the virtual city image.

[0137] In one optional embodiment, structural features of the virtual city image can be extracted to help better understand the composition and characteristics of the virtual city. Computer vision and image processing techniques can be used to extract features from the virtual city image, such as gray-level co-occurrence matrix (GLCM) and histogram of oriented gradients (HARCH), or deep learning methods such as convolutional neural networks (CNNs). Based on the extracted features, the structure of the virtual city image can be analyzed, including the layout of buildings, the distribution of roads, and the location of green spaces. The extracted structural features can be described using statistical analysis, cluster analysis, and other methods for subsequent analysis and application. Through the above steps, the structure of the first virtual city contained in the virtual city image can be extracted, obtaining the structural features of the virtual city image, thereby better understanding the composition and characteristics of the virtual city.

[0138] Step S806: Extract the style of the real city image to obtain the style features of the real city image.

[0139] In one optional embodiment, the style of real city images can be extracted to obtain style features. Deep learning techniques such as convolutional neural networks can be used to extract features from preprocessed real city image data. A pre-trained network model can be used, or a custom network structure can be designed according to specific needs. Based on feature extraction, the style features of real city images can be represented by calculating statistical data or using generative adversarial networks. The style features of images can be obtained by comparing statistical data of different images, such as color distribution and texture features. Finally, the obtained style features of real city images can be applied to the generation of virtual city images to achieve style transfer. By combining the style features of real city images with the content features of virtual city images, virtual city images with a real city style can be generated. Through these steps, the style of real city images can be extracted, style features can be obtained, and these features can be applied to the generation of virtual city images, making the virtual city images more realistic and lifelike.

[0140] Step S808: Generate a target city image based on structural features, style features, and target text data.

[0141] In one optional embodiment, structural features may include information such as the city's road layout and building distribution, which can be obtained through city map data or virtual city models. Style features may include visual features such as the city's hue, lighting, and texture, which can be extracted through statistical information from real city images or through style transfer algorithms. The target text data, i.e., the text labels of the image, can be converted into semantic information to guide the model in generating images that conform to the text description during the generation process. Natural language processing techniques can be used to convert text labels into vector representations or semantic information. In the process of generating the target city image, deep learning models such as generative adversarial networks can be used. The generator can be trained to generate images that conform to the input features, and a discriminator can be used to evaluate the realism of the generated image. During the training process, the extracted structural features, style features, and text information can be used as input. The generator will gradually adjust the generated image to approximate the target city image until a satisfactory generation effect is achieved. The process of generating the target city image based on the structural features of virtual city images, the style features of real city images, and target text data is a complex process that comprehensively utilizes deep learning and natural language processing techniques. It requires the extraction, transformation, and generation of various types of information to achieve high-quality generation of the target city image.

[0142] According to an embodiment of this application, an image generation method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0143] Figure 9 This is a flowchart of an optional image generation method according to an embodiment of this application, such as... Figure 9 As shown, the specific steps may include the following:

[0144] Step S902: Obtain the virtual avatar image, the real avatar image, and the target text data.

[0145] In one optional embodiment, virtual avatar images, real avatar images, and target text data can be acquired. The virtual avatar image is an image rendered by a virtual engine onto a first virtual avatar in the virtual world. It can be obtained by modeling and rendering a character model or scene in the virtual world. The real avatar image is an image captured by photographing a real avatar in the real world. It can be captured by images or videos taken by a camera. The target text data is used to describe the second virtual avatar of the target avatar image. It can be a textual description of the characteristics, attributes, actions, etc. of the target avatar. Specifically, the virtual avatar image can be generated by modeling and rendering the first virtual avatar in the virtual world using a virtual engine. The real avatar image can be captured by photographing in the real world using a camera. The target text data can be written according to the needs of the target avatar to describe the characteristics, attributes, actions, etc. of the target avatar. The virtual avatar image, real avatar image, and target text data can be aligned and matched to ensure consistency and relevance among them, which facilitates the subsequent generation of the target avatar image.

[0146] Step S904: Extract the structure of the first virtual image contained in the virtual image image to obtain the structural features of the virtual image image.

[0147] In one optional embodiment, during the structural feature extraction process of the virtual avatar image, the extracted structural features can be described using methods such as feature vectors and feature descriptors to mathematically express them for subsequent processing and analysis. The extracted structural features can be filtered and selected to remove redundant information and retain representative and distinguishable features. The extracted structural features can be matched with features in a database to identify and verify the structural features of the first virtual avatar in the virtual avatar image. Common matching methods include nearest neighbor matching and feature similarity matching. The matching results can be analyzed and evaluated to verify whether the structural features of the virtual avatar image have been accurately extracted, and to further analyze the composition and features of the virtual avatar. Through these steps, the structural features of the virtual avatar image can be effectively extracted, providing a foundation and support for subsequent virtual avatar generation and analysis.

[0148] Step S906: Extract the style of the real image to obtain the style features of the real image.

[0149] In one optional embodiment, the style of a real-world image can be extracted to obtain its style features. This helps generate more realistic and lifelike virtual images. A selected style extraction algorithm can be used to process the real-world image and extract its style features. These features typically include information about color, texture, and shape. The extracted style features can be represented mathematically so that subsequent virtual image generation algorithms can utilize these features to generate virtual images with similar styles. The style features of the real-world image can be fused with the image features generated by the virtual image generation algorithm to generate a virtual image with the style features of the real-world image. Through these steps, the style features of the real-world image can be effectively extracted and applied to virtual image generation, thereby producing more realistic and lifelike virtual images.

[0150] Step S908: Generate the target image based on structural features, style features, and target text data.

[0151] In one alternative embodiment, structural features, style features, and target text data can be fused together, and generative models such as generative adversarial networks or variational autoencoders can be used to generate target image images. Through the above steps, the specific implementation process of generating target image images based on structural features, style features, and target text data can be realized.

[0152] According to an embodiment of this application, an image generation method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0153] Figure 10 This is a flowchart of an optional image generation method according to an embodiment of this application, such as... Figure 10 As shown, the specific steps may include the following:

[0154] Step S1002: In response to the input command applied to the operation interface, display the virtual image, the real image, and the target text data on the operation interface.

[0155] Step S1004: In response to the generation command applied to the operation interface, the target image is displayed on the operation interface.

[0156] According to an embodiment of this application, an image generation method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0157] Figure 11 This is a flowchart of an optional image generation method according to an embodiment of this application, such as... Figure 11 As shown, the specific steps may include the following:

[0158] Step S1102: Obtain virtual image, real image and target text data by calling the first interface.

[0159] The first interface includes a first parameter, the value of which includes a virtual image, a real image, and target text data. The virtual image is an image obtained by rendering a virtual model in the virtual world using a virtual engine. The real image is an image obtained by taking pictures of the real world. The virtual world is built based on the real world. The target text data is used to describe the image content of the target image.

[0160] In one optional embodiment, a first interface can be invoked. The first interface includes a first parameter, the value of which includes a virtual image, a real image, and target text data. If a virtual image is selected, the system will invoke a virtual engine to render a virtual model in the virtual world to generate a virtual image. If a real image is selected, the system will acquire a pre-captured real-world image as input. If target text data is selected, the system will generate a corresponding image based on the image content described by the target text data. After invoking the first interface and acquiring the corresponding input data, subsequent image generation work can continue, such as image style transfer and image compositing, to generate the final virtual image. The entire process can be implemented using technologies such as deep learning models and computer vision algorithms.

[0161] Step S1104: Extract the structure of the virtual objects contained in the virtual image to obtain the structural features of the virtual image.

[0162] In one alternative embodiment, extracting the structural features of virtual objects can help to better understand and process virtual images. For each virtual image, the virtual objects in the virtual image can be segmented to separate different objects. This can be achieved through image segmentation algorithms. Through the above steps, the structure of the virtual objects contained in the virtual image can be extracted to obtain the structural features of the virtual image, thereby better understanding and utilizing the virtual image data.

[0163] Step S1106: Extract the style of the real image to obtain the style features of the real image.

[0164] In one alternative embodiment, the style of the real image can be extracted to make the virtual image more realistic and lifelike. Through the above steps, the style features of the real image can be extracted, bringing richer and more diverse style effects to the generation of virtual images.

[0165] Step S1108: Generate the target image based on structural features, style features, and target text data.

[0166] The second virtual object in the target image has the same structure as the first virtual object and the same style as the real image.

[0167] In an optional embodiment, the extracted structural features, style features, and target text data can be fused to obtain a complete feature vector representation of the target image. Through the above steps, the process of generating a target image based on structural features, style features, and target text data can be realized, which can more flexibly control the content and style of the generated image, and can also realize semantic understanding and visual expression of the target image.

[0168] Step S1110: Output the target image by calling the second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the target image.

[0169] In an optional embodiment, the target image can be output by calling a second interface. Specifically, the requirements and specifications of the generated target image can be determined, including the image size, resolution, color mode, etc. An appropriate second interface can be selected and called according to the requirements for generating the target image. The second interface can contain a series of parameters, including a second parameter. During the process of calling the second interface, the parameter value of the second parameter can be set to ensure that the parameter value includes the target image. It can include the path, name, or other relevant information of the input target image. With the set second parameter, the second interface can be called to generate the target image. The second interface can generate a target image that meets the requirements according to the parameter value setting. The generated target image can be saved in a specified path or database for later use or display. Through the above steps, the process of outputting the target image by calling the second interface can be realized, ensuring that the generated image meets the requirements and satisfies the user's needs.

[0170] According to an embodiment of this application, an image generation apparatus for implementing the above-described image generation method is also provided. Figure 12 This is a schematic diagram of an image generation apparatus according to an embodiment of this application, such as... Figure 12 As shown, the device includes: an acquisition module 1202, a first extraction module 1204, a second extraction module 1206, and a generation module 1208.

[0171] The system includes the following modules: an acquisition module for acquiring virtual images, real images, and target text data; a virtual image for rendering a first virtual object in a virtual world using a virtual engine; a real image for capturing images of the real world; a target text data for describing a second virtual object contained in the target image; a first extraction module for extracting the structure of the first virtual object in the virtual image to obtain its structural features; a second extraction module for extracting the style of the real image to obtain its style features; and a generation module for generating the target image based on the structural features, style features, and target text data, wherein the second virtual object in the target image has the same structure as the first virtual object and the same style as the real image.

[0172] The first extraction module is further used to identify the structure of the first virtual object in the virtual image using a structure graph generator, and generate a structure image corresponding to the virtual image, wherein the structure image contains the structure of the first virtual object; and to extract features from the structure image using a structure feature extractor to obtain structure features.

[0173] The structural feature extractor adopts a partial network layer structure of the image generator, and the initial network parameters of the structural feature extractor are the same as the initial network parameters of the partial network layer structure. The image generator is used to generate the target image, and the partial network layer is used to extract features from the data input to the image generator.

[0174] The second extraction module is further configured to input the real image into the style feature extractor in the image generation module, and obtain the sub-style features output by at least one network layer in the style feature extractor; and to concatenate the at least one sub-style features to obtain the style features.

[0175] The generation module is also used to guide the image generator to generate the target image by utilizing structural features, style features, and target text data.

[0176] According to an embodiment of this application, a training apparatus for an image generation model that implements the above-described image generation model training method is also provided. Figure 13 This is a schematic diagram of a training apparatus for an image generation model according to an embodiment of this application, such as... Figure 13 As shown, the device includes: an acquisition module 1302, a first training module 1304, and a second training module 1306.

[0177] The acquisition module is used to acquire first training data and second training data. The first training data includes a first training image and first text data corresponding to the first training image. The first training image includes a first virtual image or a first real image. The second training data consists of a second virtual image, a second real image corresponding to the second virtual image, and second text data. The first training module is used to train the structural feature extractor and style feature extractor in the image generation model using the first training data. The second training module is used to retrain the style feature extractor using the second training data to obtain the image generation model.

[0178] The first training module is further configured to: extract features from the first training image using a structural feature extractor to obtain the first structural features of the first training image; extract style features from the first training image using a style feature extractor to obtain the first style features of the first training image; input the first structural features, the first style features, and the first text data into an image generator to obtain the first predicted image generated by the image generator; construct a first loss function based on the first training image and the first predicted image; and adjust the parameters of the structural feature extractor and the style feature extractor based on the first loss function.

[0179] The second training module is further configured to: extract features from the second training structural image using a structural feature extractor to obtain the second structural features of the second virtual image; extract style features from the second real image using a style feature extractor to obtain the second style features of the second real image; input the second structural features, the second style features, and the second text data into an image generator to obtain the second predicted image generated by the image generator; construct a second loss function based on the second virtual image, the second real image, and the second predicted image; and adjust the parameters of the style feature extractor based on the second loss function.

[0180] The first training module is further used to construct a first sub-loss function based on the error between the second real image and the second predicted image; to extract features of the structure of the second object in the second predicted image using a structural feature extractor to obtain the predicted structural features of the second predicted image; to construct a second sub-loss function based on the error between the second structural features and the predicted structural features; and to perform weighted processing on the first sub-loss function and the second sub-loss function to obtain the second loss function.

[0181] According to an embodiment of this application, an image generation apparatus is also provided. Figure 14 This is a schematic diagram of an optional image generation apparatus according to an embodiment of this application, such as... Figure 14 As shown, the device includes: an acquisition module 1402, a first extraction module 1404, a second extraction module 1406, and a generation module 1408.

[0182] The system comprises the following modules: an acquisition module for acquiring virtual city images, real city images, and target text data; a virtual city image for rendering a first virtual city in a virtual world using a virtual engine; a real city image for photographing a real city in the real world; a virtual world constructed based on the real world; and target text data for describing a second virtual city within the target image. A first extraction module for extracting the structure of the first virtual city contained in the virtual city image, obtaining structural features of the virtual city image; a second extraction module for extracting the style of the real city image, obtaining style features of the real city image; and a generation module for generating the target city image based on the structural features, style features, and target text data.

[0183] According to an embodiment of this application, an image generation apparatus is also provided. Figure 15 This is a schematic diagram of an optional image generation apparatus according to an embodiment of this application, such as... Figure 15 As shown, the device includes: an acquisition module 1502, a first extraction module 1504, a second extraction module 1506, and a generation module 1508.

[0184] The system includes the following modules: an acquisition module for acquiring virtual avatar images, real avatar images, and target text data; a virtual avatar image for rendering a first virtual avatar in a virtual world using a virtual engine; a real avatar image for photographing a real avatar in the real world; a virtual world constructed based on the real world; and target text data for describing a second virtual avatar in the target avatar image. A first extraction module for extracting the structure of the first virtual avatar contained in the virtual avatar image to obtain the structural features of the virtual avatar image; a second extraction module for extracting the style of the real avatar image to obtain the style features of the real avatar image; and a generation module for generating the target avatar image based on the structural features, style features, and target text data.

[0185] According to an embodiment of this application, an image generation apparatus is also provided. Figure 16 This is a schematic diagram of an optional image generation apparatus according to an embodiment of this application, such as... Figure 16 As shown, the device includes: a first display module 1602 and a second display module 1604.

[0186] The first display module responds to input commands applied to the operation interface and displays virtual images, real images, and target text data on the interface. The virtual image is rendered using a virtual engine to visualize virtual models in a virtual world. The real image is captured from real-world photographs. The virtual world is constructed based on the real world. The target text data describes the image content of the target image. The second display module responds to generation commands applied to the operation interface and displays the target image. The target image is generated based on the structural features of the virtual image, the style features of the real image, and the target text data. The structural features are extracted from the structure of the virtual objects contained in the virtual image, and the style features are extracted from the style of the real image.

[0187] According to an embodiment of this application, an image generation apparatus is also provided. Figure 17 This is a schematic diagram of an optional image generation apparatus according to an embodiment of this application, such as... Figure 17 As shown, the device includes: a first calling module 1702, a first extraction module 1704, a second extraction module 1706, a generation module 1708, and a second calling module 1710.

[0188] The system comprises the following modules: a first calling module, used to obtain virtual images, real images, and target text data by calling a first interface, wherein the first interface includes a first parameter whose value includes the virtual image, real image, and target text data; the virtual image is an image rendered from a virtual model in a virtual world using a virtual engine; the real image is an image captured from the real world; the virtual world is constructed based on the real world; and the target text data describes the image content of the target image; a first extraction module, used to extract the structure of virtual objects contained in the virtual image to obtain the structural features of the virtual image; a second extraction module, used to extract the style of the real image to obtain the style features of the real image; a generation module, used to generate the target image based on the structural features, style features, and target text data; and a second calling module, used to output the target image by calling a second interface, wherein the second interface includes a second parameter whose value includes the target image.

[0189] Embodiments of this application may provide a computer terminal, which can be any computer terminal in a group of electronic devices. Optionally, in this embodiment, the aforementioned computer terminal may also be replaced by a mobile terminal or other terminal device.

[0190] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.

[0191] In this embodiment, the computer terminal described above can execute the program code in the method.

[0192] Optionally, Figure 18 This is a structural block diagram of a computer terminal according to an embodiment of this application, such as... Figure 18 As shown, the computer terminal A may include: one or more (only one is shown in the figure) processors 1802, memory 1804, memory controller, and peripheral interfaces, wherein the peripheral interfaces are connected to a radio frequency module, an audio module, and a display.

[0193] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0194] The processor can access information and applications stored in the memory via a transmission device to execute the steps in each embodiment.

[0195] According to embodiments of this application, an image processing system for implementing the above-described image processing method is also provided. Figure 19 This is a schematic diagram of an image processing system according to an embodiment of this application, such as... Figure 19 As shown, the system includes: virtual engine 1902, client 1904, and server 1906.

[0196] The system comprises: a virtual engine for rendering a first virtual object in a virtual world to obtain a virtual image; a client for sending a real image and target text data; a server connected to the virtual engine and the client for extracting the structure of the first virtual object in the virtual image to obtain its structural features, extracting the style of the real image to obtain its style features, and generating a target image based on the structural features, style features, and target text data; and a client for outputting the target image.

[0197] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0198] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0199] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods in the various embodiments of this application.

[0200] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store program code executed by the method provided in the above embodiments.

[0201] Optionally, in this embodiment, the storage medium may be located in any one of the electronic devices in the group of electronic devices in the computer network, or in any one of the mobile terminals in the group of mobile terminals.

[0202] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the steps in the various embodiments.

[0203] Embodiments of this application also provide a computer program product. Optionally, in this embodiment, the computer program product may include a computer program that, when executed by a processor, implements the methods provided in the embodiments described above.

[0204] The aforementioned computer program products can refer to software programs that have been written, tested, and released, and can run on computers or other devices. Computer program products can include application programs, operating systems, utility software, etc., used to achieve specific functions or solve specific problems.

[0205] Embodiments of this application also provide a computer program product. Optionally, the computer program product may include a non-volatile computer-readable storage medium, which can be used to store a computer program that, when executed by a processor, implements the method provided in the above embodiments.

[0206] The aforementioned non-volatile computer-readable storage medium can refer to a medium for storing data. Non-volatile computer-readable storage media can retain data without loss when power is off and can be used to store long-term data, such as operating systems, applications, and user files. Non-volatile storage media can include hard disk drives, solid-state drives, optical disks, and flash memory storage devices, etc.

[0207] Embodiments of this application also provide a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, it implements the method provided in the above embodiments.

[0208] The aforementioned computer program can refer to a set of instructions used to tell the computer to perform specific tasks or operations. Computer programs can be written by programmers using specific programming languages ​​and can include algorithms, data structures, logic, and control flow. Computer programs can be used for a variety of purposes, including application software, operating systems, etc.

[0209] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0210] It should be noted that the preferred embodiments involved in the above embodiments of this application are the same as the solutions, application scenarios and implementation processes provided in the above embodiments, but are not limited to the solutions provided in the above embodiments.

[0211] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0212] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0213] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0214] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0215] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. An image generation method, characterized in that, include: The method involves acquiring virtual images, real images, and target text data. The virtual image is an image rendered from a first virtual object in a virtual world using a virtual engine. The real image is an image captured from the real world. The virtual world is constructed based on the real world. The target text data is used to describe a second virtual object contained in the target image. The structure of the first virtual object contained in the virtual image is extracted to obtain the structural features of the virtual image; The style of the real image is extracted to obtain the style features of the real image; Based on the structural features, the style features, and the target text data, the target image is generated, wherein the second virtual object in the target image has the same structure as the first virtual object and the same style as the real image.

2. The method according to claim 1, characterized in that, The step of extracting the structure of the first virtual object contained in the virtual image to obtain the structural features of the virtual image includes: The structure of the first virtual object in the virtual image is identified using a structure diagram generator to generate a structure image corresponding to the virtual image, wherein the structure image contains the structure of the first virtual object; The structural features are obtained by extracting features from the structural image using a structural feature extractor.

3. The method according to claim 2, characterized in that, The structural feature extractor adopts a partial network layer structure of an image generator, and the initial network parameters of the structural feature extractor are the same as the initial network parameters of the partial network layer structure. The image generator is used to generate the target image, and the partial network layer is used to extract features from the data input to the image generator.

4. The method according to claim 1, characterized in that, The extraction of style from the real image to obtain style features of the real image includes: The real image is input into a style feature extractor, and sub-style features output by at least one network layer in the style feature extractor are obtained. The style feature is obtained by concatenating at least one sub-style feature.

5. The method according to claim 1, characterized in that, The process of generating the target image based on the structural features, the style features, and the target text data includes: Using the structural features, the style features, and the target text data, the image generator is guided to generate the target image.

6. A training method for an image generation model, characterized in that, The image generation model is used to execute the image generation method according to any one of claims 1 to 5, the method comprising: Acquire first training data and second training data, wherein the first training data includes a first training image and first text data corresponding to the first training image, the first training image includes a first virtual image or a first real image, and the second training data consists of a second virtual image, a second real image corresponding to the second virtual image, and second text data. The structural feature extractor and style feature extractor in the image generation model are trained using the first training data; The style feature extractor is retrained using the second training data to obtain the image generation model.

7. The method according to claim 6, characterized in that, The step of training the structural feature extractor and style feature extractor in the image generation model using the first training data includes: The structure of the first training object in the first virtual image is identified using the structure graph generator in the image generation model, and a first training structure image corresponding to the first training image is generated, wherein the training structure image contains the structure of the first training object. The structural feature extractor is used to extract features from the first training structural image to obtain the first structural feature of the first training image. The style feature extractor is used to extract style features from the first training image to obtain the first style features of the first training image. The first structural feature, the first style feature, and the first text data are input into the image generator to obtain the first predicted image generated by the image generator. A first loss function is constructed based on the first training image and the first predicted image; The parameters of the structural feature extractor and the style feature extractor are adjusted based on the first loss function.

8. The method according to claim 6, characterized in that, The step of retraining the style feature extractor using the second training data to obtain the image generation model includes: The structure graph generator in the image generation model is used to identify the second training structure image in the second virtual image to generate the second training structure image corresponding to the second virtual image, wherein the training structure image contains the structure of the first training object; The structural feature extractor is used to extract features from the second training structural image to obtain the second structural features of the second virtual image. The style feature extractor is used to extract style features from the second real image to obtain the second style features of the second real image; The second structural feature, the second style feature, and the second text data are input into the image generator to obtain the second predicted image generated by the image generator. A second loss function is constructed based on the second virtual image, the second real image, and the second predicted image; The style feature extractor is adjusted based on the second loss function.

9. The method according to claim 8, characterized in that, The construction of the second loss function based on the second virtual image, the second real image, and the second predicted image includes: Based on the error between the second real image and the second predicted image, a first sub-loss function is constructed; The structural feature extractor is used to extract features of the structure of the second object in the second predicted image to obtain the predicted structural features of the second predicted image. Based on the error between the second structural feature and the predicted structural feature, a second sub-loss function is constructed; The first sub-loss function and the second sub-loss function are weighted to obtain the second loss function.

10. An image generation method, characterized in that, include: The process involves acquiring virtual city images, real city images, and target text data. The virtual city images are rendered using a virtual engine to depict a first virtual city in a virtual world. The real city images are photographs of real cities in the real world. The virtual world is constructed based on the real world. The target text data describes a second virtual city within the target image. The structure of the first virtual city contained in the virtual city image is extracted to obtain the structural features of the virtual city image; The style of the real city image is extracted to obtain the style features of the real city image; Based on the structural features, the style features, and the target text data, a target city image is generated, wherein the second virtual city in the target city image has the same structure as the first virtual city and the same style as the real city image.

11. An image generation method, characterized in that, include: The process involves acquiring a virtual avatar image, a real avatar image, and target text data. The virtual avatar image is an image rendered by a virtual engine to depict a first virtual avatar in a virtual world. The real avatar image is an image captured from a real avatar in the real world. The virtual world is constructed based on the real world. The target text data is used to describe a second virtual avatar of the target avatar image. The structure of the first virtual image contained in the virtual image is extracted to obtain the structural features of the virtual image; The style of the real image is extracted to obtain the style features of the real image; Based on the structural features, the style features, and the target text data, the target image is generated, wherein the second virtual image in the target image has the same structure as the first virtual image and the same style as the real image.

12. An image generation method, characterized in that, include: The virtual image, real image, and target text data are obtained by calling a first interface. The first interface includes a first parameter, the value of which includes the virtual image, the real image, and the target text data. The virtual image is an image obtained by rendering a virtual model in a virtual world using a virtual engine. The real image is an image obtained by photographing the real world. The virtual world is constructed based on the real world. The target text data is used to describe the image content of the target image. The structure of the virtual objects contained in the virtual image is extracted to obtain the structural features of the virtual image; The style of the real image is extracted to obtain the style features of the real image; Based on the structural features, the style features, and the target text data, the target image is generated, wherein the second virtual object in the target image has the same structure as the first virtual object and the same style as the real image; The target image is output by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the target image.

13. A computer terminal, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method according to any one of claims 1 to 12.

14. An image processing system, characterized in that, include: A virtual engine is used to render a first virtual object in a virtual world to obtain a virtual image, wherein the virtual world is constructed based on the real world; A client is used to send real images and target text data, wherein the real images are images obtained by taking pictures of the real world, and the target text data is used to describe a second virtual object contained in the target image; The server, connected to the virtual engine and the client, is used to extract the structure of the first virtual object contained in the virtual image to obtain the structural features of the virtual image, extract the style of the real image to obtain the style features of the real image, and generate the target image based on the structural features, the style features and the target text data, wherein the second virtual object in the target image has the same structure as the first virtual object and the same style as the real image; The client is also used to output the target image.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the computer-readable storage medium is located to perform the method according to any one of claims 1 to 12.

16. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the method described in any one of claims 1 to 12.