Visual generation task processing method and visual generation model training method

WO2026200329A1PCT designated stage Publication Date: 2026-10-01CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/078726
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-25
Filing Date
2026-02-11
Publication Date
2026-10-01

Smart Images

  • Figure CN2026078726_01102026_PF_FP_ABST
    Figure CN2026078726_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a visual generation task processing method and a visual generation model training method. The visual generation task processing method comprises: acquiring task data of a target visual generation task, wherein the task data comprises a target image data pair and a target style description of a style to be generated, and the target image data pair is obtained on the basis of a target object image including a target object; superimposing the target image data pair with random noise to obtain target image noise, and on the basis of the target style description and the target image noise, generating, by means of a target visual generation model corresponding to the target visual generation task, a target style image of the target object in said style. By using a target image data pair as noise data and providing reference information on a target object to a target visual generation model, the target visual generation model can fully preserve features of the target object when generating a target style image, thereby fully meeting image generation requirements and improving generation quality.
Need to check novelty before this filing date? Find Prior Art

Description

Visual generation task processing methods and visual generation model training methods Technical Field

[0001] This disclosure relates to the field of deep learning technology, and in particular to a method for processing visual generation tasks and a method for training visual generation models. Background Technology

[0002] With the rapid development of computer technology and deep learning technology, deep learning models have shown great potential in visual generation tasks, and can perform different downstream visual generation tasks, such as generating images of different styles.

[0003] In existing technologies, when processing visual generation tasks, users typically input a text description of the image to be generated. A deep learning model then analyzes this text description to generate the corresponding image. However, this visual generation task relies on text descriptions, limiting the information that deep learning models can recognize and process. This limited information fails to fully meet the actual image generation requirements and retain the necessary objects, resulting in poor task processing quality. Therefore, a solution for visual generation tasks that can improve processing quality is urgently needed. Summary of the Invention

[0004] In view of the above, embodiments of this disclosure provide a method for processing visual generation tasks. One or more embodiments of this disclosure also relate to a method for training a visual generation model, an information processing method based on a visual generation model, a task platform, a computing device, an electronic device, a computer-readable storage medium, and a computer program product, in order to solve the technical defects existing in the prior art.

[0005] According to a first aspect of the present disclosure, a visual generation task processing method is provided, comprising:

[0006] Acquire task data for a target visual generation task, wherein the task data includes target image data pairs and a target style description of the style to be generated, and the target image data pairs are obtained based on a target object graph containing the target object;

[0007] The target image data pair is superimposed with random noise to obtain target image noise. Based on the target style description and the target image noise, the target visual generation model corresponding to the target visual generation task generates a target style image of the target object under the style to be generated. The target visual generation model is obtained by training at least one sample image data pair with a set sample style and the corresponding sample style description.

[0008] According to a second aspect of the present disclosure, a method for training a visual generative model is provided, comprising:

[0009] Obtain at least one image training sample with a defined sample style, wherein the image training sample includes sample image data pairs and corresponding sample style descriptions, the sample image data pairs are obtained based on sample object maps and sample style maps, and the sample style map is an image of a sample object in the sample object map under the defined sample style;

[0010] Noise is added to the sample image data pairs in each training sample to obtain the corresponding sample image noise;

[0011] Based on the sample style descriptions and corresponding sample image noise in each image training sample, the visual generation model to be trained is trained to obtain the target visual generation model after training.

[0012] According to a third aspect of the present disclosure, an information processing method based on a visual generative model is provided, applied to a task platform, comprising:

[0013] The device receives a model request sent by a terminal device, wherein the model request includes at least one of the following: a scene identifier of the target scene, scene input data of the target scene, and model specification parameters.

[0014] Based on the model request, a corresponding target visual generation model is determined from at least one visual generation model, wherein the at least one visual generation model is trained based on the visual generation model training method described above.

[0015] Deploy the target visual generation model, and based on the target visual generation model, construct a visual generation interface so that the terminal device can schedule the target visual generation model to execute the target visual generation task.

[0016] According to a fourth aspect of the present disclosure, a task platform is provided, including a request interface and a response unit;

[0017] The request interface is configured to receive a model request sent by a terminal device, wherein the model request includes at least one of the following: a scene identifier of the target scene, scene input data of the target scene, and model specification parameters.

[0018] The response unit is configured to determine a corresponding target visual generation model from at least one visual generation model based on the model request, wherein the at least one visual generation model is trained based on the visual generation model training method described above.

[0019] According to a fifth aspect of the present disclosure, a computing device is provided, comprising:

[0020] Memory and processor;

[0021] The memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, they implement the steps of the aforementioned visual generation task processing method, visual generation model training method, or information processing method based on visual generation model.

[0022] According to a sixth aspect of the present disclosure, an electronic device is provided, comprising:

[0023] The memory and processor are connected via a bus;

[0024] The memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, they implement the steps of the aforementioned visual generation task processing method, visual generation model training method, or information processing method based on visual generation model.

[0025] According to a seventh aspect of the present disclosure, a computer-readable storage medium is provided that stores a computer program / instructions, which, when executed by a processor, implement the steps of the above-described visual generation task processing method, visual generation model training method, or information processing method based on a visual generation model.

[0026] According to an eighth aspect of the present disclosure, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described visual generation task processing method, visual generation model training method, or information processing method based on a visual generation model.

[0027] This disclosure provides a method for processing a visual generation task, which involves acquiring task data for a target visual generation task. The task data includes target image data pairs and a target style description of the style to be generated. The target image data pairs are obtained based on a target object graph containing the target object. The target image data pairs are superimposed with random noise to obtain target image noise. Based on the target style description and the target image noise, a target style image of the target object under the style to be generated is generated using a target visual generation model corresponding to the target visual generation task. The target visual generation model is obtained by training at least one sample image data pair with a defined sample style and a corresponding sample style description.

[0028] One embodiment of this disclosure involves obtaining a target style description of the style to be generated and a target image data pair obtained based on a target object graph containing the target object. The target image data pair is then superimposed with random noise to obtain target image noise. A target visual generation model analyzes and processes the target style description and target image noise to generate a target style image of the target object under the style to be generated. Thus, in addition to the target style description, the target image data pair is treated as noise data, providing the target visual generation model with reference information about the target object. This enriches the information that the target visual generation model can analyze and process, enabling it to generate a target style image containing the target object. By fully preserving the characteristics of the target object when generating the target style image, it better meets the actual image generation requirements and improves the task processing quality of visual generation tasks. Attached Figure Description

[0029] Figure 1 is an application architecture diagram of a vision generation task processing provided in an embodiment of this disclosure;

[0030] Figure 2 is a flowchart of a visual generation task processing method provided in an embodiment of this disclosure;

[0031] Figure 3a is a schematic diagram of a target image data pair provided in an embodiment of this disclosure;

[0032] Figure 3b is a schematic diagram of a target style image of a mega-object style provided in an embodiment of this disclosure;

[0033] Figure 3c is a schematic diagram of a target style image with an underwater style provided in an embodiment of this disclosure;

[0034] Figure 4 is a schematic diagram of the inference stage of a target vision generation model provided in an embodiment of this disclosure;

[0035] Figure 5 is a flowchart of a visual generative model training method provided in an embodiment of this disclosure;

[0036] Figure 6a is a schematic diagram of a sample image data pair provided in an embodiment of this disclosure;

[0037] Figure 6b is a schematic diagram of the generation process of an image training sample provided in one embodiment of this disclosure;

[0038] Figure 6c is a schematic diagram of the inference stage of a target vision generation model provided in an embodiment of this disclosure;

[0039] Figure 7 is a flowchart of a visual generation task processing method provided in an embodiment of this disclosure;

[0040] Figure 8 is a flowchart of an information processing method based on a visual generative model provided in an embodiment of this disclosure;

[0041] Figure 9 is a schematic diagram of the structure of a task platform provided in an embodiment of this disclosure;

[0042] Figure 10 is a schematic diagram of a vision generation task processing device provided in an embodiment of the present disclosure;

[0043] Figure 11 is a schematic diagram of a visual generative model training device provided in an embodiment of the present disclosure;

[0044] Figure 12 is a structural block diagram of a computing device provided in an embodiment of the present disclosure;

[0045] Figure 13 is a structural block diagram of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0046] Numerous specific details are set forth in the following description to provide a full understanding of this disclosure. However, this disclosure can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this disclosure. Therefore, this disclosure is not limited to the specific implementations disclosed below.

[0047] The terminology used in one or more embodiments of this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this disclosure. The singular forms “a” and “the” as used in one or more embodiments of this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this disclosure refers to and includes any or all possible combinations of one or more associated listed items.

[0048] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this disclosure, and similarly, second may also be referred to as first. Depending on the context, the word “if” as used herein may be interpreted as “when”, “in response to a determination”, or “when…”.

[0049] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this disclosure are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0050] In one or more embodiments of this disclosure, a large model refers to a deep learning model with a large number of model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters. A large model can also be called a foundation model. It is pre-trained using large-scale unlabeled corpora to produce a pre-trained model with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.

[0051] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as natural language processing tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.

[0052] First, the terms and concepts involved in one or more embodiments of this disclosure will be explained.

[0053] FLUX is a generative framework based on diffusion models, a fundamental model for text-to-image generation, focusing on efficient and high-quality image and video generation. It significantly improves the speed and quality of generation tasks by optimizing the model architecture and inference process. FLUX employs an improved diffusion model architecture, reducing computational complexity during generation while maintaining high-quality results. Furthermore, by optimizing parallel computing and memory management in the inference stage, FLUX significantly improves generation speed, especially in high-resolution image and long video generation tasks. Additionally, FLUX supports distributed computing across multiple machines and GPUs, dynamically adjusting parallel strategies based on hardware configuration for efficient resource utilization.

[0054] Diffusion models are a type of generative model that has made significant progress in image, video, and audio generation tasks in recent years. Their core idea is to simulate the diffusion process in physics to gradually transform random noise into high-quality data samples (such as images or audio). The core idea of ​​diffusion models consists of two processes: the forward diffusion process, which gradually adds noise to the input data (such as an image) until it is transformed into completely random Gaussian noise; and the reverse generation process, which learns the distribution of noise and gradually recovers the original data from the random noise. This process is the inverse of the forward diffusion process, using a neural network to predict the noise at each step and progressively remove it.

[0055] VL (Vision-Language) models are a type of multimodal model that combines visual (e.g., image or video) and linguistic (e.g., text) elements, and are commonly used for multimodal understanding. The goal of VL models is to understand and generate cross-modal content, such as generating images from textual descriptions or extracting textual information from images.

[0056] SAM (Segment Anything Model) is a general-purpose image segmentation model designed to achieve the goal of "segmenting everything." Through large-scale data training and a flexible hint mechanism, SAM can perform high-quality segmentation of objects in any image without additional training. In other words, the core goal of SAM is to achieve zero-shot segmentation, that is, to segment objects in any image without fine-tuning for a specific task or dataset.

[0057] RMBG (Remove Background) is a technology or tool focused on removing backgrounds from images. It is also a type of image segmentation model and is widely used in image editing, design, e-commerce, and other fields. The goal of RMBG is to automatically or semi-automatically separate foreground objects from the background in an image to generate high-quality transparent background images.

[0058] ID: Identification, identity / attribute recognition.

[0059] LoRA (Low-Rank Adaptation) is a technique for fine-tuning large pre-trained models, particularly in Natural Language Processing (NLP). It reduces the number of parameters by introducing a low-rank matrix into the weight matrix of the pre-trained model, thereby significantly reducing computational and storage costs while maintaining model performance. The core idea of ​​LoRA is the assumption that the parameter changes required during fine-tuning can be approximated by a low-rank matrix.

[0060] Variational Autoencoder (AE) is a generative model that combines the advantages of probabilistic graphical models and deep learning. It is used to compress or decompress image features and is widely applied to generating data (such as images and text) and representation learning. The core idea of ​​VAE is to model the distribution of data by introducing latent variables and to optimize the model through variational inference.

[0061] xDiT (Scalable Inference Engine for Diffusion Transformers) is a scalable inference engine designed for Diffusion Transformers (DiTs) on large-scale multi-GPU clusters. It is a distributed inference engine specifically designed for DiTs, providing a set of efficient parallel methods and GPU kernel acceleration technology to significantly improve the performance of diffusion models in the inference stage to meet the needs of real-time inference. It aims to solve the latency and scalability problems faced by DiTs models in the inference process, especially the computational resources required when generating high-quality images and videos.

[0062] ControlNet is a control framework for generative models (such as Stable Diffusion) designed to precisely control the generation process by introducing additional conditional inputs (such as edge maps, depth maps, pose maps, etc.). The core objective of ControlNet is to provide a higher level of controllability and flexibility while maintaining generation quality.

[0063] It's important to note that in visual image generation tasks, users typically input a textual description of the image to be generated. A deep learning model then analyzes this textual description to generate the corresponding image. However, since this task relies on textual descriptions, the deep learning model can only recognize and process limited information, failing to adequately meet the actual image generation needs and retain the required objects. This results in poor image quality for visual image generation tasks. To better meet real-world image generation requirements, it's often necessary to preserve the object features of specific objects to generate corresponding stylistic images.

[0064] In practice, some text-based image generation models can achieve commercial-grade quality, such as Stable Diffusion (a generative model based on a diffusion model) and FLUX. These models enable controllable generation of styled images, such as ControlNet. Taking e-commerce scenarios as an example, in generating creative images while preserving product IDs, these methods currently face two main problems: first, they require collecting a large number of image data pairs, including the original product image and its corresponding creatively styled image; second, they require sufficient training resources. These two points significantly hinder the development of related task processing capabilities.

[0065] This embodiment of the disclosure can leverage the text-to-image capability of the FLUX model to implement lightweight product ID retention training and inference, eliminating the need for massive image data pairs. This enables low-cost generation of creative style images for product ID retention, such as giant object or underwater styles. Specifically, large model tools can be used to batch process the data required for training, automating the acquisition of creative product image data pairs. Furthermore, the LoRA-based product ID retention training and inference method shortens the development process for creative style images for product ID retention, reducing manpower and time costs, and enabling industrial application of creative product image generation. Additionally, parallel inference can be accelerated on multi-GPU machines, achieving a generation rate in seconds on the H2O inference machine.

[0066] To address the aforementioned technical problems, this disclosure provides a visual generation task processing method, a visual generation model training method, an information processing method based on a visual generation model, a task platform, a computing device, an electronic device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0067] Considering the large number of model parameters in large models and the limited computing resources of mobile terminals, the visual generation task processing method provided in this disclosure can be applied to the application architecture diagram shown in Figure 1, but is not limited thereto. In the application architecture shown in Figure 1, the large model is deployed on server 10. Server 10 can connect to one or more client devices 20 through a local area network (LAN), a wide area network (WAN), the Internet, or other types of data networks. Client devices 20 may include, but are not limited to, smartphones, tablets, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices. Client devices 20 can interact with users through a graphical user interface to invoke the large model, thereby implementing the method provided in this disclosure.

[0068] In this embodiment of the disclosure, the system consisting of a client device and a server can perform the following steps: The client device sends a target style description of the style to be generated and a target object image containing the target object to the server. The server acquires task data for the target visual generation task, wherein the task data includes target image data pairs and a target style description of the style to be generated, and the target image data pairs are obtained based on the target object image containing the target object; the target image data pairs are superimposed with random noise to obtain target image noise, and based on the target style description and the target image noise, a target style image of the target object under the style to be generated is generated through the target visual generation model corresponding to the target visual generation task, wherein the target visual generation model is obtained by training at least one sample image data pair with a defined sample style and a corresponding sample style description.

[0069] In one optional embodiment of this disclosure, the server is further configured to return the generated target style image to the client device; the client device is further configured to receive the target style image sent by the server.

[0070] By applying the scheme of this disclosure embodiment, a target style description of the style to be generated and a target image data pair obtained based on a target object map containing the target object are obtained. The target image data pair is superimposed with random noise to obtain target image noise. The target visual generation model analyzes and processes the target style description and target image noise to generate a target style image of the target object under the style to be generated. Thus, in addition to the target style description, the target image data pair is treated as noise data, providing the target visual generation model with reference information about the target object. This enriches the information that the target visual generation model can analyze and process, enabling it to generate a target style image containing the target object. When generating the target style map, the model fully preserves the characteristics of the target object, thereby fully meeting the actual image generation requirements and improving the task processing quality of the visual generation task.

[0071] It should be noted that, provided that the operating resources of the client device can meet the deployment and operation conditions of the large model, the embodiments of this disclosure can be performed on the client device.

[0072] Referring to Figure 2, Figure 2 shows a flowchart of a vision generation task processing method according to an embodiment of the present disclosure, which specifically includes the following steps.

[0073] Step 202: Obtain the task data for the target visual generation task, wherein the task data includes target image data pairs and target style descriptions of the style to be generated, and the target image data pairs are obtained based on the target object map containing the target object.

[0074] This disclosure applies to applications, websites, or mini-programs with visual generation task processing capabilities. The visual generation task processing capabilities are implemented on the application, website, or mini-program. For example, a website deploying a target visual generation model can implement visual generation task processing capabilities such as style image generation and style character creation. Alternatively, a third-party application can implement corresponding visual generation task processing capabilities by calling the deployed target visual generation model through an Application Programming Interface (API).

[0075] The target visual generation task is the task to be processed. This target visual generation task is used to instruct the controlled generation based on the task data to obtain a style-generated image of the corresponding style. For example, the target visual generation task can be a style image generation task, a style image creation task, etc.

[0076] The task data may include target image data pairs and a target style description for the style to be generated. The target image data pairs are obtained based on target object graphs containing target objects. These target objects refer to the objects whose features need to be preserved in the target visual generation task. For example, in an e-commerce scenario, the target object is the product whose features need to be preserved in the generated target style graph; in a virtual avatar creation scenario, the target object can be the image element whose features need to be preserved in the created virtual avatar. Additionally, the style to be generated refers to the style of the target visual generation task that aims to control the generated target style image, such as a giant object style or an underwater style. The target style description is the descriptive text of the style to be generated, used to describe relevant information about the style.

[0077] In one implementation, users can upload the style to be generated and the target object image whose features need to be retained via a front-end. The task processing platform generates a corresponding target style description based on the user-uploaded style. Furthermore, the task processing platform can copy the user-uploaded target object image and stitch it together in a set direction to obtain target image data pairs, thereby obtaining the task data for the target visual generation task (target image data pairs and target style description), for subsequent target style image generation.

[0078] In another implementation, users can also directly upload the target style description corresponding to the style to be generated, as well as the target object image whose features need to be retained, through the front end. The task processing platform can copy the target object image uploaded by the user and stitch it together in a set direction to obtain the target image data pair, thereby obtaining the task data (target image data pair, target style description) for the target visual generation task, and then perform subsequent target style image generation.

[0079] In another implementation, users can also directly upload target image data pairs and target style descriptions of the style to be generated through the front end to initiate a target visual generation task. That is, the copying and stitching of the target object image is implemented on the front end client. The task processing platform directly parses the target visual generation task to obtain the target image data pairs and target style descriptions of the style to be generated, and then performs subsequent target style image generation.

[0080] It should be noted that the target image data pair is obtained based on the target object image containing the target object. The target image data pair is a stitched image with two sub-images, each of which is a target object image containing the target object. That is, the target image data pair can be obtained by copying and stitching the target object image containing the target object. The target object image containing the target object uploaded by the user can have a solid color background or a non-solid color background. If it is a non-solid color background, it can be processed into a solid color background first, and then the target object image with a solid color background can be copied and stitched according to the set direction to obtain the corresponding target image data pair, thus avoiding the background content from interfering with the subsequent target style image generation process.

[0081] In one optional implementation of this embodiment, acquiring task data for the target vision generation task includes:

[0082] The target image data pair uploaded from the front end and the target style description of the style to be generated are used as the task data for the target visual generation task; or, the object image to be processed uploaded from the front end and the target style description of the style to be generated are obtained; the object image to be processed is segmented using an image segmentation model to obtain the target object image with a solid color background; the target object image with a solid color background is copied, and the copied target object image is stitched together according to a set direction to obtain the target image data pair; the target image data pair and the target style description are used as the task data for the target visual generation task.

[0083] In one implementation, the user can directly upload target image data pairs and a target style description of the style to be generated via the front end. In this case, the uploaded target image data pairs and target style description can be directly used as task data for the target visual generation task. In another implementation, the user can upload an image of the object to be processed and a target style description of the style to be generated via the front end. The image of the object to be processed can include the target object and background content. After receiving the image of the object to be processed, the task processing platform segments the image of the object to be processed using an image segmentation model to obtain a target object image with a solid color background. Then, it copies the target object image with the solid color background and stitches the copied target object image according to a set direction to obtain the target image data pairs in the task data. Here, the set direction refers to the arrangement direction of the different images in the target image data pairs, such as horizontal or vertical arrangement.

[0084] In practice, the image segmentation model can be any model capable of separating the foreground and background in an image and removing the background. For example, the image segmentation model could be a Rectangular Image Processing (RMBG) model. The image of the object to be processed uploaded from the front end is input into the RMBG model, and the RMBG model can output a target object image with the background content removed, i.e., a target object image with a solid color background. Of course, in practice, other background removal algorithms can also be used to segment the image of the object to be processed and obtain a target object image with a solid color background, such as BEN (Background Erasing Network), BiRefNet-HR (Bilateral Reference Network-High Resolution), and InSPyReNet (Interactive Spatial Pyramid Recurrent Network). The appropriate image segmentation model can be selected based on specific requirements (accuracy, speed, complexity, etc.).

[0085] For example, Figure 3a is a schematic diagram of a target image data pair provided in an embodiment of this disclosure. The target object is "high heels". The image of the object to be processed uploaded by the user through the front end is the image of the "high heels". However, the image of the "high heels" includes cluttered background content. The image of the "high heels" is input into the image segmentation model to obtain a target object image with a solid color background. The target object image with a solid color background is copied and stitched left and right to obtain the target image data pair shown in Figure 3a.

[0086] In this embodiment of the disclosure, the image of the object to be processed uploaded by the front end can be segmented by an image segmentation model to obtain a target object image with a solid color background. Then, the target image data pair in the task data can be obtained by copying and splicing. This can avoid the interference caused by cluttered background content, reduce the requirements for the image uploaded by the front end, adapt to most scenarios, and facilitate user operation.

[0087] Step 204: Superimpose the target image data pair with random noise to obtain target image noise, and based on the target style description and target image noise, generate a target style image of the target object in the style to be generated through the target visual generation model corresponding to the target visual generation task. The target visual generation model is obtained by training at least one sample image data pair with a set sample style and the corresponding sample style description.

[0088] It should be noted that the target visual generation model is a large model, requiring only a small number of samples to fine-tune the pre-trained model before it can be applied to the visual generation task of this disclosure embodiment. Specifically, the target visual generation model can be a text-to-image model, generating corresponding images based on input text. More specifically, the target visual generation model can be a diffusion model, such as FLUX. The core idea of ​​the diffusion model consists of two processes: the forward diffusion process gradually adds noise to the input image, ultimately transforming it into completely random Gaussian noise; the reverse generation process learns the distribution of noise and gradually recovers the original data from the random noise. This process is the reverse of the forward diffusion process, using a neural network to predict the noise at each step and gradually remove it.

[0089] In practice, the trained target visual generation model has the ability to receive text and generate corresponding images based on the input text, but it lacks the ability to receive images. To enable the target visual generation model to receive input images and thus preserve the features of the target object when generating target-style images, the inference process of the target visual model can be modified. The inference process of the target visual generation model is the inverse generation process of the diffusion model. In the inference stage, a latent space noise needs to be randomly initialized. The target visual generation model gradually denoises to obtain the final target-style image. Therefore, in this process, the randomly initialized latent space noise can be modified by combining the target image data pairs, so that the initialized latent space noise has prior knowledge of the target image data pairs, and thus has prior knowledge of the target object image uploaded from the front end, preserving the features of the target object when generating target-style images.

[0090] Specifically, the target image data can be superimposed with random noise to obtain target image noise. Then, the target style description and the target image noise are input into the target visual generation model corresponding to the target visual generation task. The noise is gradually removed to restore the target style image of the target object under the style to be generated. By superimposing the target image data with random noise, the features of the target object are provided to the target visual generation model, generating a target style image that retains the target object, resulting in higher quality and better meeting user needs.

[0091] In practical implementation, the target image data pairs can be encoded using an encoder to obtain latent space features. These latent space features are then superimposed with random noise to obtain the corresponding target image noise. The encoder can be any encoder capable of encoding image features, such as a VAE encoder (Variational Autoencoder), a key component of the Variational Autoencoder (VAE) model responsible for encoding the input image into a low-dimensional representation in the latent space, thus obtaining the latent space features. Of course, other encoders can also be used in practice, such as autoencoders (AEs) and generative adversarial networks (GANs).

[0092] It should be noted that the text-to-image capability of the FLUX model can be combined to implement lightweight object-preserving training and inference, without the need for massive image data pairs as support, and to achieve low-cost object-preserving creative style image generation, such as giant object style, underwater style, etc.

[0093] In one optional implementation of this embodiment, the target image data pairs are superimposed with random noise to obtain target image noise. Based on the target style description and the target image noise, a target style image of the target object under the style to be generated is generated through the target visual generation model corresponding to the target visual generation task, including:

[0094] The updated image data pair of the nth round is superimposed with random noise to obtain the updated image noise of the nth round, where n is initially set to 1, the updated image data pair is initially the target image data pair, and the updated image noise is initially the target image noise;

[0095] The target style description and updated image noise are input into the target visual generation model. The target visual generation model denoises the updated image noise according to the target style description and restores the corresponding candidate image data pair.

[0096] Increment n by 1, use the candidate image data pair as the updated image data pair for the nth round, return to the step of superimposing the updated image data pair for the nth round with random noise to obtain the updated image noise for the nth round, until the denoising stopping condition is met;

[0097] The image at the first position in the segmented candidate image data pair is taken as the target style image, where the first position is the position of the sample style map of the target visual generation model in the sample image data pair during the training phase.

[0098] In practice, the inference stage of the target visual generation model can obtain the final target style image through at least one round of denoising. The number of denoising rounds or the required accuracy of the target image data can be set in advance based on actual needs. For example, it can be configured to perform 20 rounds of denoising or 15 rounds of denoising, or the required accuracy of the target image data can be configured.

[0099] Specifically, the target image data pair is used as the first round of updated image data pair. The corresponding latent space features are encoded and superimposed with random noise to obtain the first round of updated image noise (i.e., target image noise). The target style description and this first round of updated image noise are input into the target visual generation model. The target visual generation model denoises the updated image noise according to the target style description, restoring the corresponding candidate image data pair. Then, this candidate image data pair is used as the second round of updated image data pair. The corresponding latent space features are encoded and superimposed with random noise to obtain the second round of updated image noise. The target style description and this second round of updated image noise are input into the target visual generation model. The target visual generation model denoises the updated image noise according to the target style description, restoring the corresponding candidate image data pair. This process continues until the denoising stopping condition is met, thus obtaining the target image data pair.

[0100] The denoising stopping condition can be pre-configured, such as reaching a set number of denoising processing rounds, or the accuracy of the candidate image data pairs reaching a set accuracy requirement.

[0101] In practice, after obtaining the target image data pair through at least one round of denoising, the first position in the target image data pair is the target style image that retains the target object. This first position is the position of the sample style map of the target visual generation model in the sample image data pair during the training phase. Therefore, when the denoising stopping condition is met, the image encoding features output by the target visual generation model can be decoded by the decoder to obtain the final candidate image data pair. The image at the first position in the candidate image data pair is then segmented as the target style image.

[0102] Following the previous example, the target image data pair is superimposed with random noise and subjected to multiple rounds of denoising. When the denoising stopping condition is met, the corresponding candidate image data pair can be obtained. Since the first position is the position of the sample style map of the target visual generation model in the sample image data pair during the training phase, the first position in the obtained candidate image data pair is the restored target style image. The image at the first position in the final candidate image data pair (such as the image on the right) is extracted as the target style image.

[0103] In this embodiment of the disclosure, by performing at least one round of denoising processing through the target visual generation model, the final target style image can be obtained, realizing the controllable generation of the target style image that retains the target object. Moreover, the number of rounds of denoising processing can be customized based on the processor performance of the task processing platform, that is, the number of rounds of denoising processing is adapted to the processor performance of the task processing platform, avoiding overload of the task processing platform.

[0104] In one optional implementation of this embodiment, the updated image data pair of the nth round is superimposed with random noise to obtain the updated image noise of the nth round, including:

[0105] Based on the set weights of the denoising process in the nth round, the updated image data pair in the nth round is superimposed with random noise to obtain the updated image noise in the nth round.

[0106] In the nth round of denoising, the weight of random noise in the set weight is greater than that in the (n+1)th round of denoising.

[0107] In actual implementation, weight values ​​for multiple rounds of denoising can also be configured. Based on the weight set for the nth round of denoising, the updated image data pair of the nth round is superimposed with random noise to obtain the updated image noise of the nth round.

[0108] In one implementation, the weight values ​​for different rounds of denoising can be the same. For example, if the weight of random noise is set to 0.8, then the weight of the updated image data pair is 0.2. In each round of denoising, the updated image data pair and the random noise are superimposed according to the weight value of 0.2 for the updated image data pair and the weight value of 0.8 for random noise.

[0109] In another implementation, since the trained target visual generation model only adapts to random noise input at the beginning of denoising, different weight values ​​can be configured for different rounds of denoising. Initially, the weight value of random noise is the largest and decreases gradually. That is, the weight value of random noise in the set weight of the nth round of denoising is greater than the weight value of random noise in the set weight of the (n+1)th round of denoising. In other words, at the beginning of denoising, random noise accounts for a large proportion, and as the denoising process progresses, the proportion of random noise gradually decreases, while the proportion of updated image data gradually increases.

[0110] For example, assuming 20 rounds of denoising, the weight values ​​for random noise can be set as follows: 0.9982, 0.9941, 0.9881, 0.9803, 0.9705, 0.9584, 0.9438, 0.9262, 0.9049, 0.8793, 0.8483, 0.8106, 0.7646, 0.7082, 0.6386, 0.5534, 0.4501, 0.3285, 0.1951, and 0.0702.

[0111] In this embodiment of the disclosure, the weight value of random noise can be configured to be the largest at the beginning and then reduced in each round, so that the target visual generation model can better adapt to the adjusted updated image noise, thereby obtaining more accurate denoising results, that is, more accurate candidate image data pairs, thereby improving the accuracy of the target style image.

[0112] In an optional implementation of this embodiment, the candidate image data pair is used as the updated image data pair for the nth round, including:

[0113] The image features at the second position in the candidate image data pair are updated to update the image noise, while the image features at the first position in the candidate image data pair remain unchanged, to obtain the updated image data pair in the nth round. Here, the second position is the position of the sample object image in the sample image data pair during the training phase of the target visual generation model.

[0114] In practice, the target style description and updated image noise are input into the target visual generation model. This model then performs denoising on the updated image noise based on the target style description to recover the corresponding candidate image data pairs. Afterward, the image features at the second position in the candidate image data pair (i.e., the position of the sample object image in the sample image data pair during training, without affecting the position of the target style image to be cropped) are updated with new image noise, while the image features at the first position in the candidate image data pair remain unchanged. This results in the updated image data pair, as mentioned above. In other words, the image features at the second position in the candidate image data pair generated in the previous round are updated with updated image noise, while the image features at the first position in the candidate image data pair remain unchanged, to obtain the updated image data pair for the next round.

[0115] Following the previous example, after each round of denoising, the left half of the candidate image data pair can be replaced with the noise of the updated image, while keeping the right half of the candidate image data pair unchanged. Subsequently, the right half of the final output candidate image data pair can be extracted as the final generated target style image.

[0116] In this embodiment of the disclosure, by replacing the image features at the second position in the selected image data pair (the position of the sample object image in the sample image data pair during the training phase) with the updated image noise, richer original image information can be preserved to a certain extent, avoiding information loss, and without affecting the image information at the first position in the candidate image data pair (that is, the position of the target style image to be cropped), the stability and generation effect of the target style image are improved, and it is more flexible and efficient.

[0117] For example, Figure 3b is a schematic diagram of a target style image in a mega-object style according to an embodiment of the present disclosure. As shown in Figure 3b, the target object image is an image of a "bottle". The target style image of the target object "bottle" in mega-object style can be generated through the above process, and the characteristics of the "bottle" are preserved in the target style image. The target object image is an image of a "high heel". The target style image of the target object "high heel" in mega-object style can be generated through the above process, and the characteristics of the "high heel" are preserved in the target style image.

[0118] Figure 3c is a schematic diagram of a target style image in an underwater style according to an embodiment of the present disclosure. As shown in Figure 3c, the target object image is an image of a "bottle". The target style image of the "bottle" in the underwater style can be generated through the above process, and the features of the "bottle" are preserved in the target style image. The target object image is an image of a "high heel". The target style image of the "high heel" in the underwater style can be generated through the above process, and the features of the "high heel" are preserved in the target style image.

[0119] It's important to clarify that during the inference phase of the target visual generation model, we combine target image data pairs with random noise. Based on the textural image model (i.e., the diffusion model), by introducing noise to incorporate the desired object features, we can achieve controllable generation of creative-style images while preserving object features without a complex training process. This method provides a low-cost solution for generating creative-style images that retain object features.

[0120] In one optional implementation of this embodiment, based on the target style description and target image noise, a target style image of the target object under the style to be generated is generated using the target visual generation model corresponding to the target visual generation task, including:

[0121] The target style description and target image noise are input into the target visual generation model. Under the distributed inference framework, the target visual generation model generates a target style image of the target object under the style to be generated. The inference process of the target visual generation model to generate the target style image is based on at least one parallel strategy under the distributed inference framework. The parallel strategy refers to the inference strategy in which the inference task is decomposed into multiple sub-tasks and allocated to different computing units for simultaneous execution through task decomposition and distributed computing resource coordination during the model inference process.

[0122] Parallel strategies are a key performance optimization method for deep learning systems, and can be used to improve system efficiency or solve resource constraints.

[0123] It should be noted that during the inference phase, a target style image of the target object in the style to be generated can be generated through a target visual generation model under a distributed inference framework. The inference process of the target visual generation model generating the target style image is based on at least one parallel strategy under the distributed inference framework.

[0124] In practical implementation, the distributed inference framework can be xDiT. xDiT provides a set of efficient parallel methods and GPU kernel acceleration technologies to meet real-time inference requirements. Its main features include: a hybrid parallel strategy, supporting the combined use of multiple parallel methods, including PipeFusion, Sequence Parallel, Data Parallel, and CFG Parallel. These methods can be flexibly combined according to hardware configuration and task requirements to achieve better performance; Parallel VAE: addressing the memory overflow problem of the decoding VAE module in the post-processing of diffusion models when generating high-resolution images, xDiT implements a Patch Parallel version of VAE, significantly improving the ability to process high-resolution images; and a development interface, xDiT provides a simple and flexible development interface to help users quickly support new DiT models. By reusing the logic of the open-source community diffusers library, developers can easily implement complex hybrid parallel strategies.

[0125] xDiT supports various parallel strategies, including: PipeFusion, a pipelined parallel method based on the characteristics of diffusion models, particularly suitable for multi-machine, multi-card environments with weak interconnects (such as PCIe / Ethernet), which significantly reduces communication by utilizing the input time redundancy between diffusion steps; Sequence Parallel, a hybrid sequence parallel method combining DeepSpeed-Ulysses and Ring-Attention, suitable for processing long sequences of DiT models; CFG Parallel, which maintains a constant parallelism of 2 when using classifier-free guidance (CFG), significantly improving inference efficiency; and Data Parallel, which allows xDiT to process multiple prompts or generate multiple images in parallel, further improving throughput.

[0126] In practical implementation, the PipeFusion parallel strategy can be used to leverage the temporal redundancy of the target visual generation model (i.e., the diffusion model) to reduce communication in a weakly interconnected multi-machine, multi-GPU environment. The Sequence Parallel parallel strategy, combined with DeepSpeed-Ulysses and Ring-Attention techniques, can accelerate long sequence inference tasks. The CFG Parallel parallel strategy separates and parallelizes the unconditional and conditional generation processes in classifier-free generation (CFG), improving inference efficiency. Patch Parallel VAE technology can be used to parallelize the post-processing decoding module of the target visual generation model to support the generation of high-resolution images. Based on hardware configuration and task requirements, the above parallel strategies can be dynamically combined to achieve distributed acceleration of the target visual generation model's inference process.

[0127] It should be noted that the PipeFusion parallel strategy can reduce the communication volume between diffusion steps through pipeline parallelism, making it suitable for multi-machine, multi-card environments with PCIe or Ethernet interconnects. Patch Parallel VAE technology avoids memory overflow issues by segmenting high-resolution images into multiple patches for parallel processing. Dynamically combined parallel strategies include flexible combinations of PipeFusion, Sequence Parallel, CFG Parallel, and Data Parallel to achieve near-linear scalability.

[0128] In the embodiments of this disclosure, at least one parallel strategy based on a distributed inference framework is used to accelerate the inference process of the target visual generation model, thereby generating target style images and significantly improving the inference efficiency of the target visual generation model.

[0129] Figure 4 is a schematic diagram of the inference stage of a target visual generation model provided in an embodiment of this disclosure. As shown in Figure 4, target image data pairs (obtained by copying and stitching a target object graph containing the target object) and target style descriptions are obtained. The target image data pairs are superimposed with random noise to obtain target image noise. Then, the target style description and target image noise are input into the target visual generation model (accelerated by a distributed inference framework). After multiple rounds of denoising processing, candidate image data pairs are output. The right image in the candidate image data pairs is extracted as the target style image that retains the target object.

[0130] In an optional implementation of this embodiment, after generating a target style image of the target object under the style to be generated using a target visual generation model corresponding to the target visual generation task, based on the target style description and target image noise, the method further includes:

[0131] Feed the target style image back to the front end;

[0132] Receive task feedback information sent by the front end, where the task feedback information is the information provided by the front end for the target style image;

[0133] Based on task feedback information, construct optimization sample data;

[0134] The target visual generation model is optimized and trained based on the optimized sample data.

[0135] In practice, task feedback information is provided after the target style image is generated and displayed to the front end. This feedback is based on subjective evaluations and requirements regarding the quality of the target objects retained in the image, the generated style content, or the degree of task completion. This feedback information provides detailed descriptions that help improve the model's output. Task feedback information is a crucial component of human-computer interaction, reflecting genuine feelings and expectations about the processing results of the visual generation task. By collecting and utilizing this feedback information, the performance of the visual generation model can be continuously optimized to more accurately meet practical needs and improve the quality and accuracy of image generation results.

[0136] It should be noted that an interactive feedback mechanism was used to construct optimization sample data. Based on the optimization sample data, the target visual generation model can be further optimized and trained, thereby improving the model performance of the target visual generation model and thus improving the processing accuracy and quality of the visual generation task.

[0137] This disclosure provides a visual generation task processing method. In addition to the target style description, the target image data pairs are treated as noise data to provide reference information of the target object for the target visual generation model. This enriches the information that the target visual generation model can analyze and process, enabling the target visual generation model to generate target style images containing the target object. When generating the target style map, the characteristics of the target object are fully preserved, thereby fully meeting the actual image generation needs and improving the task processing quality of the visual generation task.

[0138] It should be noted that the target visual generation model used in the embodiment shown in Figure 2 above can be trained using the visual generation model training method shown in Figure 5 below. Referring to Figure 5, Figure 5 shows a flowchart of a visual generation model training method provided according to an embodiment of this disclosure, specifically including the following steps.

[0139] Step 502: Obtain at least one image training sample with a set sample style, wherein the image training sample includes sample image data pairs and corresponding sample style descriptions, the sample image data pairs are obtained based on the sample object map and the sample style map, and the sample style map is the image of the sample object in the sample object map under the set sample style.

[0140] Specifically, the defined sample style is the style that the target visual generation model needs to learn during the training phase. In other words, it's the style that the target visual generation model needs to be able to generate images of the target style. That is, if we want to generate a certain style of creative image, we can use that style as the defined sample style, collect corresponding image training samples, and train the visual generation model. Once trained, the target visual generation model will be able to generate images of the corresponding target style. For example, the defined sample style could be a giant object style, an underwater style, etc.

[0141] In practice, at least one set of image training samples can be obtained. These training samples include sample image data pairs and corresponding sample style descriptions. The sample style description is a detailed description of the sample image data pair, indicating the content of the desired generated image. The sample image data pair is obtained by stitching together a sample object map and a sample style map; that is, a sample image data pair is a stitched image containing information from both types of images. The sample style map is the image of the sample object map under the set sample style. In other words, the sample object map and the sample style map can be stitched together in a set direction to form a sample image data pair. The sample object refers to the object whose features need to be preserved during the training phase.

[0142] In practice, image training samples can be uploaded by the front end. This means the user searches for sample style maps, extracts sample objects from them to obtain sample object images, and then concatenates the style map and the sample object image to obtain corresponding sample image data pairs. These sample image data pairs are then recognized to obtain the corresponding sample style description. The front end then uploads the sample image data pairs and their corresponding sample style descriptions to the training platform. Alternatively, the user can input the desired sample style into the front end, and the training platform automatically searches for and retrieves the corresponding sample style map, extracts sample objects from it to obtain sample object images, and then concatenates the style map and the sample object image to obtain corresponding sample image data pairs. These sample image data pairs are then recognized to obtain the corresponding sample style description.

[0143] In one optional implementation of this embodiment, obtaining at least one image training sample with a defined sample style includes:

[0144] Obtain at least one initial sample style map with a defined sample style;

[0145] Based on the candidate initial sample style map, a corresponding second sample image data pair is constructed, wherein the candidate initial sample style map is the initial sample style map corresponding to the candidate sample style, and the candidate sample style is any one of at least one set sample style;

[0146] Image analysis of the second sample image data pair is performed using the first visual language model to obtain the second sample style description of the second sample image data pair.

[0147] Based on the second sample image data pairs and the second sample style description, construct the second image training samples corresponding to the candidate sample styles.

[0148] In practice, the training platform can obtain at least one initial sample style map for a given sample style. This initial sample style map can be uploaded from the front end or obtained through searching. Then, for each given sample style, a corresponding sample image data pair is constructed based on its initial sample style map. For each given sample style, at least one corresponding initial sample style map can be obtained, and for each initial sample style map, a corresponding sample image data pair can be constructed. Then, the first visual language model performs image analysis on the sample image data pair to obtain the corresponding sample style description. The sample image data pair and the corresponding sample style description can then be used to construct the corresponding image training sample. In other words, a given sample style can construct at least one corresponding image training sample.

[0149] The first visual language model is any visual language model (VL) capable of recognizing and analyzing input images and outputting corresponding descriptive text. The second sample image data pair is input into the first visual language model, which performs image analysis on the pair and outputs a corresponding image description. This image description is then used as the second sample style description for the second sample image data pair, thus achieving annotation of the second sample image data pair.

[0150] In this embodiment of the disclosure, by obtaining an initial sample style map with at least one set sample style, and constructing corresponding sample image data pairs based on it, the first visual language model is then used to deeply analyze the sample image data pairs to obtain corresponding sample style descriptions. This successfully builds a bridge between images and language, integrates image data and semantic descriptions, and automatically constructs corresponding image training samples, providing rich training data for model training. It eliminates the need for manual collection and annotation of training data, automatically acquires a large amount of image training data, and greatly improves the model training speed and training effect.

[0151] In one optional implementation of this embodiment, obtaining at least one initial sample style map with a defined sample style includes:

[0152] Receive at least one styled sample uploaded from the front end;

[0153] Based on at least one set sample style, search for at least one corresponding initial sample style map.

[0154] In practice, users can upload at least one specified sample style via the front end, specifying the style that the visual generation model to be trained should learn. The trained target visual generation model then generates creative images in the specified sample style. The training platform can then automatically search for at least one initial sample style image corresponding to each specified sample style.

[0155] It should be noted that after the model is trained, a visual generative model can be obtained that can generate creative images according to a specified style. This model can not only capture the essence of the sample style, but also retain the features of a specific object.

[0156] In this embodiment, the training platform can automatically search for the corresponding initial sample style map based on the specified set sample style, thereby automatically constructing the corresponding image training samples. This eliminates the need for users to collect a large number of images, providing an automated image training sample construction process. It automatically completes the acquisition and construction from the set sample style to the initial sample style map and then to the image training samples, relieving users of the heavy task of collecting sample data. As a result, the requirements for the uploaded training data are also greatly reduced. Users are no longer required to prepare a large amount of comprehensive and finely classified training data in advance. They only need to simply specify the required set sample style to easily start the subsequent training process, which undoubtedly greatly lowers the training threshold.

[0157] Of course, in actual implementation, users can also directly upload the initial sample style map corresponding to each set sample style through the front end, and specify the image content that the visual generation model to be trained needs to learn. This embodiment does not limit this.

[0158] In one optional implementation of this embodiment, based on the candidate initial sample style map, a corresponding second sample image data pair is constructed, including:

[0159] The initial description text of the candidate initial sample style map is obtained by analyzing the style map of the candidate initial sample through the second visual language model. The second visual language model may be the same as or different from the first visual language model.

[0160] The initial description text is expanded using a language model to obtain the expanded description text.

[0161] The updated sample style map corresponding to the expanded descriptive text is generated by the candidate visual generation model, wherein the candidate visual generation model is the same as or different from the visual generation model to be trained.

[0162] The sample objects in the updated sample style map are concatenated with the updated sample style map to obtain the second sample image data pair.

[0163] The second visual language model is any visual language model (VL) capable of recognizing and analyzing input images and outputting corresponding descriptive text. The second visual language model can be the same as or different from the first visual language model. That is, the visual language model can be used to analyze the style map of candidate initial samples to obtain the initial descriptive text of the candidate initial sample style map. Subsequently, the same visual language model can be used to analyze the constructed second sample image data pairs to obtain the second sample style description of the second sample image data pairs. Alternatively, other visual language models can be used to extract the sample style description.

[0164] The language model is any model that can imitate and expand the input text, such as the GPT series (Generative Pre-trained Transformer) and BERT (Bidirectional Encoder Representations from Transformers) models.

[0165] The candidate visual generation model may be the same as or different from the visual generation model to be trained. That is to say, other trained visual generation models can be used to generate the corresponding updated sample style map based on the expanded descriptive text, or the current visual generation model to be trained can be used to generate the corresponding updated sample style map based on the expanded descriptive text.

[0166] In this embodiment, a second visual language model can be used to describe the candidate initial sample style map, accurately capturing key style features, compositional elements, and color matching rules in the candidate initial sample style map, transforming the complex information originally presented in visual form into a structured and organized text description. Then, the language model is used to imitate and expand the description, ensuring that the expanded content does not deviate from the style by imitating the original description, maintaining the consistency and coherence of the style. The expansion can add more details, context, and embellishments, so that the generated updated sample style map retains the original style while having richer information. Afterwards, a text-generated image model (candidate visual generation model) is used to generate an updated sample style map based on the enriched information, which can construct a brand-new image that is consistent with the original candidate initial sample style map but also has its own unique features. Then, the sample objects in the updated sample style map are spliced ​​with the updated sample style map to form a second sample image data pair, realizing the automated construction of rich and diverse sample image data pairs.

[0167] Of course, in actual implementation, the sample object map can also be directly extracted from the candidate initial sample style map, and the sample object map and the candidate initial sample style map can be stitched together to form the corresponding second sample image data pair. This disclosure does not limit this.

[0168] In an optional implementation of this embodiment, after generating the updated sample style map corresponding to the expanded descriptive text through the candidate visual generation model, the method further includes:

[0169] The updated sample style map is segmented using an image segmentation model to obtain sample objects with solid color backgrounds.

[0170] The sample objects in the updated sample style map are concatenated with the updated sample style map to obtain a second sample image data pair, including:

[0171] The sample object with a solid color background is stitched together with the updated sample style map according to the set direction to obtain the second sample image data pair.

[0172] In practice, after generating the updated sample style map corresponding to the expanded descriptive text, the updated sample style map can be segmented into its main body using an image segmentation model to obtain sample objects with solid color backgrounds. Then, these solid color background sample objects are concatenated with the updated sample style map according to a predefined direction to obtain a second sample image data pair. This predefined direction refers to a pre-configured arrangement of the sample objects and the updated sample style map, such as left-right or top-bottom arrangement.

[0173] The image segmentation model can be any model that can extract the main object from the updated sample style map, such as SAM, CenterNet (a single-stage keypoint detection method), etc.

[0174] For example, Figure 6a is a schematic diagram of a sample image data pair provided in an embodiment of the present disclosure. As shown in Figure 6a, the sample object "package" in the updated sample style map is extracted to obtain an image of the sample object "package" with a solid color background. This image is then stitched together with the updated sample style map to obtain the corresponding sample image data pair.

[0175] Figure 6b is a schematic diagram of an image training sample generation process provided in one embodiment of this disclosure. The process involves receiving at least one set sample style uploaded from the front end, and searching for at least one corresponding initial sample style map based on each set sample style. The initial sample style map is analyzed using a Visual Language Model (VL) to obtain initial descriptive text. The initial descriptive text is expanded using a language model to obtain expanded descriptive text. An updated sample style map corresponding to the expanded descriptive text is generated using a text-to-image model. The updated sample style map is then segmented using an image segmentation model to obtain sample objects with solid-color backgrounds. Finally, the sample objects with solid-color backgrounds are concatenated with the updated sample style map to obtain sample image data pairs.

[0176] It should be noted that by using various open-source models to perform the above-mentioned automated image training sample acquisition process, multiple image training samples can be obtained. These image training samples can then be used to fine-tune the visual generation model to be trained.

[0177] In this embodiment of the disclosure, the sample objects in the updated sample style map are spliced ​​together with the updated sample style map in a set direction to obtain corresponding sample image data pairs. The sample image data pairs provide the features of the sample objects that need to be retained and the features of the style image to be generated, thereby training the visual generation model to be trained. This enables the trained target visual generation model to have the ability to reconstruct the target style image of the corresponding style from noise, thus realizing the controllable generation of style images.

[0178] Step 504: Add noise to the sample image data pairs in each image training sample to obtain the corresponding sample image noise.

[0179] It should be noted that, based on the specific needs and objectives of the training process for the visual generative model to be trained, an appropriate noise type can be selected. Common noise types include, but are not limited to: Gaussian noise, which mimics random fluctuations in nature, such as thermal noise in electronic devices; salt-and-pepper noise, which simulates sudden interference that may occur during the transmission of image signals, manifested as pixels in the image randomly turning black or white; and Poisson noise, used to mimic statistical fluctuations caused by natural phenomena such as photon counting. Then, the determined noise type is applied to the sample image data pair. Specifically, based on the selected noise type, a noise matrix of the same size as the sample image data pair (original image) can be generated. For example, for Gaussian noise, the mean and variance need to be defined to generate noise values ​​that conform to a normal distribution; for salt-and-pepper noise, it randomly determines which pixel positions will become extreme values ​​(brightest or darkest). The generated noise matrix is ​​then added to the sample image data pair (original image), changing the pixel values ​​in the sample image data pair, thereby producing a noisy image, i.e., sample image noise.

[0180] In addition to adding noise, other data augmentation techniques, such as rotation, flipping, and scaling, can be combined to further increase the diversity of noise in the sample images. This is very important for improving the generalization ability and robustness of the model.

[0181] In one optional implementation of this embodiment, noise is added to the sample image data pairs in each image training sample to obtain the corresponding sample image noise, including:

[0182] Noise is added to the sample image data pair based on at least one noise dimension to obtain sample image noise of at least one size. Different noise dimensions correspond to sample image noise of different sizes, and sample image noise of different sizes is used to predict prediction image data pairs of different resolutions for the same sample image data pair.

[0183] In practice, the noise dimension refers to the number of dimensions of the noise vector input into the visual generation model to be trained. It determines the size of the noise in the corresponding sample image, allowing for the prediction of image data pairs at different resolutions from the same sample image data pair. By adjusting the number and size of the noise dimensions, the size and resolution of the generated sample image noise can be flexibly controlled to adapt to different visual generation task requirements. By adding noise to the sample image data pair using at least one noise dimension, sample image noise of different sizes can be obtained. Different noise dimensions can create noise samples of varying sizes. Thus, for a given sample image data pair, prediction of image data pairs at different resolutions can be achieved, enabling subsequent model training. Training at diverse resolutions helps improve the model's generalization ability, allowing it to better handle image data of different resolutions.

[0184] For example, for the same sample image data pair, predict image data pairs with resolutions of 512, 768, 1024, 1280, 1536, and 2048 respectively. By using multiple prediction image data pairs with different resolutions, the visual generative model to be trained can be trained.

[0185] Step 506: Based on the sample style descriptions and corresponding sample image noise in each image training sample, train the visual generation model to be trained to obtain the trained target visual generation model.

[0186] It should be noted that LoRA fine-tuning training can be performed on the visual generation model to be trained by using the sample style descriptions and corresponding sample image noise from multiple automatically obtained image training samples to obtain the trained target visual generation model.

[0187] LoRA fine-tuning training reduces the number of parameters by introducing a low-rank matrix into the weight matrix of the visual generation model to be trained, thereby significantly reducing computation and storage costs while maintaining model performance. The core idea is that the parameter changes required by the visual generation model to be trained during fine-tuning can be approximated by a low-rank matrix.

[0188] In practice, LoRA training parameters can be configured as follows: lora_r set to 16, lora_alpha set to 16, batch size set to 4, learning rate 1e-4, training using bf16 precision, training iterations of 3500, training on an 80G single GPU, and the objective function being FlowMatching. Of course, in practice, LoRA training parameters can be configured to other parameters based on actual needs. The above is merely an example, and this disclosure does not impose limitations.

[0189] In one optional implementation of this embodiment, the visual generation model to be trained is trained based on the sample style descriptions and corresponding sample image noise in each image training sample to obtain the trained target visual generation model, including:

[0190] The first sample style description and the first sample image noise are input into the visual generation model to be trained. The visual generation model to be trained performs denoising on the first sample image noise according to the first sample style description to restore the corresponding predicted image data pair. Here, the first sample style description and the first sample image noise are the sample style description and sample image noise included in the first image training sample. The first image training sample is any one of the image training samples.

[0191] Based on the predicted image data pair and the first sample image data pair in the first image training sample, the prediction loss of the visual generation model to be trained is determined. The model parameters of the visual generation model to be trained are adjusted based on the prediction loss. Then, the sample style description of the next image training sample of the first image training sample and the corresponding sample image noise are input into the adjusted visual generation model to be trained. The visual generation model to be trained is trained until the training stopping condition is met, and the target visual generation model is obtained after training.

[0192] In practice, during training, stopping conditions may include, but are not limited to, reaching a preset number of training epochs, the prediction loss converging to a preset range, or the model performance no longer showing significant improvement on the validation set. Once these stopping conditions are met, the trained target visual generation model is obtained. This target visual generation model can denoise noisy sample images based on the input sample style description, restoring high-quality target style images.

[0193] It should be noted that, to further improve the model's generalization ability, data augmentation techniques, such as random pruning, rotation, and flipping, can be introduced during training to increase the diversity of training samples. At the same time, regularization methods, such as L1 regularization and L2 regularization, can also be used to prevent overfitting.

[0194] Furthermore, after model training is complete, performance evaluation can be performed to verify its effectiveness in practical applications. Performance evaluation can be accomplished by calculating various performance metrics of the model on the test set, such as accuracy, recall, and F1 score. Through performance evaluation, we can understand the model's strengths and weaknesses, providing guidance for subsequent model optimization.

[0195] Figure 6c is a schematic diagram of the inference stage of a target visual generation model according to an embodiment of this disclosure. As shown in Figure 6c, sample image data pairs (sample object image and sample style image) with a set sample style and corresponding sample style descriptions are obtained. Noise is added to the sample image data pairs to obtain corresponding sample image noise. The sample style description and the corresponding sample image noise are input into the visual generation model to be trained. The visual generation model to be trained denoises the sample image noise according to the instructions of the sample style description to restore the corresponding predicted image data pairs. Based on the prediction loss between the predicted image data pairs and the sample image data pairs, the model parameters of the visual generation model to be trained are adjusted, and the model is trained so that the trained target visual generation model has the ability to restore style images that conform to the corresponding style and include the set object, thereby improving the model training effect and training efficiency.

[0196] In this embodiment, the sample data required for training can be batch-processed using large model tools, automatically acquiring sample image data pairs and corresponding sample style descriptions. These sample image data pairs, obtained by stitching together sample object images and sample style images, provide the visual generation model with features of the sample objects and the style images to be generated, guiding the model to reconstruct a style image that conforms to the corresponding sample style and includes the sample object. The trained target visual generation model can generate a style image based on the input target style description. This style image includes the target object image and a target style image with the corresponding style and including the target object. This ensures that the trained target visual generation model fully preserves the features of the target object when generating the target style image, thus closely meeting actual image generation needs and improving the task processing quality of visual generation. Furthermore, fine-tuning training based on LoRA can be performed, shortening the development process, reducing manpower and time costs, and enabling the industrial application of style image generation.

[0197] The following description, in conjunction with Figure 7, uses the application of the visual generation task processing method provided in this disclosure in a creative style product image generation scenario as an example to further illustrate the visual generation task processing method. Figure 7 shows a flowchart of the processing procedure of a visual generation task processing method provided in an embodiment of this disclosure, specifically including the following steps.

[0198] Step 702: Obtain the sample style image of the giant object style, extract the products from the sample style image of the giant object style, and stitch the product image with the solid color background to the sample style image of the giant object style to obtain sample image data pairs; analyze the sample image data pairs through the visual language model (VL) to determine the corresponding sample style description.

[0199] Step 704: Add noise to the sample image data pairs to obtain the corresponding sample image noise; input the sample style description and sample image noise into the raw image model to be trained, and use the raw image model to denoise the sample image noise according to the sample style description to restore the corresponding predicted image data pairs. The left side of the predicted image data pair is the predicted product image, and the right side is the predicted style image; based on the prediction loss between the predicted image data pair and the sample image data pair, adjust the model parameters of the raw image model to be trained, and continue to input the next giant object style sample style description and corresponding sample image noise into the raw image model to be trained, and train the raw image model to be trained until the training stopping condition is met, and obtain the target raw image model that has been trained.

[0200] Step 706: Obtain the style description of the giant object style and the product image containing the target product. Copy the product image and stitch it left and right to obtain the target image data pair, where both sides of the target image data pair are the product images.

[0201] Step 708: Superimpose the target image data with random noise to obtain target image noise; input the style description of the giant object style and the target image noise into the target text image model. The target text image model performs stepwise denoising processing on the input target image noise according to the style description of the giant object style to restore the target style image of the target product under the giant object style.

[0202] Step 710: Feed back the target style image to the front-end user.

[0203] It should be noted that steps 702-704 above describe the training process of the target text-based image model, while steps 706-708 describe the inference process of the target text-based image model. Both can be implemented on the same platform or on different platforms. Furthermore, steps 702-710 are merely illustrative examples using the giant object style; in actual implementation, other styles (such as the underwater style) can also be trained to achieve controllable generation of style images for the corresponding styles.

[0204] This disclosure provides a visual generation task processing method. The trained target text-to-image model only supports input text. The left and right sub-images of the obtained target style image are the same product, with the left side being the product image and the right side being the style image including the product. However, the target text-to-image model does not have the ability to input product images. Therefore, during the inference stage, the target image data pair (obtained by copying and splicing product images) is treated as noise data to provide product features to the target text-to-image model. This enriches the information that the target text-to-image model can analyze and process, enabling the target text-to-image model to generate a target style image that retains the product. When generating the target style image, the product features are fully preserved, thereby fully meeting the actual image generation needs and improving the generation quality of the style image.

[0205] Referring to Figure 8, Figure 8 shows a flowchart of an information processing method based on a visual generative model according to an embodiment of the present disclosure, applied to a task platform, specifically including the following steps 802-804.

[0206] Step 802: Receive a model request sent by the terminal device, wherein the model request includes at least one of the following: scene identifier of the target scene, scene input data of the target scene, and model specification parameters.

[0207] Step 804: Based on the model request, determine the corresponding target visual generation model from at least one visual generation model, wherein at least one visual generation model is trained based on the above-described visual generation model training method.

[0208] It should be noted that, based on the model request, determining the corresponding target visual generation model from at least one visual generation model can be done in several ways. One option is to search for the corresponding target visual generation model from at least one visual generation model included in the model library based on the model request. Another option is to train the target visual generation model based on the model request. Yet another option is to construct the target visual generation model based on the model request, which is not limited here.

[0209] For example, based on the scene identifier of the target scene, at least one pre-trained visual generation model can be found in the model library. Then, based on the model specification parameters, a visual generation model of the corresponding size can be selected from the at least one visual generation model. Finally, based on the scene input data of the target scene, the visual generation model of the corresponding size can be trained to obtain a target visual generation model suitable for user needs.

[0210] At least one visual generation model is obtained by training according to the above-described visual generation model training method. The embodiments of this disclosure and the embodiments in the specification of FIG5 are based on the same inventive concept. For specific methods, please refer to the above-described model training content, which will not be repeated here.

[0211] In one optional implementation of this embodiment, the model request includes a scene identifier of the target scene; based on the model request, determining the corresponding target visual generation model from at least one visual generation model includes:

[0212] Based on the scene identifier of the target scene, a target visual generation model suitable for the target scene is searched from the model library. The model library stores at least one visual generation model suitable for different visual generation scenes.

[0213] It should be noted that the model library is a database for storing and managing various pre-trained deep learning models. Multiple vision generation models adapted to different vision generation scenarios cover different application scenarios and needs. The model library allows users to select the appropriate model according to their needs, or directly use the model for vision generation tasks through API calls.

[0214] Multiple visual generation models adapted to different visual generation scenarios are stored in the model library, each specifically designed for different visual generation scenarios. Each model is optimized for a specific application environment, and any given visual generation model is trained using the aforementioned visual generation model training method. For example, based on the scene identifier "style image generation" of the target scene, a target visual generation model adapted to the style image generation scenario can be found in the model library.

[0215] In this embodiment of the disclosure, based on scene requirements, a target visual generation model that is suitable for the scene is accurately found through scene identification, so that the generated target style image is more accurate and fits the scene, thereby improving the user experience and the processing quality of visual generation tasks.

[0216] As an example, the task platform can provide target visual generation models for various scenarios, such as style image generation scenarios. It can provide corresponding target visual generation models based on model requests sent by terminal devices. Since the target visual generation model is selected from at least one visual generation model trained based on the above-mentioned visual generation model training method, the target visual generation model has undergone precise fine-tuning training and can achieve customized generation of creative style images that retain the target object.

[0217] In one optional implementation of this embodiment, the model request includes scene input data of the target scene; based on the model request, determining the corresponding target visual generation model from at least one visual generation model includes:

[0218] From at least one visual generation model, determine an initial visual generation model that is suitable for the target scene;

[0219] Based on the scene input data of the target scene, the initial visual generation model is trained to obtain the target visual generation model.

[0220] In actual implementation, the model request may include scene input data of the target scene, and the target visual generation model is a visual generation model suitable for the target scene.

[0221] For example, a general visual generation model is a basic visual generation model that is trained to adapt to different visual generation scenarios, but is not optimized for any specific scenario. For instance, based on scene input data of a style image generation scenario, the general visual generation model can be trained to obtain a target visual generation model adapted to the style image generation scenario.

[0222] In this embodiment of the disclosure, based on the requirements of the scene, a general visual generation model is further trained using scene input data to obtain a target visual generation model adapted to the scene, making the generated target style image more accurate and fit the scene, thereby improving the user experience and the processing quality of the visual generation task.

[0223] In one optional implementation of this embodiment, the model request includes model specification parameters; based on the model request, determining the corresponding target visual generation model from at least one visual generation model includes:

[0224] Based on the model specification parameters, the corresponding target visual generation model is searched from the model library, which stores multiple visual generation models with different model specification parameters.

[0225] Among them, the model specification parameter can be the model size, such as based on the model size: 32GB, to find the target visual generation model of the corresponding size from the model library.

[0226] In this embodiment of the disclosure, based on the model specification requirements, the corresponding target visual generation model is accurately found through the model specification parameters, which ensures the efficient and stable operation of the target visual generation model and improves the user experience.

[0227] In an optional implementation of this embodiment, after determining the corresponding target visual generation model from at least one visual generation model based on the model request, the method further includes:

[0228] Deploy the target visual generation model and build a visual generation interface based on the target visual generation model so that the terminal device can schedule the target visual generation model to execute the target visual generation task.

[0229] It should be noted that the visual generation interface is an interactive programming interface for terminal devices to schedule target visual generation models, usually provided in the form of an API. Through the visual generation interface, users can input task data for the target visual generation task, such as the target style description of the style to be generated, the target object graph containing the target object, etc., and effectively control the model's output, such as the generated target style image, etc.

[0230] In practice, one possible approach to deploying the target visual generation model is to deploy it on a distributed system of the task platform. For example, the target visual generation model can be deployed on a distributed system of the task platform, and a visual generation interface can be built based on the model and provided to the terminal device, enabling the terminal device to schedule the target visual generation model to execute the target visual generation task for the visual generation scene.

[0231] In this embodiment of the disclosure, efficient terminal invocation is achieved, the processing of visual generation tasks is optimized, and the processing quality and response speed of visual generation tasks are improved.

[0232] The information processing method based on visual generative models provided in this disclosure can be adapted to user needs to obtain target visual generative models, realize personalized model services, provide users with an efficient, flexible and easy-to-use model service method, and improve user experience.

[0233] Corresponding to the above method embodiments, this disclosure also provides a task platform embodiment. Figure 9 shows a schematic diagram of the structure of a task platform provided in one embodiment of this disclosure. As shown in Figure 9, the task platform includes: a request interface 902 and a response unit 904;

[0234] Request interface 902 is configured to receive model requests sent by terminal devices, wherein the model request includes at least one of the following: scene identifier of the target scene, scene input data of the target scene, and model specification parameters.

[0235] The response unit 904 is configured to determine a corresponding target visual generation model from at least one visual generation model based on a model request, wherein the at least one visual generation model is trained based on the visual generation model training method described above.

[0236] Optionally, the task platform also includes a visual generation interface, which is constructed based on the target visual generation model;

[0237] The vision generation interface is configured to allow terminal devices to schedule and execute target vision generation tasks.

[0238] Optionally, the model request includes a scene identifier for the target scene;

[0239] Response unit 904 is further configured as follows:

[0240] Based on the scene identifier of the target scene, a target visual generation model suitable for the target scene is searched from the model library. The model library stores at least one visual generation model suitable for different visual generation scenes.

[0241] Optionally, the model request includes scene input data for the target scene;

[0242] Response unit 904 is further configured as follows:

[0243] From at least one visual generation model, determine an initial visual generation model that is suitable for the target scene;

[0244] Based on the scene input data of the target scene, the initial visual generation model is trained to obtain the target visual generation model.

[0245] Optionally, the model request may include model specification parameters;

[0246] Response unit 904 is further configured as follows:

[0247] Based on the model specification parameters, the corresponding target visual generation model is searched from the model library, which stores multiple visual generation models with different model specification parameters.

[0248] Optionally, the task platform also includes a deployment module, configured as follows:

[0249] Deploy the target visual generation model and build a visual generation interface based on the target visual generation model so that the terminal device can schedule the target visual generation model to perform visual generation tasks.

[0250] In this embodiment of the disclosure, the task platform adapts to user needs to obtain target visual generation models, realizes personalized model services, provides users with an efficient, flexible and easy-to-use model service platform, and improves user experience.

[0251] The above is an illustrative scheme of a task platform according to this embodiment. It should be noted that the technical solution of this task platform belongs to the same concept as the technical solution of the information processing method based on the visual generative model described above. For details not described in detail in the technical solution of the task platform, please refer to the description of the technical solution of the information processing method based on the visual generative model described above.

[0252] Corresponding to the above method embodiments, this disclosure also provides an embodiment of a visual generation task processing device. Figure 10 shows a schematic diagram of the structure of a visual generation task processing device provided in one embodiment of this disclosure. As shown in Figure 10, the device includes:

[0253] The first acquisition module 1002 is configured to acquire task data for a target visual generation task, wherein the task data includes target image data pairs and a target style description of the style to be generated, and the target image data pairs are obtained based on a target object graph containing the target object.

[0254] The first generation module 1004 is configured to superimpose the target image data pair with random noise to obtain target image noise, and generate a target style image of the target object in the style to be generated based on the target style description and the target image noise, through the target visual generation model corresponding to the target visual generation task. The target visual generation model is obtained by training at least one sample image data pair with a set sample style and the corresponding sample style description.

[0255] Optionally, the first acquisition module 1002 is further configured as follows:

[0256] The target image data pair uploaded from the front end and the target style description of the style to be generated are used as the task data for the target visual generation task; or, the object image to be processed uploaded from the front end and the target style description of the style to be generated are obtained; the object image to be processed is segmented using an image segmentation model to obtain the target object image with a solid color background; the target object image with a solid color background is copied, and the copied target object image is stitched together according to a set direction to obtain the target image data pair; the target image data pair and the target style description are used as the task data for the target visual generation task.

[0257] Optionally, the first generation module 1004 is further configured as follows:

[0258] The updated image data pair of the nth round is superimposed with random noise to obtain the updated image noise of the nth round, where n is initially set to 1, the updated image data pair is initially the target image data pair, and the updated image noise is initially the target image noise;

[0259] The target style description and updated image noise are input into the target visual generation model. The target visual generation model denoises the updated image noise according to the target style description and restores the corresponding candidate image data pair.

[0260] Increment n by 1, use the candidate image data pair as the updated image data pair for the nth round, return to the step of superimposing the updated image data pair for the nth round with random noise to obtain the updated image noise for the nth round, until the denoising stopping condition is met;

[0261] The image at the first position in the segmented candidate image data pair is taken as the target style image, where the first position is the position of the sample style map of the target visual generation model in the sample image data pair during the training phase.

[0262] Optionally, the first generation module 1004 is further configured as follows:

[0263] Based on the set weights of the denoising process in the nth round, the updated image data pair in the nth round is superimposed with random noise to obtain the updated image noise in the nth round.

[0264] In the nth round of denoising, the weight of random noise in the set weight is greater than that in the (n+1)th round of denoising.

[0265] Optionally, the first generation module 1004 is further configured as follows:

[0266] The image features at the second position in the candidate image data pair are updated to update the image noise, while the image features at the first position in the candidate image data pair remain unchanged, to obtain the updated image data pair in the nth round. Here, the second position is the position of the sample object image in the sample image data pair during the training phase of the target visual generation model.

[0267] Optionally, the first generation module 1004 is further configured as follows:

[0268] The target style description and target image noise are input into the target visual generation model. Under the distributed inference framework, the target visual generation model generates a target style image of the target object in the style to be generated. The inference process of the target visual generation model to generate the target style image is based on at least one parallel strategy under the distributed inference framework. The parallel strategy refers to the inference strategy in which the inference task is decomposed into multiple sub-tasks and allocated to different computing units for simultaneous execution through task decomposition and distributed computing resource coordination during the model inference process.

[0269] Optionally, the device also includes a feedback module configured to:

[0270] Feed the target style image back to the front end;

[0271] Receive task feedback information sent by the front end, where the task feedback information is the information provided by the front end for the target style image;

[0272] Based on task feedback information, construct optimization sample data;

[0273] The target visual generation model is optimized and trained based on the optimized sample data.

[0274] This disclosure provides a visual generation task processing apparatus. It acquires a target style description of the style to be generated and a target image data pair obtained based on a target object image containing the target object. The target image data pair is superimposed with random noise to obtain target image noise. A target visual generation model analyzes and processes the target style description and target image noise to generate a target style image of the target object in the style to be generated. Thus, in addition to the target style description, the target image data pair is treated as noise data, providing the target visual generation model with reference information about the target object. This enriches the information that the target visual generation model can analyze and process, enabling it to generate a target style image containing the target object. By fully preserving the characteristics of the target object when generating the target style image, it better meets the actual image generation requirements and improves the task processing quality of visual generation tasks.

[0275] The above is a schematic scheme of a visual generation task processing device according to this embodiment. It should be noted that the technical solution of this visual generation task processing device and the technical solution of the above-described visual generation task processing method belong to the same concept. For details not described in detail in the technical solution of the visual generation task processing device, please refer to the description of the technical solution of the above-described visual generation task processing method.

[0276] Corresponding to the above method embodiments, this disclosure also provides an embodiment of a visual generative model training device. Figure 11 shows a schematic diagram of the structure of a visual generative model training device provided in one embodiment of this disclosure. As shown in Figure 11, the device includes:

[0277] The second acquisition module 1102 is configured to acquire at least one image training sample with a set sample style. The image training sample includes a sample image data pair and a corresponding sample style description. The sample image data pair is obtained based on the sample object map and the sample style map. The sample style map is the image of the sample object in the sample object map under the set sample style.

[0278] The noise addition module 1104 is configured to add noise to the sample image data pairs in each image training sample to obtain the corresponding sample image noise.

[0279] The training module 1106 is configured to train the visual generation model to be trained based on the sample style description and corresponding sample image noise in each image training sample, so as to obtain the trained target visual generation model.

[0280] Optionally, training module 1106 is further configured as follows:

[0281] The first sample style description and the first sample image noise are input into the visual generation model to be trained. The visual generation model to be trained performs denoising on the first sample image noise according to the first sample style description to restore the corresponding predicted image data pair. Here, the first sample style description and the first sample image noise are the sample style description and sample image noise included in the first image training sample. The first image training sample is any one of the image training samples.

[0282] Based on the predicted image data pair and the first sample image data pair in the first image training sample, the prediction loss of the visual generation model to be trained is determined. The model parameters of the visual generation model to be trained are adjusted based on the prediction loss. Then, the sample style description of the next image training sample of the first image training sample and the corresponding sample image noise are input into the adjusted visual generation model to be trained. The visual generation model to be trained is trained until the training stopping condition is met, and the target visual generation model is obtained after training.

[0283] Optionally, the noise-adding module 1104 is further configured as follows:

[0284] Noise is added to the sample image data pair based on at least one noise dimension to obtain sample image noise of at least one size. Different noise dimensions correspond to sample image noise of different sizes, and sample image noise of different sizes is used to predict prediction image data pairs of different resolutions for the same sample image data pair.

[0285] Optionally, the second acquisition module 1102 is further configured as follows:

[0286] Obtain at least one initial sample style map with a defined sample style;

[0287] Based on the candidate initial sample style map, a corresponding second sample image data pair is constructed, wherein the candidate initial sample style map is the initial sample style map corresponding to the candidate sample style, and the candidate sample style is any one of at least one set sample style;

[0288] Image analysis of the second sample image data pair is performed using the first visual language model to obtain the second sample style description of the second sample image data pair.

[0289] Based on the second sample image data pairs and the second sample style description, construct the second image training samples corresponding to the candidate sample styles.

[0290] Optionally, the second acquisition module 1102 is further configured as follows:

[0291] Receive at least one styled sample uploaded from the front end;

[0292] Based on at least one set sample style, search for at least one corresponding initial sample style map.

[0293] Optionally, the second acquisition module 1102 is further configured as follows:

[0294] The initial description text of the candidate initial sample style map is obtained by analyzing the style map of the candidate initial sample through the second visual language model. The second visual language model may be the same as or different from the first visual language model.

[0295] The initial description text is expanded using a language model to obtain the expanded description text.

[0296] The updated sample style map corresponding to the expanded descriptive text is generated by the candidate visual generation model, wherein the candidate visual generation model is the same as or different from the visual generation model to be trained.

[0297] The sample objects in the updated sample style map are concatenated with the updated sample style map to obtain the second sample image data pair.

[0298] Optionally, the second acquisition module 1102 is further configured as follows:

[0299] The updated sample style map is segmented using an image segmentation model to obtain sample objects with solid color backgrounds.

[0300] The sample object with a solid color background is stitched together with the updated sample style map according to the set direction to obtain the second sample image data pair.

[0301] This disclosure provides a visual generative model training device that can batch process the sample data required for training using large model tools, automatically acquiring sample image data pairs and corresponding sample style descriptions. The visual generative model to be trained is then trained using these sample image data pairs and style descriptions. These sample image data pairs are obtained by stitching together sample object images and sample style images, providing the visual generative model with features of the sample objects and the style images to be generated. This guides the visual generative model to reconstruct a style image that conforms to the corresponding sample style and includes the sample object. The trained target visual generative model can generate a style image based on the input target style description. This style image includes the target object image and a target style image with the corresponding style and including the target object. This allows the trained target visual generative model to fully retain the features of the target object when generating the target style image, thereby fully meeting the actual image generation requirements and improving the task processing quality of visual generative tasks.

[0302] The above is an illustrative scheme of a visual generative model training device according to this embodiment. It should be noted that the technical solution of this visual generative model training device and the technical solution of the visual generative model training method described above belong to the same concept. For details not described in detail in the technical solution of the visual generative model training device, please refer to the description of the technical solution of the visual generative model training method described above.

[0303] Figure 12 shows a structural block diagram of a computing device provided in one embodiment of the present disclosure.

[0304] The computing device 1200 includes:

[0305] Memory 1210 and processor 1220;

[0306] The memory 1210 is configured to store computer programs / instructions, and the processor 1220 is configured to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor 1220, they implement the steps of the above-described visual generation task processing method, visual generation model training method, or information processing method based on visual generation model.

[0307] In one or more embodiments of this disclosure, the computing device 1200 can be understood as an integrated smart terminal, including but not limited to a server, desktop computer, PC (Personal Computer), all-in-one model machine, mobile phone, tablet computer or other portable smart terminal, etc., and the computing device may have a model as described in the above embodiments of this disclosure pre-installed.

[0308] Specifically, the computing device 1200 can pre-install various types of models, including but not limited to models in natural language processing, visual processing, speech processing, code processing, and multimodal task processing, thus providing diverse model selection. In different product forms, the computing device 1200 can support one or more model usage methods, including but not limited to model training, model invocation, model fine-tuning, model deployment, model inference, and application. In some product forms, the computing device 1200 also supports model management, including but not limited to multi-type model management (supporting the management of discriminative, generative, and other types of models), model version control (supporting the control of different model versions), and model evaluation (evaluating model performance and effectiveness based on model evaluation tools). In other product forms, the computing device 1200 can also create applications based on models, providing API (Application Programming Interface) calling capabilities. Models can be called into the created applications through the API interface, and application management tools are provided to manage and monitor the applications.

[0309] Furthermore, the computing device 1200 may also include data management (supporting the creation and management of model tuning datasets), a training center (providing abundant training resources to help users learn artificial intelligence technologies), and basic control capabilities (providing enterprise-level basic control capabilities to ensure the security and efficient operation of the system). Through the above functions, it provides a comprehensive and integrated device for artificial intelligence development, training, deployment, and application.

[0310] Figure 13 shows a structural block diagram of an electronic device provided in one embodiment of the present disclosure.

[0311] The memory 1310 and the processor 1320 are connected via a bus 1330;

[0312] The memory 1310 is configured to store computer programs / instructions, and the processor 1320 is configured to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor 1320, they implement the steps of the above-described visual generation task processing method, visual generation model training method, or information processing method based on visual generation model.

[0313] Specifically, the components of the electronic device 1300 include, but are not limited to, a memory 1310 and a processor 1320. The processor 1320 and the memory 1310 can be connected via a bus 1330.

[0314] Electronic device 1300 may also include access device 1340, which enables electronic device 1300 to communicate with database 1350 storing data via one or more networks 1360. Examples of these networks 1360 include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. Access device 1340 may include one or more of any type of wired or wireless network interface (e.g., network interface controller (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, Wi-MAX (Worldwide Interoperability for Microwave Access) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, or Near Field Communication (NFC).

[0315] In one embodiment of this disclosure, the aforementioned components of the electronic device 1300, as well as other components not shown in FIG13, may also be connected to each other, for example, via bus 1330. It should be understood that the electronic device structural block diagram shown in FIG13 is merely for illustrative purposes and is not intended to limit the scope of this disclosure. Those skilled in the art can add or replace other components as needed.

[0316] Electronic device 1300 can be any type of stationary or mobile electronic device, including mobile computers or mobile electronic devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable electronic devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary electronic devices such as desktop computers or personal computers (PCs). Electronic device 1300 can also be a mobile or stationary server.

[0317] The above is an illustrative scheme of an electronic device according to this embodiment. It should be noted that the technical solution of this electronic device belongs to the same concept as the above-described visual generation task processing method, visual generation model training method, or information processing method based on visual generation model. For details not described in detail in the technical solution of the electronic device, please refer to the description of the above-described visual generation task processing method, visual generation model training method, or information processing method based on visual generation model.

[0318] An embodiment of this disclosure also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described visual generation task processing method, visual generation model training method, or information processing method based on a visual generation model.

[0319] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solutions of the above-described visual generation task processing method, visual generation model training method, or information processing method based on visual generation model. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solutions of the above-described visual generation task processing method, visual generation model training method, or information processing method based on visual generation model.

[0320] An embodiment of this disclosure also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the above-described visual generation task processing method, visual generation model training method, or information processing method based on a visual generation model.

[0321] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product belongs to the same concept as the above-mentioned visual generation task processing method, visual generation model training method, or information processing method based on visual generation model. For details not described in detail in the technical solution of the computer program product, please refer to the description of the above-mentioned technical solution of the visual generation task processing method, visual generation model training method, or information processing method based on visual generation model.

[0322] The foregoing has described specific embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0323] Computer programs / instructions include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. Computer-readable media can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in computer-readable media can be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0324] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of this disclosure are not limited to the described order of actions, because according to the embodiments of this disclosure, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this disclosure.

[0325] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0326] The preferred embodiments disclosed above are merely illustrative of this disclosure. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to specific implementations. Clearly, many modifications and variations can be made based on the embodiments of this disclosure. These embodiments are selected and specifically described in this disclosure to better explain the principles and practical applications of the embodiments of this disclosure, thereby enabling those skilled in the art to better understand and utilize this disclosure. This disclosure is limited only by the claims and their full scope and equivalents.

Claims

1. A method for processing visual generation tasks, comprising: Acquire task data for a target visual generation task, the task data including target image data pairs and a target style description of the style to be generated, the target image data pairs being obtained based on a target object graph containing the target object; The target image data pair is superimposed with random noise to obtain target image noise. Based on the target style description and the target image noise, the target visual generation model corresponding to the target visual generation task generates a target style image of the target object under the style to be generated. The target visual generation model is obtained by training at least one sample image data pair with a set sample style and the corresponding sample style description.

2. The method of claim 1, wherein, The task data for acquiring the target visual generation task includes: The target image data pairs uploaded from the front end and the target style description of the style to be generated are used as the task data for the target visual generation task; or... The process involves acquiring the image of the object to be processed uploaded from the front end, along with a target style description for the style to be generated; segmenting the image of the object to be processed using an image segmentation model to obtain a target object image with a solid color background; copying the target object image with a solid color background and stitching the copied target object image together in a set direction to obtain the target image data pair; and using the target image data pair and the target style description as task data for the target visual generation task.

3. The method of claim 1, wherein, The step of superimposing the target image data pair with random noise to obtain target image noise, and generating a target style image of the target object under the style to be generated based on the target style description and the target image noise, using the target visual generation model corresponding to the target visual generation task, includes: The updated image data pair of the nth round is superimposed with random noise to obtain the updated image noise of the nth round, wherein n is initially set to 1, the updated image data pair is initially the target image data pair, and the updated image noise is initially the target image noise; The target style description and the updated image noise are input into the target visual generation model. The target visual generation model denoises the updated image noise according to the target style description to restore the corresponding candidate image data pair. Increment n by 1, use the candidate image data pair as the updated image data pair for the nth round, return to the step of superimposing the updated image data pair for the nth round with random noise to obtain the updated image noise for the nth round, until the denoising stop condition is met; The image at the first position in the candidate image data pair is segmented as the target style image, wherein the first position is the position of the sample style map of the target visual generation model in the sample image data pair during the training phase.

4. The method according to claim 3, wherein, The step of superimposing the updated image data pair of the nth round with random noise to obtain the updated image noise of the nth round includes: Based on the set weights of the denoising process in the nth round, the updated image data pair in the nth round is superimposed with random noise to obtain the updated image noise in the nth round. In the nth round of denoising, the weight of random noise in the set weight is greater than that in the (n+1)th round of denoising.

5. The method according to claim 3, wherein, The step of using the candidate image data pair as the updated image data pair in the nth round includes: The image features at the second position in the candidate image data pair are updated with the updated image noise, while the image features at the first position in the candidate image data pair remain unchanged, to obtain the updated image data pair for the nth round, wherein the second position is the position of the sample object image in the sample image data pair during the training phase of the target visual generation model.

6. The method according to claim 1, wherein, The step of generating a target style image of the target object under the style to be generated, based on the target style description and the target image noise, using the target visual generation model corresponding to the target visual generation task, includes: The target style description and the target image noise are input into the target visual generation model. Under the distributed inference framework, the target visual generation model generates a target style image of the target object under the style to be generated. The inference process of the target visual generation model to generate the target style image is based on at least one parallel strategy under the distributed inference framework. The parallel strategy refers to the inference strategy in which the inference task is decomposed into multiple sub-tasks and allocated to different computing units for simultaneous execution through task decomposition and distributed computing resource coordination during the model inference process.

7. The method according to any one of claims 1-6, wherein, After generating the target style image of the target object under the style to be generated based on the target style description and the target image noise, using the target visual generation model corresponding to the target visual generation task, the process further includes: The target style image is fed back to the front end; Receive task feedback information sent by the front end, wherein the task feedback information is information fed back by the front end for the target style image; Based on the task feedback information, optimization sample data is constructed; The target visual generation model is tuned and trained based on the tuned sample data.

8. A method for training a visual generative model, comprising: Obtain at least one image training sample with a defined sample style, wherein the image training sample includes sample image data pairs and corresponding sample style descriptions, the sample image data pairs are obtained based on sample object maps and sample style maps, and the sample style map is an image of a sample object in the sample object map under the defined sample style; Noise is added to the sample image data pairs in each training sample to obtain the corresponding sample image noise; Based on the sample style descriptions and corresponding sample image noise in each image training sample, the visual generation model to be trained is trained to obtain the target visual generation model after training.

9. The method according to claim 8, wherein, The step of training the visual generation model to be trained based on the sample style descriptions and corresponding sample image noise in each image training sample to obtain the trained target visual generation model includes: The first sample style description and the first sample image noise are input into the visual generation model to be trained. The visual generation model to be trained performs denoising processing on the first sample image noise according to the first sample style description to restore the corresponding predicted image data pair. The first sample style description and the first sample image noise are the sample style description and sample image noise included in the first image training sample. The first image training sample is any one of the image training samples. Based on the predicted image data pair and the first sample image data pair in the first image training sample, the prediction loss of the visual generation model to be trained is determined. The model parameters of the visual generation model to be trained are adjusted based on the prediction loss. The sample style description of the next image training sample of the first image training sample and the corresponding sample image noise are then input into the adjusted visual generation model to be trained. The visual generation model to be trained is trained until the training stopping condition is met, and the target visual generation model is obtained after training.

10. The method according to claim 8 or 9, wherein, The step of adding noise to the sample image data pairs in each image training sample to obtain the corresponding sample image noise includes: The sample image data pair is denoised based on at least one noise dimension to obtain sample image noise of at least one size, wherein different noise dimensions correspond to sample image noise of different sizes, and sample image noise of different sizes is used to predict prediction image data pairs of different resolutions for the same sample image data pair.

11. The method according to claim 8, wherein, The acquisition of at least one image training sample with a defined sample style includes: Obtain at least one initial sample style map with a defined sample style; Based on the candidate initial sample style map, a corresponding second sample image data pair is constructed, wherein the candidate initial sample style map is the initial sample style map corresponding to the candidate sample style, and the candidate sample style is any one of the at least one set sample style; Image analysis of the second sample image data pair is performed using a first visual language model to obtain a second sample style description of the second sample image data pair. Based on the second sample image data pair and the second sample style description, construct the second image training sample corresponding to the candidate sample style.

12. The method according to claim 11, wherein, The process of obtaining an initial sample style map with at least one defined sample style includes: Receive at least one styled sample uploaded from the front end; Based on the at least one set sample style, search for at least one corresponding initial sample style map.

13. The method according to claim 11, wherein, The construction of the corresponding second sample image data pair based on the candidate initial sample style map includes: The candidate initial sample style map is analyzed by a second visual language model to obtain the initial description text of the candidate initial sample style map, wherein the second visual language model is the same as or different from the first visual language model; The initial description text is expanded using a language model to obtain the expanded description text; An updated sample style map corresponding to the expanded descriptive text is generated by a candidate visual generation model, wherein the candidate visual generation model is the same as or different from the visual generation model to be trained. The sample objects in the updated sample style map are concatenated with the updated sample style map to obtain the second sample image data pair.

14. The method according to claim 13, wherein, After generating the updated sample style map corresponding to the expanded descriptive text through the candidate visual generation model, the method further includes: The updated sample style map is segmented into its main body using an image segmentation model to obtain sample objects with solid color backgrounds. The step of concatenating the sample objects in the updated sample style map with the updated sample style map to obtain the second sample image data pair includes: The sample object with the solid color background is stitched together with the updated sample style map in a set direction to obtain the second sample image data pair.

15. An information processing method based on a visual generative model, applied to a task platform, comprising: The device receives a model request sent by a terminal device, wherein the model request includes at least one of the following: a scene identifier of the target scene, scene input data of the target scene, and model specification parameters. Based on the model request, a corresponding target visual generation model is determined from at least one visual generation model, wherein the at least one visual generation model is trained based on the visual generation model training method as described in any one of claims 8-14; Deploy the target visual generation model, and based on the target visual generation model, construct a visual generation interface so that the terminal device can schedule the target visual generation model to execute the target visual generation task.

16. A task platform, comprising a request interface and a response unit; The request interface is configured to receive model requests sent by the terminal device, wherein... The model request includes at least one of the following: the scene identifier of the target scene, the scene input data of the target scene, and the model specification parameters. The response unit is configured to determine a corresponding target visual generation model from at least one visual generation model based on the model request, wherein the at least one visual generation model is trained based on the visual generation model training method as described in any one of claims 8-14.

17. The task platform according to claim 16, wherein, The model request includes a scene identifier for the target scene; the response unit is further configured to: Based on the scene identifier of the target scene, a target visual generation model adapted to the target scene is searched from the model library, wherein the model library stores at least one visual generation model adapted to different visual generation task scenarios. The model request includes scene input data for the target scene; the response unit is further configured to: From at least one visual generation model, determine an initial visual generation model adapted to the target scene; Based on the scene input data of the target scene, the initial visual generation model is trained to obtain the target visual generation model; The model request includes model specification parameters; the response unit is further configured as follows: Based on the model specification parameters, the corresponding target visual generation model is searched from the model library, wherein the model library stores multiple visual generation models with different model specification parameters.

18. The task platform according to claim 16, wherein, The task platform also includes a visual generation interface, which is constructed based on the target visual generation model. The visual generation interface is configured to allow the terminal device to schedule and execute the target visual generation task.

19. A computing device, comprising: Memory and processor; The memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1-15.

20. An electronic device, comprising: A memory and a processor, the memory and the processor being connected via a bus; The memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1-15.

21. A computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1-15.