Data processing method, device, computer equipment and storage medium

Through text data processing and text-to-image conversion technology, character images are directly generated, which solves the problem of users drawing character images in the existing technology, improves the efficiency of character generation and lowers the threshold.

CN118803372BActive Publication Date: 2025-09-16TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310402350.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-11
Publication Date
2025-09-16
Estimated Expiration
2043-04-11

AI Technical Summary

Technical Problem

When creating business characters, existing animation video production software requires users to draw character pictures, which makes them dependent on painting and design skills, increases the difficulty and time of image generation, raises the threshold for character generation, and reduces efficiency.

Method used

By obtaining text data to describe the business role style, and using data processing methods and text-to-image conversion technology, role images are generated. This includes text data processing, text-to-image conversion modules, and intelligent generation services, which can directly generate role images without the need for users to draw them.

Benefits of technology

It reduces the difficulty of character image generation, reduces the drawing time and effort required to create characters, improves character generation efficiency, and lowers the threshold.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118803372B_ABST
    Figure CN118803372B_ABST
Patent Text Reader

Abstract

The present application embodiment discloses a data processing method, apparatus, computer equipment, and storage medium, which can be applied to artificial intelligence scenarios, including: obtaining object input data including text data; the text data includes a business role style for describing a business role to be generated, and the data type of the business role style is a text type; the style category in the business role style includes one or more style categories in a style category set; obtaining a data prompt template including an initial role style, adding the business role style to the data prompt template, and determining the added data prompt template as the text to be converted; the data type of the data prompt template is a text type; based on the object input data, performing text-to-image conversion on the text to be converted to obtain a role image corresponding to the object input data; the role image is used to generate the business role to be generated. Using the present application embodiment, the difficulty of image generation can be reduced and the efficiency of role generation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data processing method, apparatus, computer equipment, and storage medium. Background Art

[0002] Existing animation video production software can create business characters (for example, animation characters) by uploading pictures. However, the pictures uploaded here are generated based on characters drawn in advance based on business objects (for example, creating users). That is, this character generation method not only relies on the user's drawing and design skills, but also takes a lot of drawing time, which increases the difficulty of picture generation and greatly increases the threshold for creating animation characters, which greatly affects the efficiency of character generation. Summary of the Invention

[0003] The embodiments of the present application provide a data processing method, apparatus, computer device, and storage medium, which can reduce the difficulty of image generation and improve the efficiency of character generation.

[0004] An embodiment of the present application provides a data processing method, including:

[0005] Obtaining object input data including text data; the text data including a business role style for describing a business role to be generated, and the data type of the business role style is a text type; the style category in the business role style includes one or more style categories in a style category set; the style category set includes a virtual clothing accessory style class and a virtual role attribute style class;

[0006] Obtain a data prompt template including an initial role style, add a business role style to the data prompt template, and determine the added data prompt template as text to be converted; the data type of the data prompt template is a text type; the initial role style is the default role style configured by the pointer for the business role to be generated;

[0007] Based on the object input data, the text to be converted is converted into an image to obtain a role image corresponding to the object input data; the role image is used to generate the business role to be generated.

[0008] An embodiment of the present application provides a data processing method, including:

[0009] Display the character creation interface; the character creation interface includes a text input box and an image generation control;

[0010] In response to an input operation on the text input box, text data describing a business role style of the business role to be generated is displayed; the data type of the business role style is a text type; the style category in the business role style includes one or more style categories in a style category set; the style category set includes a virtual clothing accessory style class and a virtual role attribute style class;

[0011] When object input data including text data is obtained, responding to a trigger operation on the image generation control, switching from the character creation interface to the character generation interface;

[0012] The role image is displayed in the role generation interface; the role image is used to generate the business role to be generated; the role image is obtained after the text to be converted is converted into an image; the text to be converted is determined by the business role style and the initial role style in the data prompt template; the data type of the data prompt template is a text type; the initial role style is the default role style configured by the pointer for the business role to be generated.

[0013] An embodiment of the present application provides a data processing device, including:

[0014] A data acquisition module is configured to acquire object input data including text data; the text data includes a business role style for describing a business role to be generated, and the data type of the business role style is a text type; the style category in the business role style includes one or more style categories in a style category set; the style category set includes a virtual clothing accessory style class and a virtual character attribute style class;

[0015] The style adding module is used to obtain a data prompt template including an initial role style, add a business role style to the data prompt template, and determine the added data prompt template as the text to be converted; the data type of the data prompt template is a text type; the initial role style is the default role style configured by the pointer for the business role to be generated;

[0016] The text-image conversion module is used to perform text-image conversion on the text to be converted based on the object input data to obtain the role image corresponding to the object input data; the role image is used to generate the business role to be generated.

[0017] Among them, the style adding module includes:

[0018] The template acquisition unit is used to acquire a data prompt template including an initial character style; the data prompt template includes a first position corresponding to a virtual character attribute style class, a second position corresponding to the initial character style, and a third position corresponding to a virtual clothing and accessories style class;

[0019] a to-be-recognized text determination unit, configured to determine the to-be-recognized text based on the text format and text data of the data prompt template;

[0020] The splitting processing unit is used to split the text to be recognized based on the splitting mark in the text to be recognized to obtain N subtexts; the N subtexts include subtext X i ; N is a positive integer; i is a positive integer less than or equal to N; the business role style in the text data includes role styles corresponding to N sub-texts respectively;

[0021] Position determination unit for identifying subtext X i Corresponding character style, based on subtext X i The corresponding character style determines the subtext X from the first and third positions i Corresponding location information Y i ;

[0022] Style add unit for sub text X in data tip template i The corresponding character style is added to the position information Y i , until N role styles are added to the data prompt template respectively, and the added data prompt template is determined as the text to be converted.

[0023] The to-be-recognized text determination unit includes:

[0024] a format determination subunit, configured to use the text format of the data prompt template as the first text format and the text format of the text data as the second text format;

[0025] a format comparison subunit, configured to perform a format comparison on the first text format and the second text format to obtain a comparison result;

[0026] a first determining subunit, configured to determine the text data as text to be recognized if the comparison result indicates that the first text format is consistent with the second text format;

[0027] The second determining subunit is used to call the translation service if the comparison result indicates that the first text format is inconsistent with the second text format, and based on the translation service, translate the text format of the text data from the second text format to the first text format, and determine the translated text data as the text to be recognized.

[0028] The position determination unit includes:

[0029] Category identification subunit, used to identify subtext X i The style category of the corresponding character style;

[0030] The third determining subunit is used to determine if the subtext X iIf the style category of the corresponding character style belongs to the virtual character attribute style category, the first position is determined as the subtext X i Corresponding location information Y i ;

[0031] The fourth determining subunit is used to determine if the subtext X i If the style category of the corresponding character style belongs to the virtual clothing accessories style category, the third position is determined as the subtext X i Corresponding location information Y i .

[0032] The text-to-image conversion module includes:

[0033] A model acquisition unit is configured to call an intelligent generation service to acquire a first text-image generation model; the first text-image generation model is obtained by training a second text-image generation model based on sample text data and sample character images; the sample text data is determined based on a data prompt template;

[0034] a first data input unit, configured to input the text to be converted into the first text-to-image generation model if the object input data includes text data;

[0035] The first text-to-image conversion unit is configured to perform text-to-image conversion on the text to be converted by using a first text-to-image generation model to obtain a character image corresponding to the object input data.

[0036] The first text-image generation model includes an encoder and an image generator; the image generator includes an image information generator and an image decoder;

[0037] The first text-to-image conversion unit includes:

[0038] The encoding processing subunit is used to input the text to be converted into the encoder, and the encoder performs encoding processing on the text to be converted to obtain a text semantic vector corresponding to the text to be converted;

[0039] The information extraction subunit is used to input the text semantic vector into the image information generator, and extract information from the text semantic vector through the image information generator to obtain the image information latent vector; the image information latent vector is used to reflect the image information of the text to be converted;

[0040] The text-to-image conversion subunit is used to input the latent vector of the image information into the image decoder, and perform text-to-image conversion on the latent vector of the image information through the image decoder to obtain the initial image corresponding to the object input data;

[0041] The role picture determination subunit is used to determine the role picture corresponding to the object input data from the initial picture.

[0042] The text-to-image conversion module further includes:

[0043] a second data input unit configured to input the text to be converted and the reference image data into the first text-image generation model if the object input data includes text data and reference image data; the reference image data refers to the image data input by the business object in the role creation interface of the client;

[0044] The second text-to-image conversion unit is configured to perform text-to-image conversion on the text to be converted using the first text-to-image generation model to obtain images of key parts corresponding to the text to be converted;

[0045] The image replacement unit is used to perform image recognition on the reference image data to obtain the image to be replaced that has a part matching relationship with the key part image. In the reference image data, the image to be replaced is replaced with the key part image, and the replaced reference image data is determined as the character image corresponding to the object input data.

[0046] The device further comprises:

[0047] A sample acquisition module is used to acquire sample text data determined based on a data prompt template and a sample character image corresponding to the sample text data;

[0048] A sample input module is used to input sample text data into the second text-image generation model, perform text-image conversion on the sample text data through the second text-image generation model, and obtain a predicted character image corresponding to the sample text data;

[0049] The model training module is used to train the second text-image generation model based on the sample character images and the predicted character images to obtain the first text-image generation model.

[0050] The object input data is obtained from the image generation request sent by the client;

[0051] The device also includes:

[0052] The picture storage module is used to call the picture storage service after generating the character picture corresponding to the object input data, store the character picture through the picture storage service, and generate a picture storage link associated with the character picture;

[0053] The link sending module is used to send the picture storage link to the client, so that the client displays the character picture when responding to the trigger operation for the picture storage link.

[0054] An embodiment of the present application provides a data processing device, including:

[0055] Create an interface display module, which is used to display the character creation interface; the character creation interface includes a text input box and an image generation control;

[0056] A text data display module is configured to respond to input operations on a text input box and display text data describing a business role style of a business role to be generated; the data type of the business role style is a text type; the style category in the business role style includes one or more style categories in a style category set; the style category set includes a virtual clothing and accessories style class and a virtual role attribute style class;

[0057] A generation interface display module is used to switch from the character creation interface to the character generation interface in response to a trigger operation on the image generation control when object input data including text data is obtained;

[0058] The role picture display module is used to display the role picture in the role generation interface; the role picture is used to generate the business role to be generated; the role picture is obtained after the text to be converted is converted into a picture; the text to be converted is determined by the business role style and the initial role style in the data prompt template; the data type of the data prompt template is a text type; the initial role style is the default role style configured by the pointer to the business role to be generated.

[0059] Among them, the creation interface display module includes:

[0060] A release interface display unit, used to display a role release interface; the role release interface includes a role creation control;

[0061] The sub-interface display unit is configured to respond to a trigger operation on a character creation control and display a creation method selection sub-interface on the character publishing interface; the creation method selection sub-interface is an interface superimposed on the character publishing interface, and the interface size of the creation method selection sub-interface is smaller than the interface size of the character publishing interface; the creation method selection sub-interface includes M creation method selection controls; M is a positive integer; the M creation method selection controls include an intelligent creation method selection control;

[0062] The creation interface display unit is used to respond to the trigger operation of the intelligent creation method selection control and display the character creation interface.

[0063] Among them, the role creation interface includes a first role viewing control and a second role viewing control; the first role viewing control is used to view the historical roles created by the business object and the historical text data corresponding to the historical roles; the second role viewing control is used to view the reference roles and the reference text data corresponding to the reference roles; the reference roles are H selected roles selected from the role library; H is a positive integer.

[0064] Among them, the character creation interface includes text reference controls;

[0065] The text data display module includes:

[0066] A prompt information display unit is used to display prompt information associated with the text reference control on the character creation interface if the text input box does not display text data, in response to a trigger operation on the image generation control; the prompt information is used to instruct the business object to enter text data in the text input box based on the text reference control;

[0067] The text data display unit is used to respond to the trigger operation of the text reference control when the prompt information is closed, and display text data describing the business role style of the business role to be generated in the text input box; the text data is text data extracted from the high-quality vocabulary.

[0068] Among them, the character creation interface includes reference image drawing controls;

[0069] The device also includes:

[0070] The drawing interface display module is used to respond to the trigger operation of the reference image drawing control and switch from the character creation interface to the image drawing interface; the image drawing interface includes a drawing area and an image storage control;

[0071] A reference picture drawing module, configured to respond to a drawing operation in the drawing area and display reference picture data corresponding to the drawing operation in the drawing area;

[0072] The reference picture display module is used to respond to the trigger operation on the picture storage control, switch from the picture drawing interface to the character creation interface, and display the reference picture data in the character creation interface;

[0073] The input data determination module is used to determine the reference image data and text data as object input data.

[0074] The character generation interface is displayed in response to a trigger operation on the first image generation control; the first image generation control is an image generation control in the character creation interface;

[0075] The character picture display module includes:

[0076] A first control display unit, configured to display a first display control indicating the progress of image generation in the character generation interface;

[0077] An initial image display unit, configured to display Z initial images corresponding to the object input data in the character generation interface when the generation progress of the first display control reaches a progress threshold; Z is a positive integer; each initial image is an image with a first quality coefficient generated based on the object input data;

[0078] A second control display unit is configured to display a second picture generation control when a picture to be processed is determined from the Z initial pictures;

[0079] A third control display unit is configured to respond to a triggering operation on the second image generation control and display a second display control for indicating the image generation progress in the character generation interface;

[0080] The character picture display unit is used to display the picture to be processed with a second quality coefficient in the character generation interface when the generation progress of the second display control reaches the progress threshold, and use the displayed picture to be processed as the character picture; the second quality coefficient is higher than the first quality coefficient.

[0081] On one hand, the present application provides a computer device, including: a processor, a memory, and a network interface;

[0082] The processor is connected to the memory and the network interface, wherein the network interface is used to provide data communication functions, the memory is used to store computer programs, and the processor is used to call the computer program so that the computer device executes the method provided in the embodiment of the present application.

[0083] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program is suitable for being loaded and executed by a processor so that a computer device having the processor executes the method provided by the embodiment of the present application.

[0084] On the one hand, an embodiment of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium; a processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the method in the embodiment of the present application.

[0085] In an embodiment of the present application, a computer device with text-to-image conversion functionality, when creating a business character, does not rely on prior business object drawing and design expertise. Instead, it can directly obtain object input data comprising text data. The text data may include a business character style describing the business character to be generated, and the data type of the business character style is text. The style categories in the business character style may include one or more style categories from a set of style categories, wherein the set of style categories may include virtual clothing and accessories style classes and virtual character attribute style classes. Furthermore, the computer device may obtain a data prompt template comprising an initial character style, add the business character style to the data prompt template, and determine the added data prompt template as the text to be converted. The data type of the data prompt template is text; the initial character style refers to the default character style configured for the business character to be generated. The computer device may then perform text-to-image conversion on the text to be converted based on the object input data, obtaining a character image corresponding to the object input data; the character image is used to generate the business character to be generated. It can be seen that the role creation method provided in the embodiment of the present application does not require the business object to draw and upload a role picture, but requires the business object to determine text data that can describe the business role style of the business role to be generated. Compared with drawing a role picture, it does not need to spend a lot of drawing time and effort to clearly describe the role information of the business role. Therefore, when the computer device obtains the object input data including text data, it can directly obtain the text to be converted based on the data prompt template and text data, so as to automatically generate the role picture of the object input data later. This means that the role creation method provided in the embodiment of the present application can reduce the difficulty of image generation, greatly lower the threshold for creating roles, and thus improve the efficiency of role generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0087] Figure 1 This is a schematic diagram of a network architecture provided by an embodiment of the present application;

[0088] Figure 2 This is a schematic diagram of a scenario for generating a character image provided by an embodiment of the present application;

[0089] Figure 3 This is a flow chart of a data processing method provided in an embodiment of the present application;

[0090] Figure 4 This is a schematic diagram of an interface switching for displaying a character creation interface provided by an embodiment of the present application;

[0091] Figure 5 This is a schematic diagram of a scenario in which a business object inputs text data based on a text reference control, provided by an embodiment of the present application;

[0092] Figure 6 This is a schematic diagram of an interface switching for displaying a character image corresponding to first object input data provided by an embodiment of the present application;

[0093] Figure 7 This is a flow chart of a data processing method provided in an embodiment of the present application;

[0094] Figure 8 This is a schematic diagram of an interface switching for displaying a character image corresponding to second object input data provided by an embodiment of the present application;

[0095] Figure 9 This is a schematic diagram of a scenario for viewing business roles provided by an embodiment of the present application;

[0096] Figure 10 This is a flow chart of a data processing method provided in an embodiment of the present application;

[0097] Figure 11 is a structural diagram of a data processing device provided in an embodiment of the present application;

[0098] Figure 12 is a structural diagram of a data processing device provided in an embodiment of the present application;

[0099] Figure 13 is a structural diagram of a data processing device provided in an embodiment of the present application;

[0100] Figure 14 is a structural diagram of a data processing device provided in an embodiment of the present application;

[0101] Figure 15 This is a schematic diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0102] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0103] It should be understood that the embodiment of the present application provides a method for automatically generating character images through text data to create business roles, which can be applied to the field of artificial intelligence. Among them, the so-called artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or digital computer-controlled calculations to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0104] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and smart transportation.

[0105] Natural language processing (NLP) is a key area of ​​research in computer science and artificial intelligence. It studies the theories and methods that enable effective communication between humans and computers using natural language. Natural language processing (NLP) integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language we use in everyday life—and is closely linked to the study of linguistics. Natural language processing technologies typically include text processing, semantic understanding, machine translation, robotic question answering, and knowledge graphs.

[0106] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0107] See Figure 1 , Figure 1 This is a schematic diagram of a network architecture provided by an embodiment of the present application. Figure 1 As shown, the network architecture may include a server 10F and a terminal device cluster. The terminal device cluster may include one or more terminal devices. Figure 1 As shown, the terminal device cluster may specifically include terminal device 100a, terminal device 100b, terminal device 100c, ..., terminal device 100n. Figure 1 As shown, terminal device 100a, terminal device 100b, terminal device 100c, ..., terminal device 100n can each establish a network connection with the server 10F, so that each terminal device can exchange data with the server 10F through the network connection. The network connection here does not limit the connection method and can be directly or indirectly connected through wired communication, directly or indirectly connected through wireless communication, or through other methods, which are not limited in this application.

[0108] Each terminal device in the terminal device cluster may include: smart phones, tablet computers, laptop computers, desktop computers, smart speakers, smart watches, car terminals, smart TVs and other smart terminals with data processing functions. Figure 1 Each terminal device in the terminal device cluster shown can be installed with a client (for example, animation video production software). When the client runs in each terminal device, it can communicate with the above-mentioned Figure 1 The server 10F shown in FIG. 10F performs data interaction. The client may be an independent client or an embedded sub-client integrated in a client (eg, a social client, an educational client, a multimedia client, etc.), which is not limited here.

[0109] like Figure 1 As shown, the server 10F in the embodiment of the present application can be the server corresponding to the client. The server 10F can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The embodiment of the present application does not limit the number of terminal devices and servers.

[0110] For ease of understanding, the embodiments of the present application can be Figure 1 Select one terminal device from the multiple terminal devices shown as the terminal device used by the business object (ie, the business terminal device). Figure 1The terminal device 100a shown is used as a business terminal device, and a client can be integrated into the business terminal device. At this time, the business terminal device can realize data interaction with the server 10F through the business data platform corresponding to the client. Among them, the client here can run a trained text-image generation model (i.e., the first text-image generation model). For example, the text-image generation model (e.g., the Stable Diffusion model) can be a deep learning text-to-image generation model, which is mainly used to generate detailed images based on the description of the text. For example, the model name of the text-image generation model here can be "SV_pose_ <svp>_token_training.ckpt".

[0111] Among them, the embodiment of the present application may refer to the business data that the business object can input in the client as object input data, and the object input data may include data of a data type of text data. Optionally, the object input data may also include data of two data types of text data and image data (i.e., reference image data), which will not be limited here. Among them, the reference image data in the object input data can be graffiti drawn by the business object itself, or it can be a picture uploaded by the business object. This means that the embodiment of the present application can subsequently combine the reference image data and text data, and through the first text-image generation model, quickly and accurately generate the character picture of the object input data, thereby meeting the personalized needs of the business object and improving the accuracy and fun of the character picture generation.

[0112] The style category set in the embodiment of the present application may include multiple style categories, and each style category may accurately represent the role image of the required business role by adding adjectives such as color or personality. For example, the style category set may include a virtual clothing accessories style class and a virtual character attribute style class. Among them, the virtual clothing accessories style class may include a virtual clothing styling class and a virtual clothing accessories class. For example, the text description corresponding to the virtual clothing styling class here may include a red long skirt, a white shirt, shorts, black leather shoes, etc., and the text description corresponding to the virtual clothing accessories class here may include a sun hat, red glasses, toys (for example, a yo-yo), jewelry, etc.

[0113] The virtual character attribute style class here may include a first attribute style class (e.g., a body attribute style class), a second attribute style class (e.g., a biological attribute style class), and a third attribute style class (e.g., an other attribute style class). For example, the text description corresponding to the body attribute style class here may include tall, short, fat, thin, etc. The text description corresponding to the biological attribute style class here may include virtual characters (e.g., beautiful women, handsome men, children, babies, old people, etc.), virtual animals (e.g., cute puppies, frogs, etc.), virtual plants (e.g., flowers, grass, etc.), and virtual objects (e.g., cars, benches, etc.). The text description corresponding to the other attribute style class here may include green skin, red hair, etc.

[0114] In the embodiment of the present application, the computer device with the text-to-image conversion function can be a server or a Figure 1 Any terminal device in the terminal device cluster shown, for example, terminal device 100a, will not be limited to the specific form of the computer device. For ease of understanding, the computer device in the embodiment of the present application can be a server (for example, Figure 1 The server 10F shown in FIG. 1 is taken as an example to illustrate a specific implementation method of the computer device automatically generating a character image based on text data. The server can provide a translation service, an intelligent generation service, and a picture storage service. The translation service is used to translate the text format of the received text data into the text format of the data prompt template; the intelligent generation service is used to perform text-to-image conversion on the text to be converted determined by the text data and the data prompt template based on a trained text-to-image conversion model (i.e., the first text-to-image conversion model); the picture storage service is used to store the character image corresponding to the object input data and generate a picture storage link.

[0115] For example, when creating a business role, the computer device can obtain object input data including text data. The text data here may include a role style (i.e., a business role style) for describing the business role to be generated, and the data type of the business role style belongs to a text type. The style category in the business role style may include one or more style categories in a style category set. Furthermore, the computer device can obtain a data prompt template (prompt template) including an initial role style, and then add a business role style to the data prompt template, determine the added data prompt template as the text to be converted, and then perform text-to-image conversion on the text to be converted based on the object input data to obtain a role picture corresponding to the object input data, where the role picture is used to generate the business role indicated by the object input data.

[0116] The data prompt template is set when training the text-image generation model. The initial character style in the data prompt template is the default character style configured for the business character to be generated (for example, full-body or half-body character styles used to describe the character's body structure). The data type of the data prompt template is text type. The text format of the data prompt template can be any of the other text formats such as Chinese, English, Russian, etc., which will not be limited here. For example, the data prompt template can be " <svp>The English data template "style, XXX, full body, solo, ###" can also be a Chinese data template "style, XXX, full body, solo, ###", which will not be limited here.

[0117] It can be seen that when the computer device obtains object input data including text data, it can directly obtain the text to be converted based on the data prompt template and text data, so as to automatically generate a character picture of the object input data. This means that the character creation method provided in the embodiment of the present application can reduce the difficulty of picture generation, greatly lower the threshold for creating characters, and thus improve the efficiency of character generation.

[0118] For further understanding, please refer to Figure 2 , Figure 2 This is a schematic diagram of a scene for generating a character image provided by an embodiment of the present application. Figure 2 As shown, the computer device with text-to-image conversion function in the embodiment of the present application can be Figure 2 The server 20F shown in FIG. 2 may be the server 20F described above. Figure 1 Server 10F is shown. Figure 2 The terminal device 200a shown in the figure can be a business object terminal used by object A (ie, business object). The terminal device 200a can be the above-mentioned Figure 1 Any terminal device in the terminal device cluster shown is integrated with a client (eg, animation video production software), for example, terminal device 100a. Figure 2 The text-image generation model 20W (i.e., the first text-image generation model) shown can be a trained neural network model for text-image conversion. The text-image generation model 20W is trained based on sample text data and sample character images, where the sample text data is determined based on the data prompt template 2M.

[0119] like Figure 2 As shown, the terminal device 200a can display the character creation interface 230J provided by the client. The character creation interface 230J may include a text input box and a picture generation control (for example, Figure 2 (See control 2K, "Start Generate" control, for example.) When object A wishes to create a business role, it can perform an input operation on the text input box in role creation interface 230J. The input operation here refers to a trigger operation performed on the text input box for entering text data. This trigger operation may include contact operations such as clicking and long pressing, as well as non-contact operations such as voice and gestures, which are not limited herein.

[0120] When the client responds to the input operation on the text input box, the client can display text data (for example, Figure 2 The text data 2T shown here is of the text type. The business role style data type is text, and the style categories within the business role style may include one or more style categories from a style category set. The style category set may include a virtual clothing accessory style class and a virtual character attribute style class. For example, the text data 2T may be a text description such as "shoulder-length hair, big-eyed beauty."

[0121] It is understandable that the text data 2T here can be object A through Figure 2 The virtual keyboard shown may be used for direct input, or it may be randomly selected from a high-quality vocabulary after object A triggers a text reference control (e.g., a "random description" control) included in the text input box. The high-quality vocabulary here may be descriptive text manually selected by the audit subject (e.g., the audit user) corresponding to the client, or it may be descriptive text intelligently selected by the backend server (i.e., server 20F) corresponding to the client, and the above terms are not limited here.

[0122] After the object A input is completed, the client can obtain the object input data including the text data 2T. Then, when the object A performs a trigger operation (e.g., a click operation) on the control 2K in the character creation interface 230J, the client can respond to the trigger operation and switch the terminal interface of the terminal device 200a from the character creation interface 230J to the character generation interface to wait for the character image corresponding to the object input data to be displayed. At the same time, when responding to the trigger operation, the client can also generate a picture generation request (e.g., Figure 2 The picture generation request 2q shown can then be sent to the server 20F.

[0123] When the server 20F receives the image generation request 2q, it can obtain the object input data including the text data 2T. At the same time, the server 20F can also obtain the data prompt template including the initial character style (for example, Figure 2 The data prompt template 2M shown is a text data type, and the initial role style is a default role style configured for the business role to be generated. For example, if the data prompt template 2M here can be a Chinese data template such as "Shaping, XXX, Half-body, Alone, ###", then the text description corresponding to the initial role style can include "Shaping, Half-body, Alone".

[0124] Furthermore, the server 20F can add the business role style included in the text data 2T to the data prompt template, and then determine the added data prompt template as the text to be converted. For example, the text to be converted here can be a text description such as "style, shoulder-length hair, big-eyed beauty head, half-body, single". At this time, the server 20F can input the text to be converted into Figure 2 The text-image generation model 20W shown in FIG. 2 performs text-image conversion on the text to be converted to obtain a character image corresponding to the object input data (for example, Figure 2 The character picture 2P shown in FIG. 2A can then be returned to the terminal device 200a.

[0125] When the terminal device 20a receives the character image 2P, it can generate the character image on the character generation interface (for example, Figure 2 The role generation interface 240J) shown in FIG. 2 shows the role picture 2P, wherein the role picture 2P can be used to create a business role. Figure 2 As shown, the character generation interface 240J here can also display the work creation information corresponding to the character picture 2P, and the work creation information here can include the creation object corresponding to the character picture 2P (for example, object A), the creation timestamp corresponding to the character picture 2P (that is, the timestamp of the response to the trigger operation for the control 2K), the description word corresponding to the character picture 2P (that is, Figure 2 The text data 2T shown).

[0126] It can be seen that the character creation method provided by the embodiment of the present application does not require object A to draw and upload the character picture, but requires object A to Figure 2 In the text input box shown, text data 2T that can describe the business role style of the business role to be generated is determined. Compared with directly drawing a role picture, the role information of the business role can be clearly described without spending a lot of drawing time and effort. In this way, when the server 20F obtains the object input data including text data 2T, it can directly obtain the text to be converted according to the data prompt template 2M and the text data 2T, so as to automatically generate the role picture 2P of the object input data through the text-image generation model 20W. This means that the role creation method provided in the embodiment of the present application can reduce the difficulty of image generation, greatly reduce the threshold for creating roles, and thus improve the efficiency of role generation.

[0127] It should be noted that Figure 2 The interfaces and controls shown are merely some reference forms of expression. In actual business scenarios, developers can make relevant designs based on product requirements. The embodiments of this application do not limit the specific forms of the interfaces and controls involved.

[0128] The specific implementation method of automatically generating a character image by quickly converting the text to be converted determined by the text data and the data prompt template through the computer device through the object input data including text data can be seen below Figure 3-Figure 9 The corresponding embodiment.

[0129] Further, see Figure 3 , Figure 3 This is a flow chart of a data processing method provided by an embodiment of the present application. Figure 3 As shown, the method can be executed by a computer device with a text-to-image conversion function, which can be a terminal device (for example, the above Figure 1 Any terminal device in the terminal device cluster shown, for example, the terminal device 100a having the model application function, can also be a server (for example, the above Figure 1 For ease of understanding, the embodiment of the present application takes the method executed by the server as an example for explanation, and the method may include at least the following steps S101-S103:

[0130] Step S101: Obtain object input data including text data.

[0131] Specifically, when a business object wishes to create a business role, it can perform an input operation on a text input box in a role creation interface provided by a client. In response to this input operation, the client displays text data describing the business role style of the business role to be created in the text input box. The input operation here refers to a trigger operation performed on the text input box for entering text data. This trigger operation can include contact operations such as clicking and long pressing, or non-contact operations such as voice and gestures, which are not limited here. The data type of the business role style is text, and the style category of the business role style can include one or more style categories from a style category set. The style category set can include virtual clothing and accessory style classes and virtual character attribute style classes. Furthermore, upon obtaining object input data including text data, the client can respond to the business object's trigger operation on an image generation control in the role creation interface, generate an image generation request based on the object input data, and then send the image generation request to a server (e.g., the client backend) to enable the server to obtain the object input data including text data carried in the image generation request.

[0132] It should be understood that the business terminal device in the embodiment of the present application (i.e., the terminal device used by the business object) can display the role release interface provided by the client after starting the client. The role release interface may include a role creation control. Further, the client can respond to the triggering operation of the business object for the role creation control and display the creation method selection sub-interface on the role release interface, wherein the creation method selection sub-interface here can be an interface superimposed on the role release interface, and the interface size of the creation method selection sub-interface is smaller than the interface size of the role release interface. In addition, the creation method selection sub-interface can include M creation method selection controls; M is a positive integer; M creation method selection controls include an intelligent creation method selection control. Then, the client can respond to the triggering operation of the business object for the intelligent creation method selection control to display the role creation interface.

[0133] For further understanding, please refer to Figure 4 , Figure 4 This is a schematic diagram of an interface switching for displaying a character creation interface provided by an embodiment of the present application. Figure 4 As shown, the interfaces shown in the embodiments of the present application are provided by a client (for example, animation video production software) integrated in a business terminal device. The business terminal device may be an object terminal used by object A (ie, a business object). The business terminal device may be the above-mentioned Figure 1 Any terminal device in the terminal device cluster shown, for example, terminal device 100a.

[0134] like Figure 4 As shown, the role publishing interface 410J1 can be a terminal interface displayed by the business terminal device after starting the client. The role publishing interface 410J1 can include the business roles created by object A in history, specifically including role 1 and role 2. In addition, the role publishing interface 410J1 can also include a role creation control (for example, Figure 4 Control 4K1 shown, the "Create Character" control).

[0135] When object A wants to create a new business role, it can execute a trigger operation on control 4K1. When the client responds to the trigger operation, a creation mode selection sub-interface (for example, Figure 4 The creation mode selection sub-interface 420J shown in FIG. 4 is a sub-interface for selecting a creation mode. The creation mode selection sub-interface 420J here may be an interface superimposed on the role publishing interface 410J2, and the interface size of the creation mode selection sub-interface 420J is smaller than the interface size of the role publishing interface 410J2. In addition, the creation mode selection sub-interface 420J may include M creation mode selection controls; M is a positive integer; M here may be 2 for example, and may specifically include intelligent creation mode selection controls (for example, Figure 4 The control 4K2 shown is the "AI Create Character" control) and the image upload control (for example, the "Upload from Album" control). The image upload control can be used to directly upload the character image extracted by object A.

[0136] Furthermore, object A can perform a trigger operation on the control 4K2, and when the client responds to the trigger operation, a character creation interface can be displayed. Figure 4 The character creation interface 430J1 and the character creation interface 430J2 are two types of interfaces. The character creation interface 430J2 can be another type of character creation interface expanded on the basis of the character creation interface 430J1. Figure 4 As shown, both the character creation interface 430J1 and the character creation interface 430J2 can display a text input box 4Q1, an image generation control (eg, control 4K4, i.e., a "start generation" control), and a work display area 4Q2.

[0137] in, Figure 4 The text input box 4Q1 shown may include prompting the business object to enter text data, a character limit corresponding to the text input box 4Q1 (e.g., 500), and a text reference control (e.g., control 4K3, i.e., a "Random Description" control). For example, the prompt may be "Please enter a Chinese or English description of the character you wish to generate."

[0138] in, Figure 4 The work display area 4Q2 shown may include a first character viewing control (eg, Figure 4 My Works controls shown) and second role viewing controls (e.g. Figure 4 The "reference case" control shown). The first role viewing control here can be used to view the historical roles created by the business object and the text data corresponding to the historical roles (i.e., historical text data). The second role viewing control here can be used to view the reference roles and the text data corresponding to the reference roles (i.e., reference text data), and the reference roles are H selected roles selected from the role library; H is a positive integer. Among them, the role library here can be used to store selected roles that have passed quality screening. The selected roles here refer to those determined based on the quality parameters of the roles, and the quality parameters here can be determined based on the interactive parameters after a certain role is released (for example, the number of likes, the number of comments, the number of views, etc.). For example, the selected role can be a role with a higher interactive parameter manually screened by the audit object (for example, the audit user) corresponding to the client, or it can be a role with a higher interactive parameter intelligently screened by the server corresponding to the client, and it will not be limited here.

[0139] Unlike the character creation interface 430J1, the character creation interface 430J2 includes not only the business controls and other data displayed on the character creation interface 430J1, but also a reference image determination control. Specifically, the reference image determination control may include: Figure 4 The reference image drawing control (e.g., control 4K5, i.e., the "graffiti reference image" control) and the reference image upload control (e.g., control 4K6, i.e., the "upload reference image" control) are shown. The reference image drawing control can be used to instruct a business object to draw reference image data (e.g., graffiti) in another interface (i.e., the image drawing interface). The reference image upload control can be used to instruct a business object to upload reference image data (e.g., a picture in an album) stored on a business terminal device.

[0140] It is understandable that when the business object is a new user registered in the client, when the character creation interface is displayed from the terminal interface of the business terminal device, the work display area 4Q2 will first stay at the second character viewing control (i.e. Figure 4 The "Reference Case" control shown in the figure) is used to display H selected roles randomly distributed from the role library (for example, Figure 4 2 selected characters shown) for reference by new users to learn how to input appropriate text data, so that the client can quickly obtain valid text data to generate character images that can accurately meet user needs.

[0141] Furthermore, after the client displays the role creation interface, the business object may perform an input operation on a text input box in the role creation interface, so that the client may display text data describing a business role style of the business role to be generated in the text input box.

[0142] For example, after the business object performs a trigger operation on the text input box, the client can display a character input cursor in the text input box to specify the character input position. Then, the business object can directly perform an input operation through the keyboard (for example, a virtual keyboard displayed on the character creation interface or an external physical keyboard), and the client displays the text data entered by the business object in the text input box in response to the input operation. It is understandable that the description words (i.e., text words) that the business object can enter in the text input box can be Chinese or English, and the two description words can be separated by a separator (for example, a space, a comma, a semicolon, etc.), and the total number of characters of the description words in the text input box needs to be less than or equal to the character limit in the text input box (for example, 500).

[0143] For another example, after the business object performs a trigger operation on the text input box, it can directly reference the text control (for example, Figure 4 The control 4K3 shown executes a trigger operation. In response to the trigger operation, the client can obtain text words randomly extracted from the high-quality vocabulary. The extracted text words can then be used as text data describing the business role style of the business role to be generated, and displayed in the text input box. The text words in the high-quality vocabulary here can be descriptive text manually selected by the audit object corresponding to the client, or they can be descriptive text intelligently selected by the server corresponding to the client, and the above are not limited here.

[0144] For another example, if the business object does not perform an input operation on the text input box, but performs a trigger operation on the image generation control in the character creation interface, a pop-up prompt is required, that is, if the text input box does not display text data, and the business object performs a trigger operation on the image generation control in the character creation interface, then the client can display prompt information associated with the text reference control on the character creation interface when responding to the trigger operation. The prompt information here can be used to instruct the business object to enter text data in the text input box based on the text reference control. Furthermore, when closing the prompt information, the client can respond to the business object's trigger operation on the text reference control and display text data describing the business role style of the business role to be generated in the text input box; the text data here refers to text data extracted from a high-quality vocabulary.

[0145] For further understanding, please refer to Figure 5 , Figure 5 This is a schematic diagram of a scenario in which a business object provides a text reference control to input text data. Figure 5 As shown, the interfaces shown in the embodiments of the present application are provided by a client (for example, animation video production software) integrated in a business terminal device. The business terminal device may be an object terminal used by object A (ie, a business object). The business terminal device may be the above-mentioned Figure 1 Any terminal device in the terminal device cluster shown, for example, terminal device 100a.

[0146] like Figure 5 If the text input box of the character creation interface 530J1 does not display text data, and the object A performs a trigger operation on the image generation control (e.g., control 5K1) in the character creation interface, the client can display prompt information associated with the text reference control (e.g., prompt information 5X) on the character creation interface in response to the trigger operation. The prompt information 5X can be displayed on the character creation interface 530J2. For example, the prompt information 5X can be a text message such as "If you don't have any ideas for the time being, you can also click the "Random Description" button. We have prepared some inspiration for you."

[0147] Furthermore, the object A can execute a trigger operation for closing the prompt information (for example, clicking the "OK, I understand" button on the character creation interface 530J2) to enable the client to close the prompt information and display the prompt information. Figure 5 The character creation interface 530J3 is shown. At this point, object A can directly trigger the text reference control (e.g., control 5K2) in the text input box. In response to this trigger, the client can retrieve randomly extracted text words from a high-quality vocabulary. The extracted text words can then be used as text data describing the business role style of the business role to be created, and displayed in the text input box. The text data here can be a text description such as "shoulder-length hair, big-eyed beauty."

[0148] It is understood that each time the control 5K2 is triggered, the text words displayed in the text input box are refreshed and replaced. If the text words displayed after the client responds to the first trigger operation of object A on the control 5K2 does not meet the user's needs of object A, object A can trigger the control 5K2 again until the text words displayed in the text input box meet the user's needs.

[0149] It can be seen that this method of recommending text words for business objects based on a high-quality vocabulary can quickly display text data describing the business role style of the business role to be generated in the text input box, thereby improving the input efficiency of text data for business objects, and further improving the image generation efficiency when generating role images subsequently.

[0150] It is understandable that when text data is displayed in the text input box, the business object directly triggers the image generation control in the character creation interface, which can be understood as the completion of the input of the business object (for example, object A). Then, when the client responds to the trigger operation, the text data can be directly used as the object input data, and then the image generation request can be directly generated based on the object input data, and then the image generation request can be sent to the server (for example, the client backend) to enable the server to obtain the object input data including the text data carried by the image generation request.

[0151] Step S102 : obtaining a data prompt template including an initial role style, adding a business role style to the data prompt template, and determining the added data prompt template as the text to be converted.

[0152] Among them, the data type of the data prompt template is a text type, and the initial role style is the default role style configured by the pointer for the business role to be generated. Specifically, the server can obtain a data prompt template including the initial role style. The data prompt template may include a first position corresponding to the virtual role attribute style class, a second position corresponding to the initial role style, and a third position corresponding to the virtual clothing accessories style class. Furthermore, the server can determine the text to be recognized based on the text format and text data of the data prompt template, and then can split the text to be recognized based on the split mark in the text to be recognized to obtain N subtexts. Among them, the N subtexts here include subtext X i ; N is a positive integer; i is a positive integer less than or equal to N; the business role pattern in the text data includes role patterns corresponding to N subtexts. At this time, the server can identify the subtext X i The corresponding character style can then be based on the subtext X i The corresponding character style determines the subtext X from the first and third positions i Corresponding location information Y i , in the data tip template, change the subtext X i Add to location information Y i , until N sub-texts are added to the data prompt template respectively, and the added data prompt template is determined as the text to be converted.

[0153] Wherein, since the first text-image generation model is obtained after the server trains the second text-image generation model based on sample text data and sample character pictures, and the sample data here is determined based on the data prompt template. For example, when the server trains the second text-image generation model (i.e., the initial model), it can obtain the sample text data determined based on the data prompt template and the sample character picture corresponding to the sample text data, and then input the sample text data into the second text-image generation model, and perform text-image replacement on the sample text data through the second text-image generation model to obtain the predicted character picture corresponding to the sample text data. Furthermore, the server can train the second text-image generation model based on the sample character picture and the predicted character picture to obtain the first text-image generation model. Wherein, the model parameters associated with the second text-image generation model are adjusted according to the model loss value, and the model loss value can be determined based on the similarity between the sample character picture and the predicted character picture.

[0154] Among them, the sample character picture here can be a standing picture of a character designed independently by the development object, such as a simple line drawing style, solid color filling, and a picture of the character slightly facing left. Based on such sample text data and sample character pictures, after training the first text-image generation model (for example, the Stable Diffusion model), the final second text-image generation model can be able to output a character style with unique characteristics, and when the model is subsequently applied, it can quickly output character pictures that meet the user expectations of the business object, thereby improving the generation efficiency of character pictures.

[0155] Therefore, in order to be able to quickly use the first text-image generation model and automatically convert text data into images, the server needs to pre-process the text data based on the text format of the data prompt template when obtaining the text data determined by the business object in the text input box, that is, convert the data format of the text data into the text format of the data prompt template.

[0156] It is understood that the server may use the text format of the data prompt template as the first text format and the text format of the text data as the second text format, and then perform a format comparison on the first text format and the second text format to obtain a comparison result. If the comparison result indicates that the first text format is consistent with the second text format, the server may directly determine the text data as the text to be recognized.

[0157] Optionally, if the comparison result indicates that the first text format is inconsistent with the second text format, the server needs to call a translation service, and then based on the translation service, translate the text format of the text data from the second text format to the first text format, and determine the translated text data as the text to be recognized. For example, if the data prompt template is " <svp>style, XXX, full body, solo, "such an English data template, and the text data is "unicorn girl, colorful, dreamlike" in Chinese text. Then, the server can convert the text format of the text data into English format, and the text to be recognized obtained can be "unicorn, colorful, dreamlike" in English text.

[0158] Further, the server can split the text to be recognized based on the splitting flag in the text to be recognized to obtain N sub-texts. Among them, if the text to be recognized includes delimiter symbols (for example, space, comma, semicolon and other characters), the server can use the delimiter symbol as the splitting flag and directly split the text data. For ease of understanding, the text formats of the text to be recognized and the data prompt template in the embodiments of the present application can both take the Chinese format as an example. For example, if the text to be recognized is a text description such as "gray trousers, black glasses, frog head", the server can split the text to be recognized into three sub-texts, namely sub-text X1 (for example, gray trousers), sub-text X2 (for example, black glasses), and sub-text X3 (for example, frog head) according to the splitting flag (for example, comma).

[0159] Optionally, if the text to be recognized does not include delimiter symbols, the server can use the word nature as the splitting flag to split the text to be recognized. For example, if the text to be recognized is a text description such as "a frog wearing black glasses and gray trousers", the server filters the text words belonging to the stop word list in the text to be recognized based on the stop word list, and then splits the filtered text to be recognized according to the text words of the noun nature to obtain sub-text X1 (for example, wearing black glasses), sub-text X2 (for example, wearing gray trousers), and sub-text X3 (for example, frog).

[0160] After obtaining N sub-texts, the server can write the N sub-texts into the data prompt template respectively based on the role styles corresponding to the N sub-texts, and determine the data prompt template after addition as the text to be converted. Among them, it can be understood that the server can recognize the style category of the role style corresponding to sub-text X i If the style category of the role style corresponding to sub-text X i belongs to the virtual role attribute style category, the server can determine the first position as the position information Y i corresponding to sub-text X i . Optionally, if the style category of the role style corresponding to sub-text X i belongs to the virtual clothing accessory style category, the server can determine the third position as the position information Y i corresponding to sub-text X i .

[0161] For example, if the data prompt template is a Chinese data template such as "style, XXX, full body, separate, ###", then the positions corresponding to the three initial character styles of "style", "full body" and "separate" can be called fixed positions (i.e., the second position), and the position of the first placeholder in the data prompt template (i.e., the first position, for example, where "XXX" is located) can be used to add a character style belonging to the virtual character attribute style class, and the position of the second placeholder in the data prompt template (i.e., the third position, for example, where "###" is located) can be used to add a character style belonging to the virtual clothing accessories style class. For example, if the N sub-texts corresponding to the text to be recognized are 3, specifically including sub-text X1 (for example, wearing black glasses), sub-text X2 (for example, wearing gray pants), and sub-text X3 (for example, frog), and when the server recognizes that the character style of sub-text X1 (for example, black glasses) belongs to the virtual clothing accessory style class, the character style of sub-text X2 (gray pants) belongs to the virtual clothing accessory style class, and the character style of sub-text X3 (for example, frog head) belongs to the virtual character attribute style class, then after the server adds these 3 character styles to the data prompt template, the text to be converted obtained can be a Chinese text such as "style, frog head, full body, separate, black glasses, gray pants".

[0162] For example, if the data prompt template is " <svp>style, XXX, full body, solo, ###", and the sub-text after splitting includes sub-text X1 (for example, tiger head), then the text to be converted obtained by the server can be " <svp>style, tiger head, full body, solo". If the split subtext includes subtext X1 (for example, panda head) and subtext X2 (for example, red dress), the text to be converted obtained by the server can be " <svp>style, panda head, fullbody, solo, red dress". This means that after adding the character styles corresponding to N sub-texts to the data prompt text, if no character style is added to a certain position in the data prompt text (for example, the first position or the third position), the server can delete the placeholder at this position and use the final data prompt template as the text to be converted.

[0163] Step S103 : Based on the object input data, the text to be converted is converted into an image, and a character image corresponding to the object input data is obtained.

[0164] Specifically, the server can call an intelligent generation service (i.e., an AI generation service) to obtain a first text-image generation model, wherein the first text-image generation model here (i.e., the trained text-image generation model) is based on sample text data and sample role pictures, and is obtained after training the second text-image generation model (the text-image generation model before training), and the sample text data here is determined based on a data prompt template. Furthermore, the computer device can convert the text to be converted into a text-image through the first text-image generation model, and then determine the role picture corresponding to the object input data based on the picture obtained after the text-image conversion. The role picture can be used to generate the business role to be generated.

[0165] Among them, if the object input data includes text data, the server can directly input the text to be converted into the first text-to-image generation model, and then perform text-to-image conversion on the text to be converted through the first text-to-image generation model, and then directly determine the character image corresponding to the object input data from the initial image after the text-to-image conversion.

[0166] It is understood that since the first text-to-image generation model here may include an encoder and an image generator, and the image generator includes an image information generator and an image decoder, when the server performs text-to-image conversion on the text to be converted using the first text-to-image generation model, it may first input the text to be converted into the encoder, which then encodes the text to be converted to obtain a text semantic vector corresponding to the text to be converted. The text semantic vector is then input into the image information generator, which then extracts information from the text semantic vector to obtain an image information latent vector. The image information latent vector here can be used to reflect the image information of the text to be converted. Finally, the server may input the image information latent vector into the image decoder, which then performs text-to-image conversion on the image information latent vector to obtain an initial image corresponding to the object input data. The character image corresponding to the object input data can then be determined from the initial image. The number of initial images here may include Z, where Z is a positive integer. Each initial image refers to an image with a first quality coefficient generated based on the object input data, i.e., a low-quality image (e.g., a thumbnail).

[0167] For example, when the number Z of initial images determined by the server is one, the server can directly convert the initial image again to obtain an initial image with a second quality coefficient, and then use the initial image with the second quality coefficient as a character image to be sent to the client. When the client does not receive the character image returned by the server, the client can display a first display control for indicating the progress of image generation in the character generation interface. When the client receives the character image returned by the server, the generation progress of the first display control reaches a progress threshold (for example, 100%). At this time, the client can display the character image in the character generation interface. The second quality coefficient here is higher than the first quality coefficient, that is, the character image is a high-quality image (for example, a high-definition image) compared to the initial image.

[0168] Optionally, if the number Z of initial images determined by the server is more than one, the server can directly send these Z initial images to the client, so that the client displays these Z initial images in the character generation interface, so that the business object selects an initial image that meets the user expectations of the business object from the Z initial images, and then the selected initial image can be determined as the image to be processed. When the client does not receive the Z initial images returned by the server, the client can display a first display control for indicating the progress of image generation in the character generation interface. When the client receives the Z initial images returned by the server, the generation progress of the first display control reaches a progress threshold (for example, 100%). At this time, the client can display these Z character images in the character generation interface. Then, when the client responds to the business object's selection operation to determine the image to be processed from the Z initial images (i.e., when determining the image to be processed from the Z initial images), the client can display a second display control in the character generation interface for indicating the progress of image generation (i.e., a business control for displaying the progress of high-definition image generation). At the same time, the client can send the image to be processed to the server so that the server can convert the image to be processed to obtain an image to be processed with a second quality coefficient, and then use the image to be processed with the second quality coefficient as a character image to be sent to the client. When the client receives the character image, the generation progress of the second display control reaches the progress threshold. At this time, the client can display the character image in the character generation interface.

[0169] For further understanding, please refer to Figure 6 , Figure 6 This is a schematic diagram of an interface switching for displaying a character picture corresponding to first object input data provided by an embodiment of the present application. The first object input data here may include data of a data type such as text data. Figure 6 As shown, the interfaces shown in the embodiments of the present application are provided by a client (for example, animation video production software) integrated in a business terminal device. The business terminal device may be an object terminal used by object A (ie, a business object). The business terminal device may be the above-mentioned Figure 1 Any terminal device in the terminal device cluster shown, for example, terminal device 100a.

[0170] like Figure 6 As shown, when the business object is input, a trigger operation is executed for the image generation control (for example, control 6K1, i.e., the "start generation" control) in the character creation interface 630J. At this time, the client can obtain the object input data including text data, and then respond to the trigger operation to switch the terminal interface from the character creation interface 630J to the character generation interface (for example, Figure 6 The character generation interface 640J1 shown). The character generation interface 640J1 may include a picture display area, work creation information, and a first display control (e.g., control 6K2) for indicating the progress of picture generation. The work creation information may include a creation object (e.g., object A), a creation timestamp (i.e., a response timestamp for control 6K1), and a description word (i.e., the text data displayed in the text input box in the character creation interface 630J). At the same time, the client may also send a picture generation request determined based on the object input data to the server, so that the server may perform text-to-image conversion on the text to be converted based on the object input data including text data carried in the picture generation request and the first text-to-image generation model, and obtain Z initial pictures, where Z is a positive integer.

[0171] When the client receives the Z initial images returned by the server, the generation progress of the control 6K2 reaches a progress threshold (e.g., 100%), at which point the client can display the Z character images in the character generation interface. For ease of explanation, the number of initial images Z in the embodiment of the present application can be 4, specifically including the four thumbnails displayed in the character generation interface 640J2, namely, image P1, image P2, image P3, and image P4.

[0172] The character generation interface 640J2 may further include a page rollback control (e.g., control 6K6), a refresh control (e.g., control 6K7), and an uncontrollable image selection prompt control (e.g., control 6K3). The page rollback control is used to return from the current interface (i.e., the character generation interface 640J2) to the character creation interface 630J, and the refresh control can be used to regenerate a new initial image according to the object input data determined by object A in the character creation interface 630J.

[0173] If none of the four initial images meet the user's expectations of object A, object A can perform a trigger operation on control 6K6 to return the client to the character creation interface 630J. Object A can re-enter the text in the text input box of the character creation interface 630J to modify the original text data and obtain new text data. At this time, the client can send the object input data including the new text data to the server so that the server can perform text-to-image conversion on the text to be converted determined by the new text data. Optionally, object A can also perform a trigger operation on control 6K7. At this time, the client can send the object input data including the original text data to the server again so that the server can perform text-to-image conversion on the text to be converted determined by the original text data again.

[0174] If there is an initial image (e.g., image P3) among the four initial images that meets the user expectations of object A, then object A performs a trigger operation on image P3 among the four initial images to determine image P3 as the image to be processed. Then, the client can prominently display image P3 and replace the image selection prompt control (e.g., control 6K3) in the character generation interface to display a controllable image generation control (e.g., control 6K4 in the character generation interface 640J3). For ease of distinction, in the embodiment of the present application, the image generation control (e.g., control 6K1) in the character creation interface 630J is referred to as the first image generation control, and the image generation control (e.g., control 6K4, i.e., the "Generate HD Work" control) in the character generation interface 640J3 is referred to as the second image generation control.

[0175] Furthermore, the client can respond to the trigger operation that the object A can perform on the control 6K4, and display a second display control for indicating the progress of image generation in the character generation interface (for example, the control 6K5 in the character generation interface 640J4). At the same time, the client can generate a picture conversion request for the image to be processed (picture P3), and send the picture conversion request to the server so that the server converts the image to be processed to obtain a picture P3 with a second quality coefficient (for example, a high-definition picture), and then send the picture P3 with the second quality coefficient to the client. When the client receives the picture, the generation progress of the control 6K5 reaches the progress threshold. At this time, the client can display the picture P3 with the second quality coefficient in the character generation interface 640J5, and then use the displayed picture P3 as the character picture.

[0176] In an embodiment of the present application, the business object does not need to draw and upload a role picture. Instead, text data that can describe the business role style of the business role to be generated is entered in the text input box. Compared with drawing a role picture, the role information of the business role can be clearly described without spending a lot of drawing time and effort. Therefore, when the computer device obtains the object input data including text data, it can directly obtain the text to be converted based on the data prompt template and the text data, so as to automatically generate the role picture of the object input data later. This means that the role creation method provided in the embodiment of the present application can reduce the difficulty of image generation, greatly lower the threshold for creating roles, and thus improve the efficiency of role generation.

[0177] Further, see Figure 7 , Figure 7 This is a flow chart of a data processing method provided by an embodiment of the present application. The method involves a business terminal device and a server in a role creation system, that is, the method can be performed by a business terminal device (for example, the above Figure 1 The terminal device 100a shown) and the server (for example, the above Figure 1 The method may include at least the following steps S201 to S211:

[0178] Step S201 : The client in the service terminal device responds to the input operation performed by the service object on the text input box in the role creation interface, and obtains object input data including text data.

[0179] Specifically, when a business object wants to create a certain business role, it can perform an input operation on the text input box in the role creation interface provided by the client, so that when the client responds to the input operation, text data describing the business role style of the business role to be generated is displayed in the text input box. The input operation here refers to a trigger operation performed on the text input box for inputting text data. The trigger operation may include contact operations such as clicking and long pressing, or non-contact operations such as voice and gestures, which will not be limited here. The data type of the business role style here is a text type, and the style category in the business role style may include one or more style categories in a style category set. The style category set here may include a virtual clothing accessory style class and a virtual character attribute style class. Furthermore, the client can determine the object input data to be sent to the server based on the text data.

[0180] In this embodiment of the present application, the object input data may include not only the text data displayed in the text input box, but also the picture data (i.e., reference picture data) determined by the business object. It is understandable that when the character creation interface includes a reference picture drawing control (e.g., Figure 4 When the control 4K5 in the character creation interface 430J2 shown is displayed, the client can switch from the character creation interface to the image drawing interface in response to the triggering operation of the business object for the reference image drawing control. The image drawing interface here includes a drawing area and an image storage control. Further, the client can respond to the drawing operation of the business object in the drawing area, display the reference image data corresponding to the drawing operation in the drawing area, and then respond to the triggering operation of the business object for the image storage control, switch from the image drawing interface to the character creation interface, and display the reference image data in the character creation interface. Then, the client can respond to the triggering operation of the business object for the image generation control, and determine the reference image data and text data as object input data.

[0181] In step S202, when the client obtains the object input data, it generates an image generation request and sends the image generation request to the server, so that the server can perform a format comparison on the text format of the text data and the text format of the data prompt template when obtaining the object input data to obtain the comparison result.

[0182] Specifically, when the client obtains the object input data, it can directly generate an image generation request based on the object input data, and then send the image generation request to the server (for example, the client backend) so that the server obtains the object input data including the text data carried by the image generation request. Then, the server can obtain a data prompt template including the initial character style, and then determine the text to be recognized based on the text format and text data of the data prompt template. For example, the server can perform a format comparison on the text format of the text data and the text format of the data prompt template to obtain a comparison result.

[0183] Step S203: If the comparison result indicates that the text format of the text data is inconsistent with the text format of the data prompt template, the server calls a translation service.

[0184] In step S204 , the server translates the text format of the text data into the text format of the data prompt template based on the translation service, and uses the translated text data as the text to be recognized.

[0185] In step S205 , if the comparison result indicates that the text format of the text data is consistent with the text format of the data prompt template, the server uses the text data as text to be recognized and calls the intelligent generation service.

[0186] Specifically, since the data prompt template here can include a first position corresponding to the virtual character attribute style class, a second position corresponding to the initial character style, and a third position corresponding to the virtual clothing accessories style class, when the server obtains the text to be recognized, it can split the text to be recognized based on the split mark in the text to be recognized to obtain N subtexts. The N subtexts here include subtext X i ; N is a positive integer; i is a positive integer less than or equal to N; the business role pattern in the text data includes role patterns corresponding to N subtexts. At this time, the server can identify the subtext X i The corresponding character style can then be based on the subtext X i The corresponding character style determines the subtext X from the first and third positions i Corresponding location information Y i , in the data tip template, change the subtext X i Add to location information Y i , until N sub-texts are added to the data prompt template respectively, the added data prompt template is determined as the text to be converted, and then the intelligent generation service (i.e., AI generation service) can be called.

[0187] Step S206: The server generates a character image corresponding to the object input data based on the intelligent generation service.

[0188] Specifically, the server can obtain a first text-image generation model based on the intelligent generation service, wherein the first text-image generation model (i.e., the trained text-image generation model) is obtained by training a second text-image generation model (the pre-trained text-image generation model) based on sample text data and sample character images, wherein the sample text data is determined based on a data prompt template. Furthermore, the computer device can convert the text to be converted into an image using the first text-image generation model, and then determine the character image corresponding to the object input data based on the image obtained after the text-image conversion.

[0189] For example, if the object input data includes text data, the server can directly input the text to be converted into the first text-to-image generation model, and then perform text-to-image conversion on the text to be converted through the first text-to-image generation model, and then directly determine the character image corresponding to the object input data from the initial image after the text-to-image conversion.

[0190] For another example, if the object input data includes text data and reference image data, the server can input the text to be converted and the reference image data into the first text-image generation model together. The reference image data here refers to the image data input by the business object in the role creation interface of the client. At this time, the server can perform text-image conversion on the text to be converted through the first text-image generation model, and determine the key part image corresponding to the text to be converted from the initial image after the text-image conversion. Furthermore, the server can perform image recognition on the reference image data to obtain the image to be replaced that has a part matching relationship with the key part image, replace the image to be replaced with the key part image in the reference image data, and determine the replaced reference image data as the role image corresponding to the object input data.

[0191] For further understanding, please refer to Figure 8 , Figure 8 This is a schematic diagram of an interface switching for displaying a character picture corresponding to the second object input data provided by an embodiment of the present application. The second object input data here may include two data types: text data and reference text and image data. Figure 8 As shown, the interfaces shown in the embodiments of the present application are provided by a client (for example, animation video production software) integrated in a business terminal device. The business terminal device may be an object terminal used by object A (ie, a business object). The business terminal device may be the above-mentioned Figure 1 Any terminal device in the terminal device cluster shown, for example, terminal device 100a.

[0192] like Figure 8 As shown, the character creation interface 830J1 may include a reference image drawing control (e.g., control 8K1, i.e., a "graffiti reference image" control). When object A performs a trigger operation on control 8K1, the client may respond to the trigger operation and switch from the character creation interface 830J1 to the image drawing interface 850J. The image drawing interface 850J here may include a drawing area and an image storage control (e.g., control 8K2, i.e., a "save graffiti" control). When object A performs a drawing operation (a trigger operation for drawing reference image data) in the drawing area, the client may respond to the drawing operation when the drawing is completed. Figure 8 The drawing area shown displays reference picture data corresponding to the drawing operation (for example, Figure 8 Then, the object A may perform a trigger operation on the control 8K2, so that the client switches from the image drawing interface 850J to the character creation interface (e.g., the character creation interface 830J2) in response to the trigger operation, and then the image data 80P may be displayed in the character creation interface 830J2.

[0193] At the same time, the object A can also perform an input operation on the text input box in the role creation interface 830J2 to display text data (for example, Figure 8 Furthermore, when object A completes input, object A can trigger an image generation control (e.g., control 8K3, i.e., a "start generation" control) in the character creation interface 830J2 to determine the text data 8T and the image data 80P as object input data. Furthermore, an image generation request can be generated based on the object input data and sent to the server.

[0194] When the server obtains the object input data carried in the image generation request, it can first determine the text to be converted based on the text data 8T and the data prompt template. Then, when obtaining the first text-to-image generation model, it can input the text to be converted and the image data 80P into the first text-to-image generation model. The first text-to-image generation model performs text-to-image conversion on the text to be converted, and determines the key part image corresponding to the text to be converted based on the initial image after the text-to-image conversion. The key part image here can include a first key image of the head and a second key image of the pants.

[0195] Furthermore, the server can perform image recognition on the image data 80P to obtain images to be replaced that have a part matching relationship with the first key image (for example, Figure 8 The picture 8P1 shown), the picture to be replaced that has a part matching relationship with the second key picture (for example, Figure 8 Then, the server can replace the two to-be-replaced pictures with the corresponding key parts pictures in the picture data 80P, and determine the replaced picture data 80P as the character picture corresponding to the object input data (for example, Figure 8 Furthermore, the server may send the character image 81P to the client, so that the client displays the character image 81P in a character generation interface (eg, character generation interface 840J).

[0196] The specific implementation of steps S201 and S206 can be found in the above Figure 3 The description of steps S101 to S103 in the corresponding embodiment will not be repeated here.

[0197] Step S207: The server calls the image storage service.

[0198] In step S208, the server stores the character picture based on the picture storage service and generates a picture storage link associated with the character picture.

[0199] Step S209: The server sends the image storage link to the client.

[0200] Step S210 : The client responds to the triggering operation performed by the business object on the picture storage link, and stores and displays the character picture.

[0201] In step S211 , the client responds to the triggering operation performed by the business object on the role image and creates a business role to be generated.

[0202] It should be understood that the business object can perform a trigger operation on the character picture so that the client generates a skeleton material request for sending to the server based on the character picture, so that the server can issue the skeleton material (for example, Spine skeleton material) corresponding to the character picture to the client based on the skeleton material request. Among them, Spine is a 2D skeleton animation editing tool for game development, and provides a set of skeleton animation data definitions and runtime libraries. The skeleton animation of the present invention adopts Spine skeleton animation and uses the Spine data structure. Furthermore, when the client receives the skeleton material returned by the server, it can display the character picture in the material cutting interface. At this time, the business object can cut the character picture based on the cutting prompt information in the material cutting interface to display the cut character picture in the material editing interface. Among them, the material editing interface can include intelligent cutout controls and facial features selection controls.

[0203] If the business object performs a trigger operation on the smart cutout control, the client can intelligently cut out the background image corresponding to the cropped character image when responding to the trigger operation. At this time, the client can bind the character image after cutting out the background image with the skeleton material returned by the server to obtain the image to be edited. Furthermore, if the business object performs a trigger operation on the facial features selection control, the client can respond to the trigger operation and generate a facial features material request for sending to the server, so that the server can issue facial features mapping materials (for example, Spine facial features mapping materials) to the client based on the facial features material request. When the client receives the facial features mapping materials, it can display multiple facial features mapping materials in the material editing interface. Furthermore, the business object can perform a trigger operation on a facial features mapping material among multiple facial features mapping materials, so that the client can fit the facial features mapping material corresponding to the trigger operation to the corresponding position in the image to be edited to obtain the image to be added.

[0204] Then, the client can respond to the trigger operation for the image to be added and generate an animation material request for sending to the server, so that the server can issue animation materials (for example, Spine animation materials) for associating the skeleton and facial features to the client based on the animation material request. When the client receives the animation materials, it can display multiple animation materials in the animation adding interface. Furthermore, the business object can perform trigger operations on certain animation materials among the multiple animation materials, so that the client can generate complete custom animation materials based on the animation materials corresponding to the trigger operation and display the preview effect in the animation adding interface.

[0205] Furthermore, the business object can perform an information editing operation on the character corresponding to the custom animation material, so that when the client responds to the information editing operation, it obtains the character's corresponding character information (such as the character name and character voice, etc.), thereby obtaining the business character, and submitting the business character to the server for storage. At this point, the business character has been created, and subsequent business objects can use the business character to create stories and add the business character to their current draft for preview display.

[0206] It is understandable that the business object can view the historical roles created by the business object and the historical text data corresponding to the historical roles based on the first role viewing control included in the role creation interface. For ease of understanding, please refer to Figure 9 , Figure 9 This is a schematic diagram of a scenario for viewing business roles provided by an embodiment of the present application. Figure 9 As shown, the interfaces shown in the embodiments of the present application are provided by a client (for example, animation video production software) integrated in a business terminal device. The business terminal device may be an object terminal used by object A (ie, a business object). The business terminal device may be the above-mentioned Figure 1 Any terminal device in the terminal device cluster shown, for example, terminal device 100a.

[0207] like Figure 9 As shown, the character creation interface 930J1 may include a first character viewing control (e.g., control 9K1, i.e., the "My Works" control) and a second character viewing control (e.g., control 9K2, i.e., the "Reference Case" control). The character creation interface 930J1 may display the reference character and reference text data corresponding to the second character viewing control.

[0208] When object A has not created a character, after object A performs a trigger operation on control 9K1, the client can respond to the trigger operation and display a text message such as "Looking forward to the birth of a masterpiece..." in the work display area of ​​the character creation interface (for example, the work display area 9Q1 in the character creation interface 930J2).

[0209] Optionally, if object A has previously created characters, after object A triggers control 9K1, the client can respond to the trigger by displaying the characters created by object A and the corresponding historical text data in the work display area of ​​the character creation interface (e.g., work display area 9Q2 in character creation interface 930J3). Furthermore, the client can also display the number of historical characters created by object A (e.g., 2) in the first character viewing control.

[0210] It is understandable that when a business role (e.g. Figure 8 After the business object performs a trigger operation on the control 9K1, the client can display the picture 81P in the work display area of ​​the character creation interface.

[0211] It can be seen that after the server in the embodiment of the present application determines the text to be converted based on the text data input by the business object and the data prompt template, it can perform text-to-image conversion on the text to be converted to generate multiple groups of related pictures (i.e., Z initial pictures) that match the text data. Then the business object can also select a satisfactory initial picture from the Z initial pictures for conversion processing to obtain a character picture with a second quality coefficient (i.e., a high-definition picture). Further, the client can subsequently respond to the trigger operation of the business object, automatically generate a Spine skeleton animation character for the character picture, and automatically match it with action animation and corresponding tone configuration to obtain a personalized business role. This role creation method can quickly and accurately generate business roles that meet user needs.

[0212] Further, see Figure 10 , Figure 10 This is a flow chart of a data processing method provided by an embodiment of the present application. Figure 10 As shown, the method can be executed by the service terminal device used by the service object, and the service terminal device can be the above Figure 1 Any terminal device in the terminal device cluster shown, for example, terminal device 100a. The method may at least include the following steps S301 to S304:

[0213] Step S301: Display the character creation interface.

[0214] Specifically, after starting the client, the service terminal device may display the role publishing interface provided by the client (for example, Figure 4 The role publishing interface 410J1 shown in FIG. 4 may include a role creation control. Further, the client may respond to a trigger operation of the business object on the role creation control and display a creation mode selection sub-interface (e.g., Figure 4 The creation method selection sub-interface 420J shown in the figure) is an interface for selecting a creation method, wherein the creation method selection sub-interface here can be an interface superimposed on the character publishing interface, and the interface size of the creation method selection sub-interface is smaller than the interface size of the character publishing interface. In addition, the creation method selection sub-interface can include M creation method selection controls; M is a positive integer; and the M creation method selection controls include an intelligent creation method selection control. Then, the client can respond to the triggering operation of the business object for the intelligent creation method selection control and display the character creation interface (for example, the character creation interface 430J1 or the character creation interface 430J2). The character creation interface can include a text input box and an image generation control.

[0215] Step S302 : In response to an input operation on the text input box, text data describing a business role style of the business role to be generated is displayed.

[0216] Specifically, the business object can perform an input operation on a text input box in the character creation interface, so that the client can display text data describing the business character style of the business character to be generated in the text input box. The data type of the business character style is text; the style category of the business character style includes one or more style categories from a style category set; the style category set includes a virtual clothing and accessories style class and a virtual character attribute style class.

[0217] Step S303: When object input data including text data is acquired, the character creation interface is switched to the character generation interface in response to a triggering operation on the image generation control.

[0218] Specifically, when the business object is input, the client can obtain the object input data including text data, and then when the business object performs a trigger operation on the image generation control in the character creation interface, the client can respond to the trigger operation and switch from the character creation interface to the character generation interface.

[0219] Step S304: displaying the character image in the character generation interface.

[0220] Among them, the role image is used to generate the business role to be generated (that is, the business role indicated by the object input data); the role image is obtained after the text to be converted is converted into a picture; the text to be converted here is jointly determined by the business role style and the initial role style in the data prompt template; the data type of the data prompt template is a text type, and the initial role style in the data prompt template is a pointer to the default role style configured for the business role to be generated.

[0221] The specific implementation of steps S301 to S304 can be found in the above Figure 3 The description of steps S101 to S103 in the corresponding embodiment will not be repeated here.

[0222] Further, see Figure 11 , Figure 11 This is a structural diagram of a data processing device provided in an embodiment of the present application. Figure 11 As shown, the data processing device 1 may include: a data acquisition module 100 , a style adding module 200 and a text-to-image conversion module 300 .

[0223] The data acquisition module 100 is configured to acquire object input data including text data; the text data includes a business role style for describing a business role to be generated, and the data type of the business role style is a text type; the style category in the business role style includes one or more style categories in a style category set; the style category set includes a virtual clothing accessory style class and a virtual character attribute style class;

[0224] The style adding module 200 is used to obtain a data prompt template including an initial role style, add a business role style to the data prompt template, and determine the added data prompt template as the text to be converted; the data type of the data prompt template is a text type; the initial role style is the default role style configured for the business role to be generated;

[0225] The text-to-image conversion module 300 is used to perform text-to-image conversion on the text to be converted based on the object input data, and obtain a role image corresponding to the object input data; the role image is used to generate the business role to be generated.

[0226] The specific implementation of the data acquisition module 100, the style adding module 200 and the text-to-image conversion module 300 can be found in the above Figure 3 The description of steps S101 to S103 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here either.

[0227] Further, see Figure 12 , Figure 12 This is a structural diagram of a data processing device provided in an embodiment of the present application. Figure 12 As shown, the data processing device 2 may include: a data acquisition module 10, a style adding module 20, a text-to-image conversion module 30, a sample acquisition module 40, a sample input module 50, a model training module 60, an image storage module 70 and a link sending module 80.

[0228] The data acquisition module 10 is configured to acquire object input data including text data; the text data includes a business role style for describing a business role to be generated, and the data type of the business role style is a text type; the style category in the business role style includes one or more style categories in a style category set; the style category set includes a virtual clothing accessory style class and a virtual character attribute style class;

[0229] The style adding module 20 is used to obtain a data prompt template including an initial role style, add a business role style to the data prompt template, and determine the added data prompt template as the text to be converted; the data type of the data prompt template is a text type; the initial role style is the default role style configured by the pointer to the business role to be generated.

[0230] The style adding module 20 includes: a template acquiring unit 201 , a to-be-recognized text determining unit 202 , a splitting processing unit 203 , a position determining unit 204 and a style adding unit 205 .

[0231] The template acquisition unit 201 is used to acquire a data prompt template including an initial character style; the data prompt template includes a first position corresponding to a virtual character attribute style class, a second position corresponding to the initial character style, and a third position corresponding to a virtual clothing and accessories style class;

[0232] The to-be-recognized text determining unit 202 is configured to determine the to-be-recognized text based on the text format and text data of the data prompt template.

[0233] The to-be-recognized text determination unit 202 includes: a format determination subunit 2021 , a format comparison subunit 2022 , a first determination subunit 2023 and a second determination subunit 2024 .

[0234] The format determination subunit 2021 is configured to use the text format of the data prompt template as the first text format and the text format of the text data as the second text format;

[0235] The format comparison subunit 2022 is used to perform a format comparison on the first text format and the second text format to obtain a comparison result;

[0236] The first determining subunit 2023 is configured to determine the text data as text to be recognized if the comparison result indicates that the first text format is consistent with the second text format;

[0237] The second determining subunit 2024 is configured to call a translation service if the comparison result indicates that the first text format is inconsistent with the second text format, and based on the translation service, translate the text format of the text data from the second text format to the first text format, and determine the translated text data as the text to be recognized.

[0238] The specific implementation of the format determination subunit 2021, the format comparison subunit 2022, the first determination subunit 2023 and the second determination subunit 2024 can be found in the above Figure 3 The description of the text to be recognized in the corresponding embodiment will not be repeated here.

[0239] The splitting processing unit 203 is used to split the text to be recognized based on the splitting mark in the text to be recognized to obtain N subtexts; the N subtexts include subtext X i ; N is a positive integer; i is a positive integer less than or equal to N; the business role style in the text data includes role styles corresponding to N sub-texts respectively;

[0240] The position determination unit 204 is used to identify the subtext X i Corresponding character style, based on subtext X i The corresponding character style determines the subtext X from the first and third positions i Corresponding location information Y i .

[0241] The location determination unit 204 includes: a category identification subunit 2041 , a third determination subunit 2042 and a fourth determination subunit 2043 .

[0242] The category identification subunit 2041 is used to identify the subtext X i The style category of the corresponding character style;

[0243] The third determining subunit 2042 is used to determine if the subtext X i If the style category of the corresponding character style belongs to the virtual character attribute style category, the first position is determined as the subtext X i Corresponding location information Y i ;

[0244] The fourth determining subunit 2043 is used to determine if the subtext X i If the style category of the corresponding character style belongs to the virtual clothing accessories style category, the third position is determined as the subtext X i Corresponding location information Y i .

[0245] The specific implementation of the category identification subunit 2041, the third determination subunit 2042 and the fourth determination subunit 2043 can be found in the above Figure 3 The description of the location information in the corresponding embodiment will not be repeated here.

[0246] The style adding unit 205 is used to add the subtext X in the data prompt template. i The corresponding character style is added to the position information Y i , until N role styles are added to the data prompt template respectively, and the added data prompt template is determined as the text to be converted.

[0247] The specific implementation of the template acquisition unit 201, the to-be-recognized text determination unit 202, the splitting processing unit 203, the position determination unit 204 and the style adding unit 205 can be found in the above Figure 3 The description of step S102 in the corresponding embodiment will not be repeated here.

[0248] The text-to-image conversion module 30 is used to perform text-to-image conversion on the text to be converted based on the object input data to obtain a role image corresponding to the object input data; the role image is used to generate the business role to be generated.

[0249] The text-to-image conversion module 30 includes: a model acquisition unit 301 , a first data input unit 302 , a first text-to-image conversion unit 303 , a second data input unit 304 , a second text-to-image conversion unit 305 and an image replacement unit 306 .

[0250] The model acquisition unit 301 is used to call the intelligent generation service to obtain a first text-image generation model; the first text-image generation model is obtained by training a second text-image generation model based on sample text data and sample character images; the sample text data is determined based on a data prompt template;

[0251] The first data input unit 302 is configured to input the text to be converted into the first text-to-image generation model if the object input data includes text data;

[0252] The first text-to-image conversion unit 303 is configured to perform text-to-image conversion on the text to be converted using a first text-to-image generation model to obtain a character image corresponding to the object input data.

[0253] The first text-image generation model includes an encoder and an image generator; the image generator includes an image information generator and an image decoder;

[0254] The first text-to-image conversion unit 303 includes: an encoding processing subunit 3031 , an information extraction subunit 3032 , a text-to-image conversion subunit 3033 and a character picture determination subunit 3034 .

[0255] The encoding processing subunit 3031 is used to input the text to be converted into the encoder, and the encoder performs encoding processing on the text to be converted to obtain a text semantic vector corresponding to the text to be converted;

[0256] The information extraction subunit 3032 is used to input the text semantic vector into the image information generator, and extract information from the text semantic vector through the image information generator to obtain the image information latent vector; the image information latent vector is used to reflect the image information of the text to be converted;

[0257] The text-to-image conversion subunit 3033 is used to input the latent vector of the image information into the image decoder, and perform text-to-image conversion on the latent vector of the image information through the image decoder to obtain the initial image corresponding to the object input data;

[0258] The character picture determination subunit 3034 is configured to determine the character picture corresponding to the object input data from the initial picture.

[0259] The specific implementation of the encoding processing subunit 3031, the information extraction subunit 3032, the text-to-image conversion subunit 3033 and the character image determination subunit 3034 can be found in the above Figure 3 The description of the first text-image generation model in the corresponding embodiment will not be repeated here.

[0260] The second data input unit 304 is configured to input the text to be converted and the reference image data into the first text-image generation model if the object input data includes text data and reference image data; the reference image data refers to the image data input by the business object in the role creation interface of the client;

[0261] The second text-to-image conversion unit 305 is configured to perform text-to-image conversion on the text to be converted using the first text-to-image generation model to obtain images of key parts corresponding to the text to be converted;

[0262] The picture replacement unit 306 is used to perform picture recognition on the reference picture data to obtain a picture to be replaced that has a part matching relationship with the key part picture. In the reference picture data, the picture to be replaced is replaced with the key part picture, and the replaced reference picture data is determined as the character picture corresponding to the object input data.

[0263] The specific implementation of the model acquisition unit 301, the first data input unit 302, the first text-to-image conversion unit 303, the second data input unit 304, the second text-to-image conversion unit 305 and the image replacement unit 306 can be found in the above Figure 3 The description of step S103 in the corresponding embodiment will not be repeated here.

[0264] The sample acquisition module 40 is used to acquire sample text data determined based on the data prompt template and a sample character image corresponding to the sample text data;

[0265] The sample input module 50 is used to input the sample text data into the second text-image generation model, perform text-image conversion on the sample text data through the second text-image generation model, and obtain a predicted character image corresponding to the sample text data;

[0266] The model training module 60 is used to train the second text-image generation model based on the sample character images and the predicted character images to obtain the first text-image generation model.

[0267] The object input data is obtained from the image generation request sent by the client;

[0268] The picture storage module 70 is used to call the picture storage service after generating the character picture corresponding to the object input data, store the character picture through the picture storage service, and generate a picture storage link associated with the character picture;

[0269] The link sending module 80 is used to send the picture storage link to the client, so that the client displays the character picture when responding to the trigger operation on the picture storage link.

[0270] The specific implementation of the data acquisition module 10, the style adding module 20, the text-to-image conversion module 30, the sample acquisition module 40, the sample input module 50, the model training module 60, the image storage module 70 and the link sending module 80 can be found in the above Figure 7 The description of steps S201 to S211 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here either.

[0271] Further, see Figure 13 , Figure 13 This is a structural diagram of a data processing device provided in an embodiment of the present application. Figure 13 As shown, the data processing device 3 may include: a creation interface display module 400, a text data display module 500, a generation interface display module 600 and a character picture display module 700.

[0272] The creation interface display module 400 is used to display the character creation interface; the character creation interface includes a text input box and a picture generation control;

[0273] The text data display module 500 is configured to display text data describing a business role style of a business role to be generated in response to an input operation on the text input box; the data type of the business role style is a text type; the style category of the business role style includes one or more style categories in a style category set; the style category set includes a virtual clothing and accessories style class and a virtual character attribute style class;

[0274] The generation interface display module 600 is used to switch from the character creation interface to the character generation interface in response to a trigger operation on the image generation control when object input data including text data is obtained;

[0275] The role picture display module 700 is used to display the role picture in the role generation interface; the role picture is used to generate the business role to be generated; the role picture is obtained after the text to be converted is converted into an image; the text to be converted is determined by the business role style and the initial role style in the data prompt template; the data type of the data prompt template is a text type; the initial role style is a pointer to the default role style configured for the business role to be generated.

[0276] The specific implementation of the creation interface display module 400, the text data display module 500, the generation interface display module 600 and the character image display module 700 can be found in the above Figure 10 The description of steps S301 to S304 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here either.

[0277] Further, see Figure 14 , Figure 14 This is a structural diagram of a data processing device provided in an embodiment of the present application. Figure 14 As shown, the data processing device 4 may include: a creation interface display module 91, a text data display module 92, a generation interface display module 93, a character picture display module 94, a drawing interface display module 95, a reference picture drawing module 96, a reference picture display module 97 and an input data determination module 98.

[0278] The creation interface display module 91 is used to display the character creation interface; the character creation interface includes a text input box and a picture generation control.

[0279] The creation interface display module 91 includes: a publishing interface display unit 911 , a sub-interface display unit 912 and a creation interface display unit 913 .

[0280] The release interface display unit 911 is used to display the role release interface; the role release interface includes a role creation control;

[0281] The sub-interface display unit 912 is configured to display a creation mode selection sub-interface on the character publishing interface in response to a trigger operation on the character creation control; the creation mode selection sub-interface is an interface superimposed on the character publishing interface, and the interface size of the creation mode selection sub-interface is smaller than the interface size of the character publishing interface; the creation mode selection sub-interface includes M creation mode selection controls; M is a positive integer; the M creation mode selection controls include an intelligent creation mode selection control;

[0282] The creation interface display unit 913 is used to respond to the triggering operation of the intelligent creation mode selection control and display the character creation interface.

[0283] The specific implementation of the publishing interface display unit 911, the sub-interface display unit 912 and the creation interface display unit 913 can be found in the above Figure 4 The description of the character creation interface in the corresponding embodiment will not be repeated here.

[0284] The text data display module 92 is used to respond to input operations on the text input box and display text data used to describe the business role style of the business role to be generated; the data type of the business role style is a text type; the style category in the business role style includes one or more style categories in a style category set; the style category set includes a virtual clothing accessory style class and a virtual role attribute style class.

[0285] Among them, the character creation interface includes text reference controls;

[0286] The text data display module 92 includes a prompt information display unit 921 and a text data display unit 922 .

[0287] The prompt information display unit 921 is used to display prompt information associated with the text reference control on the character creation interface in response to a trigger operation on the image generation control if the text input box does not display text data; the prompt information is used to instruct the business object to enter text data in the text input box based on the text reference control;

[0288] The text data display unit 922 is used to respond to the trigger operation of the text reference control when the prompt information is closed, and display text data describing the business role style of the business role to be generated in the text input box; the text data is text data extracted from the high-quality vocabulary.

[0289] The specific implementation of the prompt information display unit 921 and the text data display unit 922 can be found in the above Figure 5 The description of the text reference control in the corresponding embodiment will not be repeated here.

[0290] The generation interface display module 93 is used to switch from the character creation interface to the character generation interface in response to a trigger operation on the image generation control when object input data including text data is obtained;

[0291] The role picture display module 94 is used to display the role picture in the role generation interface; the role picture is used to generate the business role to be generated; the role picture is obtained after the text to be converted is converted into a picture; the text to be converted is determined by the business role style and the initial role style in the data prompt template; the data type of the data prompt template is a text type; the initial role style is a pointer to the default role style configured for the business role to be generated.

[0292] The character generation interface is displayed in response to a trigger operation on the first image generation control; the first image generation control is an image generation control in the character creation interface;

[0293] The character picture display module 94 includes: a first control display unit 941 , an initial picture display unit 942 , a second control display unit 943 , a third control display unit 944 and a character picture display unit 945 .

[0294] The first control display unit 941 is used to display a first display control for indicating the progress of image generation in the character generation interface;

[0295] The initial image display unit 942 is configured to display Z initial images corresponding to the object input data in the character generation interface when the generation progress of the first display control reaches a progress threshold; Z is a positive integer; each initial image is an image generated based on the object input data and has a first quality coefficient;

[0296] The second control display unit 943 is configured to display a second picture generation control when a picture to be processed is determined from the Z initial pictures;

[0297] The third control display unit 944 is configured to respond to a triggering operation on the second image generation control and display a second display control for indicating the image generation progress in the character generation interface;

[0298] The character picture display unit 945 is used to display the to-be-processed picture with the second quality coefficient in the character generation interface when the generation progress of the second display control reaches the progress threshold, and use the displayed to-be-processed picture as the character picture; the second quality coefficient is higher than the first quality coefficient.

[0299] The specific implementation of the first control display unit 941, the initial image display unit 942, the second control display unit 943, the third control display unit 944 and the character image display unit 945 can be found in the above Figure 6 The description of the character pictures in the corresponding embodiment will not be repeated here.

[0300] Among them, the role creation interface includes a first role viewing control and a second role viewing control; the first role viewing control is used to view the historical roles created by the business object and the historical text data corresponding to the historical roles; the second role viewing control is used to view the reference roles and the reference text data corresponding to the reference roles; the reference roles are H selected roles selected from the role library; H is a positive integer.

[0301] Among them, the character creation interface includes reference image drawing controls;

[0302] The drawing interface display module 95 is used to switch from the character creation interface to the image drawing interface in response to a trigger operation on the reference image drawing control; the image drawing interface includes a drawing area and an image storage control;

[0303] The reference picture drawing module 96 is configured to respond to a drawing operation in the drawing area and display reference picture data corresponding to the drawing operation in the drawing area;

[0304] The reference picture display module 97 is used to respond to the trigger operation on the picture storage control, switch from the picture drawing interface to the character creation interface, and display the reference picture data in the character creation interface;

[0305] The input data determination module 98 is used to determine the reference image data and text data as object input data.

[0306] The specific implementation of the creation interface display module 91, the text data display module 92, the generation interface display module 93, the character image display module 94, the drawing interface display module 95, the reference image drawing module 96, the reference image display module 97 and the input data determination module 98 can be found in the above Figure 7 The description of steps S201 to S211 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here either.

[0307] Further, see Figure 15 , Figure 15 This is a schematic diagram of a computer device provided in an embodiment of the present application. Figure 15 As shown, the computer device 1000 may include: at least one processor 1001, such as a CPU, at least one network interface 1004, a memory 1005, and at least one communication bus 1002. The communication bus 1002 is used to implement connection and communication between these components. The network interface 1004 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage. The memory 1005 may optionally also be at least one storage device located away from the aforementioned processor 1001. As Figure 15 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module and a device control application. In some embodiments, the computer device may also include Figure 15 The user interface 1003 shown, for example, if the computer device is Figure 1 The terminal device with data processing function shown (for example, terminal device 100b) may further include the user interface 1003, wherein the user interface 1003 may include a display screen (Display), a keyboard (Keyboard), etc.

[0308] exist Figure 15 In the computer device 1000 shown, the network interface 1004 is mainly used for network communication; the user interface 1003 is mainly used for providing an input interface for the user; and the processor 1001 can be used to call the device control application stored in the memory 1005.

[0309] It should be understood that the computer device 1000 described in the embodiment of the present application can execute the above Figure 3 、 Figure 7 or Figure 10 The description of the data processing method in the corresponding embodiment can also be performed as described above. Figure 11 In the corresponding embodiment, the data processing device 1, Figure 12 In the corresponding embodiment, the data processing device 2, Figure 13 In the corresponding embodiment, the data processing device 3 or Figure 14 The description of the data processing device 4 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here either.

[0310] The present invention also provides a computer-readable storage medium that stores a computer program. The computer program includes program instructions that are executed by a processor to implement Figure 3 、 Figure 7 or Figure 10 For details on the data processing methods provided in each step, please refer to Figure 3 、 Figure 7 or Figure 10 The implementation methods provided by each step will not be repeated here.

[0311] The computer-readable storage medium may be the data transmission device provided in any of the aforementioned embodiments or the internal storage unit of the computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Furthermore, the computer-readable storage medium may also include both the internal storage unit of the computer device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.

[0312] The present application also provides a computer program product, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, enabling the computer device to perform the data processing methods or apparatuses described in the preceding embodiments, which are not further detailed here. Furthermore, the beneficial effects of the same methods are not further detailed here.

[0313] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The above-described program can be stored in a computer-readable storage medium. When executed, the program can include the processes in the above-described method embodiments. The above-described storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0314] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.< / svp> < / svp> < / svp> < / svp> < / svp> < / svp>

Claims

1. A data processing method, characterized in that: include: Obtaining object input data including text data; the text data including a business role style for describing a business role to be generated, and the data type of the business role style being a text type; the style category in the business role style including one or more style categories in a style category set; the style category set including a virtual clothing and accessories style class and a virtual character attribute style class; Obtaining a data prompt template including an initial role style, adding the business role style to the data prompt template, and determining the added data prompt template as text to be converted; the data type of the data prompt template is a text type; the initial role style is a default role style configured for the business role to be generated; Based on the object input data, performing text-to-image conversion on the text to be converted to obtain a character image corresponding to the object input data; The role image is used to generate the business role to be generated.

2. The method according to claim 1, characterized in that The step of obtaining a data prompt template including an initial role style, adding the business role style to the data prompt template, and determining the added data prompt template as the text to be converted includes: Acquire a data prompt template including an initial character style; the data prompt template includes a first position corresponding to the virtual character attribute style class, a second position corresponding to the initial character style, and a third position corresponding to the virtual clothing and accessories style class; Determining the text to be recognized based on the text format of the data prompt template and the text data; Based on the splitting mark in the text to be recognized, the text to be recognized is split to obtain N subtexts; the N subtexts include subtext X i ; N is a positive integer; i is a positive integer less than or equal to N; the business role style in the text data includes the role styles corresponding to the N sub-texts respectively; Identify the subtext X i The corresponding character style, based on the subtext X i Corresponding character style, from the first position and the third position, determine the subtext X i Corresponding location information Y i ; In the data prompt template, change the subtext to X i The corresponding character style is added to the position information Y i , until N character styles are respectively added to the data prompt template, and the added data prompt template is determined as the text to be converted.

3. The method according to claim 2, characterized in that The determining of the text to be recognized based on the text format of the data prompt template and the text data includes: Using the text format of the data prompt template as the first text format and using the text format of the text data as the second text format; performing a format comparison on the first text format and the second text format to obtain a comparison result; If the comparison result indicates that the first text format is consistent with the second text format, determining the text data as text to be recognized; If the comparison result indicates that the first text format is inconsistent with the second text format, a translation service is called, and based on the translation service, the text format of the text data is translated from the second text format to the first text format, and the translated text data is determined as the text to be recognized.

4. The method according to claim 2, characterized in that The identification of the subtext X i The corresponding character style, based on the subtext X i Corresponding character style, from the first position and the third position, determine the subtext X i Corresponding location information Y i ,include: Identify the subtext X i The style category of the corresponding character style; If the subtext X i If the style category of the corresponding character style belongs to the virtual character attribute style category, the first position is determined as the subtext X i Corresponding location information Y i ; If the subtext X i If the style category of the corresponding character style belongs to the virtual clothing accessories style category, the third position is determined as the subtext X i Corresponding location information Y i .

5. The method according to claim 1, wherein The step of performing text-to-image conversion on the text to be converted based on the object input data to obtain a character image corresponding to the object input data includes: Invoke the intelligent generation service to obtain a first text-image generation model; the first text-image generation model is obtained by training a second text-image generation model based on sample text data and sample character images; the sample text data is determined based on the data prompt template; If the object input data includes text data, inputting the text to be converted into the first text-to-image generation model; The text to be converted is converted into an image through the first text-to-image generation model to obtain a character image corresponding to the object input data.

6. The method according to claim 5, characterized in that The first text-image generation model includes an encoder and an image generator; the image generator includes an image information generator and an image decoder; The step of performing text-to-image conversion on the text to be converted by using the first text-to-image generation model to obtain a character image corresponding to the object input data includes: Inputting the text to be converted into the encoder, encoding the text to be converted by the encoder to obtain a text semantic vector corresponding to the text to be converted; Inputting the text semantic vector into the image information generator, extracting information from the text semantic vector through the image information generator to obtain an image information latent vector; the image information latent vector is used to reflect the image information of the text to be converted; Inputting the image information latent vector into the image decoder, and performing text-to-image conversion on the image information latent vector by the image decoder to obtain an initial image corresponding to the object input data; A character picture corresponding to the object input data is determined from the initial picture.

7. The method according to claim 5, characterized in that The method further comprises: If the object input data includes text data and reference image data, the text to be converted and the reference image data are input into the first text-image generation model; the reference image data refers to the image data input by the business object in the role creation interface of the client; Performing text-to-image conversion on the text to be converted using the first text-to-image generation model to obtain images of key parts corresponding to the text to be converted; Perform image recognition on the reference image data to obtain a picture to be replaced that has a part matching relationship with the key part picture. In the reference image data, replace the picture to be replaced with the key part picture, and determine the replaced reference image data as the character picture corresponding to the object input data.

8. The method according to claim 5, characterized in that The method further comprises: Obtaining sample text data determined based on a data prompt template and a sample character image corresponding to the sample text data; Inputting the sample text data into the second text-to-image generation model, performing text-to-image conversion on the sample text data through the second text-to-image generation model to obtain a predicted character image corresponding to the sample text data; Based on the sample character pictures and the predicted character pictures, the second text-image generation model is trained to obtain the first text-image generation model.

9. The method according to claim 1, characterized in that The object input data is obtained from the image generation request sent by the client; The method further comprises: After generating a character picture corresponding to the object input data, calling a picture storage service, storing the character picture through the picture storage service, and generating a picture storage link associated with the character picture; The picture storage link is sent to the client, so that the client displays the character picture when responding to a trigger operation on the picture storage link.

10. A data processing method, characterized in that: include: Display the character creation interface; The character creation interface includes a text input box and a picture generation control; In response to an input operation on the text input box, text data describing a business role style of the business role to be generated is displayed; the data type of the business role style is a text type; the style category in the business role style includes one or more style categories in a style category set; the style category set includes a virtual clothing and accessories style class and a virtual character attribute style class; When object input data including text data is acquired, responding to a trigger operation on the image generation control, switching from the character creation interface to the character generation interface; Displaying a role picture in the role generation interface; the role picture is used to generate the business role to be generated; The character image is obtained after performing text-to-image conversion on the text to be converted; The text to be converted is determined by the business role style and the initial role style in the data prompt template; the data type of the data prompt template is a text type; The initial role style refers to a default role style configured for the business role to be generated.

11. The method according to claim 10, characterized in that The display character creation interface includes: Displaying a character publishing interface; the character publishing interface includes a character creation control; In response to a trigger operation on the character creation control, a creation method selection sub-interface is displayed on the character publishing interface; the creation method selection sub-interface is an interface superimposed on the character publishing interface, and the interface size of the creation method selection sub-interface is smaller than the interface size of the character publishing interface; the creation method selection sub-interface includes M creation method selection controls; M is a positive integer; the M creation method selection controls include an intelligent creation method selection control; In response to a triggering operation on the intelligent creation mode selection control, a character creation interface is displayed.

12. The method according to claim 10, characterized in that The role creation interface includes a first role viewing control and a second role viewing control; the first role viewing control is used to view the historical roles created by the business object and the historical text data corresponding to the historical roles; the second role viewing control is used to view the reference role and the reference text data corresponding to the reference role; The reference characters are H selected characters from the character library; H is a positive integer.

13. The method according to claim 10, characterized in that The character creation interface includes a text reference control; The response to the input operation on the text input box is to display text data for describing the business role style of the business role to be generated, including: If the text input box does not display text data, then in response to a trigger operation on the image generation control, prompt information associated with the text reference control is displayed on the character creation interface; the prompt information is used to instruct the business object to enter text data in the text input box based on the text reference control; When the prompt information is closed, in response to a triggering operation on the text reference control, text data describing a business role style of the business role to be generated is displayed in the text input box; the text data is text data extracted from a high-quality vocabulary.

14. The method according to claim 10, characterized in that The character creation interface includes a reference image drawing control; The method further comprises: In response to a triggering operation on the reference image drawing control, switching from the character creation interface to an image drawing interface; the image drawing interface includes a drawing area and an image storage control; In response to a drawing operation in the drawing area, displaying reference picture data corresponding to the drawing operation in the drawing area; In response to a triggering operation on the image storage control, switching from the image drawing interface to the character creation interface, and displaying the reference image data in the character creation interface; The reference picture data and the text data are determined as object input data.

15. The method according to claim 10, characterized in that The character creation interface is displayed in response to a trigger operation on a first image generation control; the first image generation control is an image generation control in the character creation interface; The displaying of the character picture in the character generation interface includes: Displaying a first display control for indicating the progress of picture generation in the character generation interface; When the generation progress of the first display control reaches a progress threshold, Z initial images corresponding to the object input data are displayed in the character generation interface; Z is a positive integer; each initial image is an image with a first quality coefficient generated based on the object input data; When a picture to be processed is determined from the Z initial pictures, a second picture generation control is displayed; In response to a triggering operation on the second image generation control, displaying a second display control for indicating image generation progress in the character generation interface; When the generation progress of the second display control reaches a progress threshold, a to-be-processed picture with a second quality coefficient is displayed in the character generation interface, and the displayed to-be-processed picture is used as the character picture; the second quality coefficient is higher than the first quality coefficient.

16. A data processing device, characterized in that: include: A data acquisition module is configured to acquire object input data including text data; the text data includes a business role style for describing a business role to be generated, and the data type of the business role style is a text type; the style category in the business role style includes one or more style categories in a style category set; the style category set includes a virtual clothing and accessories style class and a virtual character attribute style class; a style adding module, configured to obtain a data prompt template including an initial role style, add the business role style to the data prompt template, and determine the added data prompt template as text to be converted; the data type of the data prompt template is a text type; the initial role style is a default role style configured for the business role to be generated; The text-to-image conversion module is used to perform text-to-image conversion on the text to be converted based on the object input data to obtain a role image corresponding to the object input data; the role image is used to generate the business role to be generated.

17. A data processing device, characterized in that: include: Create an interface display module, used to display the character creation interface; The character creation interface includes a text input box and a picture generation control; a text data display module for displaying text data describing a business role style of a business role to be generated in response to an input operation on the text input box; the data type of the business role style is a text type; the style category in the business role style includes one or more style categories in a style category set; the style category set includes a virtual clothing and accessories style class and a virtual character attribute style class; a generation interface display module for, upon acquiring object input data including text data, responding to a triggering operation on the image generation control, switching from the character creation interface to the character generation interface; A role picture display module, configured to display a role picture in the role generation interface; the role picture is used to generate the business role to be generated; The character image is obtained after performing text-to-image conversion on the text to be converted; The text to be converted is determined by the business role style and the initial role style in the data prompt template; the data type of the data prompt template is a text type; The initial role style refers to a default role style configured for the business role to be generated.

18. A computer device, characterized in that: include: processors and memory and network interfaces; The processor is connected to the memory and the network interface, wherein the network interface is used to provide a data communication function, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1 to 15.

19. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which is suitable for being loaded and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1 to 15.

20. A computer program product, characterized in that The computer program product comprises a computer program stored in a computer-readable storage medium, wherein the computer program is suitable for being read and executed by a processor, so as to enable a computer device having the processor to perform the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Automated creation and design of presentation charts

    US11200715B1

  • Generating descriptive text for images

    US20150161086A1