Digital human generation method and computer readable storage medium
By semantic analysis and parameter library matching of the text data entered by users, a digital person image that conforms to the characteristics of the industry is generated, and the problems of complex processes and cumbersome operations in the existing technology are solved, and efficient and high-quality digital person generation is achieved.
Patent Information
- Application Number
- CN202510555585.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-01
AI Technical Summary
The existing generative digital human production process is complex, and it is difficult to quickly summarize the prompt words for complex context understanding to assist in the generation. Users need to manually handle multiple links, the operation steps are cumbersome, which affects the production efficiency. It takes an average of 17.3 human-computer interactions to generate a live streaming digital human, which mainly consumes time to focus on the feature parameter calibration link.
By obtaining the text data input by the user, using the preset analytical module for semantic analysis, extracting industry-related features, and determining digital life generation parameters from the preset parameter library, constructing structured generation instructions, and driving the image generation module to generate digital person images that meet industry characteristics, and integrating dynamic characteristics to output the final digital person.
The operation steps are simplified, the efficiency of digital human production is improved, and the number of human-computer interactions is reduced. The generated digital humans are more in line with user needs, have higher quality and are highly scalable.
Smart Images

Figure CN120409497A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic information technology, and particularly to a method for generating a digital human and a computer-readable storage medium. Background Art
[0002] The existing production process of generative digital humans is relatively complex, and it is difficult to quickly summarize the prompt words for complex context understanding to assist in generation. Users need to manually process multiple links, and the operation steps are cumbersome, affecting the production efficiency. On average, it takes 17.3 times of human-computer interaction to generate a live-streaming digital human, and the main time-consuming is concentrated in the feature parameter calibration link. Summary of the Invention
[0003] The purpose of the present invention is to solve the above problems and provide a method for generating a digital human and a computer-readable storage medium.
[0004] The technical solution of this application is implemented as follows: The present invention provides a method for generating a digital human, including: obtaining text data input by a user; performing semantic analysis on the text data through a preset parsing module to extract industry-related features; determining corresponding digital human generation parameters from a preset parameter library according to the industry-related features; constructing a structured generation instruction based on the digital human generation parameters; driving an image generation module through the generation instruction to generate a digital human image conforming to the industry-related features; integrating the digital human image with preset dynamic characteristics and outputting a final digital human.
[0005] As a further improvement, the performing semantic analysis on the text data through a preset parsing module to extract industry-related features includes: inputting the text data into a preset semantic parsing model; identifying industry identifiers and key attributes in the text data through the semantic parsing model; generating feature tags corresponding to the industry according to the industry identifiers and key attributes; mapping the feature tags to a preset industry feature library to obtain industry-related features matching the text data; outputting the industry-related features as the basis for subsequent parameter determination.
[0006] As a further improvement, the determining corresponding digital human generation parameters from a preset parameter library according to the industry-related features includes: obtaining parameter requirements corresponding to the industry-related features; screening parameter combinations matching the parameter requirements from the preset parameter library; adjusting the parameter combinations through a preset optimization algorithm to generate optimized parameters adapted to the industry-related features; applying a preset enhancement rule to the optimized parameters to generate final digital human generation parameters; storing the digital human generation parameters as the input of subsequent generation instructions.
[0007] As a further improvement, constructing a structured generation instruction based on the digital life generation parameters includes: obtaining the core elements in the digital life generation parameters; converting the core elements into a structured description text through a preset instruction generation model; applying a preset priority sorting rule to the description text to adjust the weights of each element; optimizing the expression order of the description text according to the weights to generate a final generation instruction; and transmitting the generation instruction to an image generation module as input.
[0008] As a further improvement, driving an image generation module through the generation instruction to generate a digital human image conforming to the industry-related characteristics includes: obtaining the image generation elements in the generation instruction; parsing the image generation elements through a preset image generation model to extract corresponding visual features; generating an initial digital human image according to the visual features; applying a preset image enhancement technology to the initial digital human image to adjust the image details; and outputting a final digital human image conforming to the industry-related characteristics.
[0009] As a further improvement, integrating the digital human image with preset dynamic characteristics and outputting a final digital human includes: obtaining the static feature data of the digital human image; determining corresponding dynamic characteristic parameters according to the industry-related characteristics; fusing the dynamic characteristic parameters with the static feature data through a preset mapping model; adjusting the coordination of actions and expressions for the fused data to generate a dynamic digital human image; and outputting the dynamic digital human image as the final digital human.
[0010] As a further improvement, obtaining the text data input by the user includes: receiving the text data provided by the user through a preset input interface; performing format verification on the text data to generate a standardized input text; applying a preset semantic completion model to the input text to supplement implicit information; generating complete text data according to the supplemented input text; and transmitting the complete text data to a parsing module as input.
[0011] The present invention further provides a computer-readable storage medium storing a computer program, which can be executed by a processor of the device where the computer-readable storage medium is located to implement the above method.
[0012] The advantages or beneficial effects in the above technical solutions at least include: Compared with the prior art, the present invention can reduce the prompt-assisted generation link for complex context understanding, reduce the user's need to manually process multiple links, simplify the operation steps, and thus significantly improve the efficiency of digital human production. For example, the prior art requires an average of 17.3 human-computer interactions to generate a live streaming digital human, while the present invention can effectively reduce this number of interactions and save time.
[0013] Furthermore, the present invention performs semantic analysis on text data through a preset parsing module, can more accurately extract industry-related features, provide a more reliable basis for subsequent parameter determination, and generate a digital human image that better meets user needs.
[0014] In addition, the present invention can further screen matching parameter combinations from a preset parameter library according to industry-related features, and generate the final digital human generation parameters through an optimization algorithm and enhanced rules to ensure the adaptability of the parameters and the quality of the generated digital human.
[0015] Even further, the present invention drives an image generation module through generated instructions, parses image generation elements and extracts visual features. After generating an initial digital human image, image enhancement technology is applied to adjust image details, and a higher-quality digital human image can be output.
[0016] Finally, the method of the present invention can also be applied in various scenarios and has strong scalability. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings illustrate exemplary embodiments of the present application and, together with the description thereof, are used to explain the principles of the present application. These drawings are included to provide a further understanding of the present application and are included in this specification and form a part of this specification.
[0018] Figure 1 Shows a flowchart of the method for generating a digital human provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] Embodiments of the present application will be described in more detail below with reference to the drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not used to limit the protection scope of the present application.
[0020] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.
[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present application are for illustrative purposes only and are not used to limit the scope of these messages or information.
[0022] Please refer to Figure 1 As shown, an embodiment of the present invention provides a method for generating a digital human, including: S1, obtaining text data input by a user; S2, performing semantic analysis on the text data through a preset parsing module to extract industry-related features; S3, determining corresponding digital human generation parameters from a preset parameter library according to the industry-related features; S4, constructing a structured generation instruction based on the digital human generation parameters; S5, driving an image generation module through the generation instruction to generate a digital human image that conforms to the industry-related features; S6, integrating the digital human image with preset dynamic characteristics and outputting a final digital human.
[0023] In step S1, as a further improvement, in one embodiment, the obtaining of the text data input by the user includes: receiving the text data provided by the user through a preset input interface, and the description of the text is not limited. For example, it can be a short description of the user such as "professional lecturer in the financial industry", "professional lecturer in law", "intellectual property lecturer", etc.; performing format verification on the text data to generate a standardized input text; applying a preset semantic completion model to the input text to supplement implicit information; generating complete text data according to the supplemented input text; and transmitting the complete text data to the parsing module as input.
[0024] In step S2, as a further improvement, in one of the embodiments, the semantic analysis of the text data by a preset parsing module to extract industry-related features specifically includes: inputting the text data into a preset semantic parsing model, where the semantic parsing model is not limited and existing models can be used without limitation here, such as DeepSeek, KIMI, etc.; identifying industry identifiers and key attributes in the text data through the semantic parsing model. For example, the industry identifiers of professional lecturers in the financial industry mainly include: the image of knowledge outputters (lecturers / consultants); high information density, emphasizing "clear logic" and "distinct viewpoints"; often involving data display, financial terms, and trend analysis; the scene is biased towards the workplace, conference, or studio style. And the key attributes of professional lecturers in the financial industry mainly include: middle-aged, professional attire, and a steady tone; rational expression: partial to cool colors, controlled gestures, and an explanatory posture; background exclusivity: K-line charts, financial information screens, and office environments; copywriting strategy: question-guided hook, such as "Do you know why the market index dropped today?"; subtitle style: clear structure, keyword highlighting, and upper and lower column display for logical sorting, etc.; generating feature tags corresponding to the industry according to the industry identifiers and key attributes; mapping the feature tags to a preset industry feature library to obtain industry-related features matching the text data; outputting the industry-related features as the basis for subsequent parameter determination.
[0025] In step S3, as a further improvement, in one embodiment, the corresponding digital human generation parameters are determined from the preset parameter library according to the industry-related characteristics, including: obtaining parameter requirements corresponding to the industry-related characteristics, the parameters including style, age group, gender, regional characteristics, oral posture, screen ratio, copywriting style, character details, background and subtitle style, etc.; screening parameter combinations matching the parameter requirements from the preset parameter library, wherein the parameters in the preset parameter library include at least: 1. style, that is, the style at least includes positioning professionals, or positioning daily social interaction; 2. age, that is, the age group at least includes teenagers, young people, middle-aged people or the elderly; 3. gender, that is, the gender includes males and females; 4. area, that is, the regional characteristics at least include China, Europe, Africa, South Asia, East Asia, the Middle East, South America, or North America; 5. gesture, i.e., the oral posture at least includes focusing the eyes on the camera or relaxing the body naturally; 6. frame_ratio, i.e., the screen ratio at least includes 9:16 or 16:9; 7. text, i.e., the copywriting style includes references to the corresponding copywriting styles in the industry; 8. details, i.e., the character details at least include facial expressions, hairstyles, clothing, etc.; 9. background, i.e., the background industry includes the background that enhances the sense of the network, etc.; 10. caption, i.e., the subtitle style includes the subtitle style that matches the image; the parameter combination is adjusted through a preset optimization algorithm to generate optimized parameters that adapt to the relevant characteristics of the industry; based on the optimized parameters, the preset enhancement rules are applied to generate the final digital human generation parameters; The following is a detailed description of each enhancement rule to ensure that the generated digital human image is highly fit and expressive online: Style enhancement, which is used to enhance professionalism and rationality. For example, cool colors, straight lines, and calm light and shadow can be used to enhance the image of a professional lecturer in the financial industry. Age matching enhancement, which is used to align with industry trust expectations. For example, a middle-aged image with authentic mature features and no excessive beautification can be used to enhance the image of a professional lecturer in the financial industry. Enhanced gender expression, which is used to conform to the industry's conventional role preferences. For example, male: strong authority or female: high affinity can be adopted. This can be selected based on specific communication goals to enhance professional lecturers in the financial industry; Regional adaptation enhancement, which is used to synchronize clothing with culture / climate. For example, it can be combined with Chinese spring and autumn business attire, with more stable colors to enhance the professional lecturers in the financial industry; Gesture semantic enhancement, which is used to ensure that the delivery is clear, natural and credible. For example, a slightly sideways posture, gestures below the elbow, and eye-focused shots can be used to enhance the performance of professional lecturers in the financial industry. Enhanced frame logic, which is used in conjunction with chart / logical explanations. For example, a 16:9 landscape screen can be adopted, supporting left-right information partitioning, facilitating the display of trend charts / text and pictures side by side to enhance professional lecturers in the financial industry; Enhanced copywriting style, which is used to attract attention and convey value. For example, a thought-provoking question hook + concise structure can be adopted to enhance professional lecturers in the financial industry; Enhanced character details, which are used to build a highly internet-savvy and credible digital human. For example, a refined business hairstyle + dark suit + focused expression can be adopted to enhance professional lecturers in the financial industry; Enhanced background semantics, which are used to strengthen the immersion in the financial industry. For example, industry elements such as data charts, financial walls, and bookshelves can be included in the background to enhance professional lecturers in the financial industry; Enhanced subtitle style, which is used to improve reading efficiency and logical follow-up. For example, black text on a white background + highlighted keyword color blocks + upper and lower column layered display can be adopted to enhance professional lecturers in the financial industry. Store the digital human generation parameters as the input for subsequent generation instructions.
[0026] In step S4, as a further improvement, in one embodiment, constructing a structured generation instruction based on the digital human generation parameters includes: obtaining the core elements in the digital human generation parameters; converting the core elements into structured descriptive text through a preset instruction generation model; applying a preset priority sorting rule to the descriptive text to adjust the weights of each element; optimizing the expression order of the descriptive text according to the weights to generate the final generation instruction; transmitting the generation instruction to the image generation module as input.
[0027] In step S5, as a further improvement, in one embodiment, driving the image generation module through the generation instruction to generate a digital human image that conforms to the industry-related characteristics includes: obtaining the image generation elements in the generation instruction; parsing the image generation elements through a preset image generation model to extract the corresponding visual features; generating an initial digital human image according to the visual features; applying a preset image enhancement technique to the initial digital human image to adjust the image details; outputting the final digital human image that conforms to the industry-related characteristics.
[0028] In step S6, as a further improvement, in one of the embodiments, integrating the digital human image with the preset dynamic characteristics and outputting the final digital human includes: obtaining the static feature data of the digital human image; determining the corresponding dynamic characteristic parameters according to the industry-related characteristics; fusing the dynamic characteristic parameters with the static feature data through a preset mapping model; adjusting the coordination of actions and expressions for the fused data to generate a dynamic digital human image; and outputting the dynamic digital human image as the final digital human.
[0029] Specifically, define the feature vector extracted from the static image of the digital human as s ∈ R n , where n is the dimension of the feature, and this vector can include visual features such as color histograms, edge detection results, depth information, etc.
[0030] The dynamic characteristic parameters determined according to the industry characteristics can be represented as a vector d ∈ R m , where m is the dimension of the dynamic characteristics. These dynamic characteristic parameters can include action frequency, expression intensity, intonation change, etc.
[0031] Then fuse the static feature data and the dynamic characteristic parameters as follows: .
[0032] Where, W s ∈R k×n and W d ∈R k×m are weight matrices respectively used to adjust the influence of static features and dynamic characteristics on the fusion result, b ∈ R k is the bias vector, and f ∈ R k is the fused feature vector used to represent the digital human features combining static and dynamic information.
[0033] To generate a natural dynamic digital human image, it is necessary to ensure the coordination of actions and expressions. A time-series-based model can be used to adjust actions and expressions: .
[0034] Where, a(t) ∈ R p represents the action feature at time t; e(t) ∈ R q represents the expression feature at time t; b a ∈R p and b e ∈R q are bias vectors respectively; σ is the activation function, such as sigmoid or ReLU, used to introduce non-linearity.
[0035] In other embodiments, to ensure the coordination of actions and expressions, a coordination loss function can be further defined: 。
[0036] Among them, by minimizing this loss function, the changes in actions and expressions can be made smoother and more natural.
[0037] The final dynamic digital human image can be obtained by inputting the fused feature vector f into a generation model, for example, using a generative adversarial network (GAN) or a variational autoencoder (VAE) to generate a dynamic image sequence: 。
[0038] Where G is the generation model, such as a generator network based on CNN; I(t) is the dynamic digital human image generated at time t. Through these mathematical models, the generation from static feature data to the dynamic digital human image can be achieved, and the coordination of actions and expressions can be ensured.
[0039] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, which can be executed by a processor of the device where the computer-readable storage medium is located to implement the method described above.
[0040] The present invention can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0041] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0042] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions in the process Figure 1One process or multiple processes and / or boxes Figure 1 The functions specified in one box or multiple boxes.
[0043] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 One process or multiple processes and / or boxes Figure 1 The steps of the functions specified in one box or multiple boxes.
[0044] As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for generating a digital human, characterized in that, Including: Obtain the text data input by the user; perform semantic analysis on the text data through a preset parsing module to extract industry-related features; According to the industry-related features, determine the corresponding digital human generation parameters from a preset parameter library; based on the digital human generation parameters, construct a structured generation instruction; drive an image generation module through the generation instruction to generate a digital human image that conforms to the industry-related features; integrate the digital human image with preset dynamic characteristics and output the final digital human.
2. The method for generating a digital human according to claim 1, wherein, The performing semantic analysis on the text data through a preset parsing module to extract industry-related features includes: inputting the text data into a preset semantic parsing model; identifying industry identifiers and key attributes in the text data through the semantic parsing model; generating feature tags corresponding to the industry according to the industry identifiers and key attributes; mapping the feature tags to a preset industry feature library to obtain industry-related features matching the text data; outputting the industry-related features as the basis for subsequent parameter determination.
3. The method for generating a digital human according to claim 1, wherein The determining the corresponding digital human generation parameters from a preset parameter library according to the industry-related features includes: obtaining the parameter requirements corresponding to the industry-related features; screening parameter combinations matching the parameter requirements from the preset parameter library; adjusting the parameter combinations through a preset optimization algorithm to generate optimized parameters adapted to the industry-related features; applying a preset enhancement rule to the optimized parameters to generate the final digital human generation parameters; storing the digital human generation parameters as the input for subsequent generation instructions.
4. The method for generating a digital human according to claim 1, wherein The constructing a structured generation instruction based on the digital human generation parameters includes: obtaining the core elements in the digital human generation parameters; converting the core elements into structured description text through a preset instruction generation model; applying a preset priority sorting rule to the description text to adjust the weights of each element; optimizing the expression order of the description text according to the weights to generate the final generation instruction; transmitting the generation instruction to the image generation module as input.
5. The method for generating a digital human according to claim 1, wherein, The driving an image generation module through the generation instruction to generate a digital human image that conforms to the industry-related features includes: obtaining the image generation elements in the generation instruction; parsing the image generation elements through a preset image generation model to extract corresponding visual features; generating an initial digital human image according to the visual features; applying a preset image enhancement technology to the initial digital human image to adjust the image details; outputting the final digital human image that conforms to the industry-related features.
6. The method for generating a digital human according to claim 1, wherein The integrating the digital human image with preset dynamic characteristics and outputting the final digital human includes: obtaining the static feature data of the digital human image; determining the corresponding dynamic characteristic parameters according to the industry-related features; fusing the dynamic characteristic parameters with the static feature data through a preset mapping model; adjusting the coordination of actions and expressions for the fused data to generate a dynamic digital human image; outputting the dynamic digital human image as the final digital human.
7. The method for generating a digital human according to claim 1, wherein The obtaining of the text data input by the user includes: receiving the text data provided by the user through a preset input interface; performing format verification on the text data to generate a standardized input text; applying a preset semantic completion model to the input text to supplement implicit information; generating complete text data according to the supplemented input text; and transmitting the complete text data to the parsing module as input.
8. A computer-readable storage medium, characterized in that, A computer program is stored, and the computer program can be executed by the processor of the device where the computer-readable storage medium is located to implement the method according to any one of claims 1 to 7.