A method, apparatus, device, medium, and program product for cartoon digital human modal interaction.

By acquiring user multimodal data and matching the skeletal deformation rules and emotional characteristics of cartoon digital humans, the problem of the contradiction between realistic models and cartoon expression features was solved, realizing the natural fluency and style adaptability of cartoon digital human modal interaction and improving the interactive experience.

CN122134885APending Publication Date: 2026-06-02SHANGHAI INVESTIGATION DESIGN & RES INST CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI INVESTIGATION DESIGN & RES INST CO LTD
Filing Date
2026-02-26
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing technologies, 3D motion models of realistic human designs are directly applied to the digital human interaction of cartoon characters, resulting in stiff and disjointed movements of the cartoon characters, which seriously weakens the expressive effect of the cartoon style and lacks a natural and anthropomorphic feel.

Method used

By acquiring multimodal input data from users, the system matches the skeletal deformation rules of the target cartoon digital human from a pre-built style feature library based on style information. It then inputs emotional feature information into the action mapping model and outputs target action information, ensuring accurate correlation between actions and emotions and adaptation to skeletal movement patterns.

Benefits of technology

It achieves natural and smooth animation and emotional consistency in cartoon digital human movements, enhances the expressive effect of cartoon style, and improves the natural anthropomorphic feel and overall adaptability of interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134885A_ABST
    Figure CN122134885A_ABST
Patent Text Reader

Abstract

This invention relates to the field of artificial intelligence technology and discloses a method, apparatus, device, medium, and program product for modal interaction of a cartoon digital human. The method includes: acquiring multimodal data input by a user, the multimodal data including style information, text information, and / or voice information of the target cartoon digital human; determining the target skeletal deformation rules of the target cartoon digital human based on the style information in a pre-constructed style feature library; identifying the emotional feature information of the target cartoon digital human based on the text information and / or voice information; inputting the emotional feature information into a pre-constructed motion mapping model, so that the motion mapping model outputs the target motion information of the target cartoon digital human, the motion mapping model being used to characterize the correlation between different emotional feature information and motion information; and outputting the interaction result of the target cartoon digital human based on the target skeletal deformation rules and the target motion information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to a cartoon digital human modal interaction method, apparatus, device, medium, and program product. Background Technology

[0002] Digital human interaction refers to the anthropomorphic intelligent human-computer interaction between users and digital humans through multimodal means, with the digital humans providing real-time feedback through language and body expressions. The demand for this form of interaction is increasing in the field of cartoon characters, but significant problems with motion expression adaptation exist in practical implementation. In related technologies, 3D motion models are mostly designed based on real people. Their skeletal structure, movement patterns, and facial expressions are developed around the logic of real human body movements and facial expression characteristics, conforming to the natural movement presentation needs of realistic characters. However, the core of cartoon character interaction lies in exaggerated facial expressions and elastic body movements. These exclusive expressive features are fundamentally contradictory to the design logic of realistic motion models. Realistic models cannot support the natural display of exaggerated expressions such as large eyebrow raises and wide smiles in cartoon characters, nor can they achieve the elastic body movement effects such as limb stretching during jumps and elastic transitions between movements. Directly applying these 3D motion models designed for realistic characters to the digital human interaction of cartoon characters results in stiff and disjointed movements, severely weakening the expressive effect of the cartoon style and making the overall interactive experience lack a natural anthropomorphic feel. Summary of the Invention

[0003] This invention provides a method, device, equipment, medium, and program product for cartoon digital human modal interaction, in order to solve the problem in related technologies where 3D motion models of realistic human designs are directly applied to the digital human interaction of cartoon characters, resulting in stiff and disjointed movements of the cartoon characters, which seriously weakens the expressive effect of the cartoon style.

[0004] In a first aspect, the present invention provides a method for modal interaction of a cartoon digital human. The method includes: acquiring multimodal data input by a user, the multimodal data including style information, text information, and / or voice information of the target cartoon digital human; determining target skeletal deformation rules for the target cartoon digital human based on the style information in a pre-constructed style feature library, the style feature library containing skeletal deformation rules corresponding to different styles of cartoon digital humans; identifying emotional feature information of the target cartoon digital human based on the text information and / or voice information; inputting the emotional feature information into a pre-constructed motion mapping model, so that the motion mapping model outputs target motion information of the target cartoon digital human, the motion mapping model being used to characterize the correlation between different emotional feature information and motion information; and outputting the interaction result of the target cartoon digital human based on the target skeletal deformation rules and the target motion information.

[0005] The cartoon digital human modal interaction method provided by this invention first acquires user multimodal input data containing style information, and then accurately matches the skeletal deformation rules corresponding to the target cartoon digital human from a pre-built style feature library based on the style information. This method abandons the realistic skeletal motion rules commonly used in related technologies, and specifically creates skeletal deformation criteria that fit the unique expressive needs of different cartoon styles. It adapts to the core requirements of exaggerated facial expressions and flexible limb movements of cartoon characters from the underlying skeletal design level, fundamentally solving the essential problem of the contradiction between realistic models and cartoon expressive feature design logic. Furthermore, by recognizing emotional feature information in text or speech and inputting it into a motion mapping model to output matching target motion information, it achieves precise correlation between emotional features and motion information. This technology integrates the target skeletal deformation rules with the target action information driven by emotions, allowing the animation output to match the real-time emotional state of the cartoon digitizer. Further integration of these rules with the target action information outputs the final interactive result. This ensures that the cartoon digitizer's movements not only follow its unique skeletal movement patterns but also closely match its emotional state. This effectively avoids the stiffness and disjointed transitions found in other technologies, while also enhancing the expressive effect of the cartoon style. Furthermore, it ensures that the cartoon digitizer's interactive performance combines style adaptability with emotional resonance, making the overall interactive result more consistent with the expressive characteristics of the cartoon character. This addresses the shortcomings of other technologies that lack a natural, anthropomorphic feel in the interactive experience of cartoon digitizers, achieving natural and smooth modal interaction and significantly improving the overall interactive experience.

[0006] In one optional implementation, the style feature library is constructed through the following steps: acquiring information on cartoon characters of different styles; extracting key parameters from the information on cartoon characters of each style to determine the initial key parameters for the corresponding style, wherein the key parameters are used to characterize the skeletal proportions and joint range of motion of the cartoon character; acquiring motion simulation data of cartoon characters of each style in motion capture experiments; determining the true key parameters of the corresponding style based on the motion simulation data of cartoon characters of each style; optimizing the initial key parameters using the true key parameters of cartoon characters of each style to obtain the optimized key parameters of cartoon characters of each style; determining the skeletal deformation rules of the corresponding style of cartoon digitizer based on the optimized key parameters of cartoon characters of each style; and associating the style information and skeletal deformation information of the cartoon digitizer of each style to obtain the style feature library.

[0007] The method provided in this optional implementation decomposes core initial parameters such as skeletal proportions and joint range of motion based on information about cartoon characters of different styles, and optimizes these parameters by combining real motion acquisition test data. This allows the skeletal deformation rules to fit the actual motion expression needs of each cartoon style, accurately supporting the presentation of exaggerated expressions and elastic limb movements of cartoon characters. At the same time, by linking style information with skeletal deformation information to build a database, the skeletal deformation rules of the target cartoon digital human can be quickly matched, ensuring the adaptability and scientific nature of the skeletal deformation rules to the cartoon style. This fundamentally solves the problem of stiff and disjointed movements in cartoon digital humans, allowing the movement performance of cartoon digital humans to fit their own style characteristics, enhancing the expression effect of the cartoon style, and laying a core foundation for the natural and smooth modal interaction of subsequent cartoon digital humans.

[0008] In one optional implementation, the action mapping model is constructed through the following steps: obtaining emotional feature information and action combination information of multiple cartoon animation emotion samples, wherein the cartoon animation emotion samples are determined by emotional expression segments in cartoon works; extracting features from each action combination information to obtain corresponding action information; associating the emotional feature information of each cartoon animation emotion sample with the corresponding action information to obtain an associated dataset; and training a preset deep learning hybrid model using the associated dataset until the model accuracy reaches the preset requirements to obtain the action mapping model.

[0009] The method provided in this optional implementation involves extracting features and associating emotion and action information to form a standardized associated dataset. This dataset is then used to train a deep learning hybrid model to a preset accuracy. This allows the model to accurately establish the correspondence between cartoon emotion features and action information, ensuring that the output target action is highly compatible with the cartoon digitizer's emotion and avoiding the awkwardness of a disconnect between emotion and action. At the same time, the model, trained based on cartoon-specific samples, outputs actions that also match the expressive characteristics of the cartoon character. This, in conjunction with skeletal deformation rules, supports the natural emotional action presentation of the cartoon digitizer, further enhancing the anthropomorphism and fluency of the cartoon digitizer's modal interaction.

[0010] In one optional implementation, the step of inputting emotional feature information into a pre-constructed motion mapping model so that the motion mapping model outputs target motion information of the target cartoon digital human includes: obtaining target application scenario information of the target cartoon digital human; determining the motion mapping model of the target application scenario based on the application scenario information, wherein different application scenarios correspond to different motion mapping models; and inputting emotional feature information into the motion mapping model of the target application scenario so that the motion mapping model outputs target motion information of the target cartoon digital human.

[0011] The method provided by this optional implementation method, by combining target application scenario information with a dedicated motion mapping model, overcomes the limitation of a single model adapting to all scenarios. It allows the output of the motion mapping model to take into account both the emotional characteristics of the cartoon digital human and the expressive needs of specific application scenarios, avoiding the problem of incongruous expression of the same emotion in different scenarios. The output target action not only fits the real-time emotion of the cartoon digital human, but also accurately adapts to the interactive characteristics of different scenarios. After coordinating with the target skeletal deformation rules, the action presentation of the cartoon digital human is more in line with the actual needs of the scenario and the cartoon style itself, further improving the scenario adaptability and overall naturalness of the cartoon digital human modal interaction, and making the interactive performance more in line with the actual use needs in different scenarios.

[0012] In one optional implementation, the step of outputting the interaction result of the target cartoon digital human based on the target skeletal deformation rules and target action information includes: extracting key information based on text information, and determining user intent information based on the key information; inputting style information, emotional feature information, target application scenario information, and user intent information into a pre-constructed large language model so that the large language model outputs dialogue information and dialogue action information; determining the rendering information of the target cartoon digital human based on the emotional feature information, target skeletal deformation rules, target application scenario information, dialogue action information, and target action information; optimizing the target action information and dialogue action information based on the emotional feature information, target application scenario information, user intent information, and dialogue information to determine the contextual action sequence; and outputting the interaction result of the target cartoon digital human based on the dialogue information, target skeletal deformation rules, target action information, and contextual action information.

[0013] The method provided by this optional implementation first clarifies the user's intent through key text information, then synchronously outputs adapted dialogue information and dialogue action information based on a large language model to achieve initial linkage between language and action. Next, it integrates multi-dimensional information to determine rendering information, optimizes actions, and forms a contextual action sequence that fits the interactive scenario. This ensures that the action expression not only follows the skeletal deformation rules of the cartoon style but also highly adapts to the dialogue content, user intent, and application scenario. At the same time, it solves the problems of disconnect between action and dialogue and lack of contextual adaptability of action sequences through multi-element collaborative optimization, making the cartoon digital human's action presentation more logical and fluid. Ultimately, it achieves comprehensive collaborative output of dialogue, action, and rendering, making the interaction result have style adaptability, emotional fit, scene matching, and intent responsiveness, further enhancing the anthropomorphic effect and overall immersiveness of the cartoon digital human modal interaction.

[0014] In one optional implementation, the large language model is constructed through the following steps: a three-level specialized corpus is constructed, wherein the first level of the three-level specialized corpus contains interactive corpora corresponding to different styles, the second level contains interactive corpora under different emotions, and the third level includes interactive corpora under different application scenarios and user intentions, as well as dialogue action information; the initial large model is fine-tuned based on the three-level specialized corpus to obtain the large language model.

[0015] The method provided in this optional implementation fine-tunes the initial large model based on a three-level specialized corpus of style, emotion, and scene-intent-dialogue action. This allows the large language model to break free from the adaptation limitations of general models. Its output dialogue information can accurately match the style, emotion, and specific application scenario of the cartoon digital human. It can also simultaneously output dialogue action information that matches the user's intent, realizing the linkage generation of dialogue content and action commands of the cartoon digital human. This avoids the problem of dialogue and action being disconnected. At the same time, it allows the language and action commands output by the model to deeply collaborate with the output of the style feature library and the action mapping model. This lays the core language and action command foundation for the subsequent multi-dimensional information fusion to generate natural interactive results, and greatly improves the adaptability of the large model to the multimodal interaction of the cartoon digital human.

[0016] Secondly, the present invention provides a cartoon digital human modal interaction device, the device comprising: an acquisition module for acquiring multimodal data input by a user, the multimodal data including style information, text information, and / or voice information of a target cartoon digital human; a first determination module for determining target skeletal deformation rules of the target cartoon digital human based on the style information in a pre-constructed style feature library, the style feature library containing skeletal deformation rules corresponding to different styles of cartoon digital humans; a recognition module for recognizing emotional feature information of the target cartoon digital human based on the text information and / or voice information; a second determination module for inputting the emotional feature information into a pre-constructed motion mapping model, so that the motion mapping model outputs target motion information of the target cartoon digital human, the motion mapping model being used to characterize the correlation between different emotional feature information and motion information; and an output module for outputting the interaction result of the target cartoon digital human based on the target skeletal deformation rules and the target motion information.

[0017] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the cartoon digital human modal interaction method of the first aspect or any corresponding embodiment described above.

[0018] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the cartoon digital human modal interaction method of the first aspect or any corresponding embodiment described above.

[0019] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to execute the cartoon digital human modal interaction method described in the first aspect or any corresponding embodiment thereof. Attached Figure Description

[0020] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of an application scenario according to an embodiment of the present invention; Figure 2 This is a schematic flowchart of the first type of cartoon digital human modal interaction method according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the second type of cartoon digital human modal interaction method according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the third process of the cartoon digital human modal interaction method according to an embodiment of the present invention; Figure 5 This is a flowchart illustrating a specific example of the cartoon digital human modal interaction method in the embodiments of this application; Figure 6 This is a flowchart illustrating yet another specific example of the cartoon digital human modal interaction method in the embodiments of this application; Figure 7 This is a structural block diagram of a cartoon digital human modal interaction device according to an embodiment of the present invention; Figure 8 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.

[0024] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0025] As an optional application scenario of this invention, considering the specific application environment architecture or specific hardware architecture upon which the execution of the cartoon digital human modal interaction method depends, the specific application environment architecture or specific hardware architecture is described herein. For example... Figure 1 As shown, the architecture system may include at least one terminal device and at least one server. Figure 1 The system is illustrated in the example, which includes a computer 101, a mobile terminal 102, and a server 103, and the terminal devices such as the computer 101 and the mobile terminal 102 are connected to the server 103 through a network 110.

[0026] Specifically, the terminal device can be a smartphone, tablet, laptop, PDA, desktop computer, game console, smart TV, smart wearable device, in-vehicle terminal, VR (Virtual Reality) device, AR (Augmented Reality) device, etc. Server 103 can be a standalone physical server, a server cluster, a distributed system, or a cloud server providing cloud services. Network 110 can be a wired or wireless network, examples of which include, but are not limited to, the Internet, corporate intranet, local area network, wide area network, mobile communication network, and combinations thereof.

[0027] In related technologies, 3D motion models are mostly designed based on real people. Their skeletal structure, movement patterns, and facial expressions are all developed around the logic of real human body movements and facial expression characteristics, conforming to the natural movement presentation needs of realistic characters. However, the core of cartoon character interaction lies in exaggerated facial expressions and elastic body movements. These exclusive expressive features are fundamentally contradictory to the design logic of realistic motion models. Realistic models cannot support the natural display of exaggerated expressions such as large eyebrow raises and wide smiles in cartoon characters, nor can they achieve the elastic body movement effects such as body stretching when jumping and elastic transitions between movements. Directly applying these 3D motion models designed for realistic characters to the digital human interaction of cartoon characters will result in stiff and disjointed movements of the cartoon characters, severely weakening the expressive effect of the cartoon style and making the overall interactive experience lack a natural anthropomorphic feel.

[0028] In view of this, this application provides a cartoon digital human modal interaction method, which can be applied to a server to realize the interaction of the cartoon digital human. The method provided in this application first acquires user multimodal input data containing style information, and then accurately matches the skeletal deformation rules corresponding to the target cartoon digital human from a pre-built style feature library based on the style information. It abandons the realistic skeletal action rules commonly used in related technologies, and specifically creates skeletal deformation criteria that fit the exclusive expression needs of different cartoon styles. From the underlying skeletal design level, it adapts to the core requirements of exaggerated facial expressions and elastic limb movements of cartoon characters, fundamentally solving the essential problem of the contradiction between realistic models and cartoon expression feature design logic. Furthermore, by recognizing emotional feature information in text or speech and inputting it into an action mapping model to output matching target action information, it achieves a precise association between emotional features and action information, allowing… The motion output is aligned with the real-time emotional state of the cartoon digitizer. Further integration of the target skeletal deformation rules with the emotion-driven target motion information results in a final interactive outcome. This allows the cartoon digitizer's motion to both follow its unique skeletal movement patterns and highly adapt to its emotional state. This effectively avoids the stiffness and disjointed transitions found in related technologies, while also enhancing the expressive effect of the cartoon style. Simultaneously, the interactive performance of the cartoon digitizer possesses both style adaptability and emotional resonance, making the overall interactive result more consistent with the expressive characteristics of the cartoon character. This improves upon the lack of natural anthropomorphism in the interactive experience of cartoon digitizers in related technologies, achieving natural and smooth modal interaction for the cartoon digitizer and significantly enhancing the overall interactive experience.

[0029] According to an embodiment of the present invention, a cartoon digital human modal interaction method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0030] This embodiment provides a cartoon digital human modal interaction method, which can be used in the aforementioned server. Figure 2 This is a flowchart of a cartoon digital human modal interaction method according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps: Step S201: Obtain multimodal data input by the user. The multimodal data includes style information, text information, and / or voice information of the target cartoon digital human.

[0031] For example, style information is used to characterize and define the unique action and visual expression style of the target cartoon digital human, such as Q version, Chinese style cartoon, Disney cartoon, etc.; the target cartoon digital human is a specific cartoon image designated by the user in this interaction that needs to complete anthropomorphic interactive feedback, and all subsequent processing such as skeletal deformation, action generation, and dialogue output revolves around this image; text information is the interactive content entered by the user in text form, such as questions, dialogues, and instructions, which is the core data source for identifying emotional features and extracting user intentions; voice information: interactive content entered by the user in voice form, which can extract emotional features from dimensions such as tone, intonation, and speech rate, and can also help determine user intentions, and is an important supplement to text information. "And / or" means that the user can enter text alone, enter voice alone, or enter both at the same time.

[0032] Step S202: Determine the target skeletal deformation rules of the target cartoon digital human based on the style information in the pre-built style feature library. The style feature library contains skeletal deformation rules corresponding to different styles of cartoon digital humans.

[0033] For example, the style feature library is a pre-built structured database that stores a one-to-one correspondence between different cartoon styles and their respective exclusive skeletal deformation rules. It serves as the fundamental data carrier for quickly retrieving style adaptation rules. Skeletal deformation rules are the core rules used to standardize the skeletal proportions, joint range of motion, and limb movement patterns of a specific style of cartoon digitizer, determining key expressive features such as the range of motion, limb elasticity, and facial exaggeration that the cartoon character of that style can achieve. The target skeletal deformation rule refers to the skeletal deformation rule precisely matched from the style feature library that perfectly matches the style of the target cartoon digitizer in this interaction. It is the underlying action guideline that must be followed for all subsequent action generation and execution.

[0034] Step S203: Identify the emotional characteristics of the target cartoon digital human based on text information and / or voice information.

[0035] For example, in this embodiment of the application, the input text / voice information is first preprocessed. The voice information is first converted into text using speech-to-text technology to achieve data format unification. Then, features are extracted separately. On the text side, natural language processing technology is used to extract text features such as emotional words and sentence structure and tone, and basic emotions are identified through a text emotion recognition model. On the voice side, acoustic features such as tone, speech rate, pitch, and loudness are extracted, and emotional tendencies are analyzed through a voice emotion recognition model. Finally, the recognition results of text and voice are fused to output the emotional feature information of the target cartoon digital human, which includes the emotion type and corresponding intensity.

[0036] Step S204: Input the pre-built motion mapping model based on the emotion feature information so that the motion mapping model outputs the target motion information of the target cartoon digital human. The motion mapping model is used to represent the correlation between different emotion feature information and motion information.

[0037] For example, in this embodiment of the application, emotional feature information is used as the core input and passed to the action mapping model that has been pre-trained based on cartoon-specific samples. The model has established a precise correlation between different emotional features and cartoon action information through training. The model will match and output the target action information that is suitable for the specific emotional features input.

[0038] Step S205: Output the interactive results of the target cartoon digital human based on the target skeleton deformation rules and target motion information.

[0039] For example, in this embodiment of the application, the target action information for emotion matching is adapted to the skeletal rules specific to the cartoon style and transformed into a presentable interactive effect. Specifically, firstly, the target action information output by the action mapping model is adapted and adjusted according to the target skeletal deformation rules, correcting the amplitude of the action, the range of motion of the joints, the trajectory of the limb movement, etc., so that the action perfectly matches the style characteristics of the target cartoon digital human; then, the adjusted action information is transformed into skeletal motion parameters that can drive the cartoon digital human model; finally, based on these parameters, the cartoon digital human model is driven to complete the action rendering, and the final interactive result containing the limb and facial expressions that fit the style and emotion is output, realizing a high degree of adaptation between the action performance and the cartoon character.

[0040] The cartoon digital human modal interaction method provided in this embodiment first acquires user multimodal input data containing style information, and then accurately matches the skeletal deformation rules corresponding to the target cartoon digital human from a pre-built style feature library based on the style information. It abandons the realistic skeletal motion rules commonly used in related technologies, and specifically creates skeletal deformation criteria that fit the unique expressive needs of different cartoon styles. From the underlying skeletal design level, it adapts to the core requirements of exaggerated facial expressions and elastic limb movements in cartoon characters, fundamentally solving the essential problem of the contradiction between realistic models and cartoon expressive feature design logic. Furthermore, by recognizing emotional feature information in text or speech and inputting it into a motion mapping model to output matching target motion information, it achieves precise correlation between emotional features and motion information. This technology integrates the target skeletal deformation rules with the target action information driven by emotions, allowing the animation output to match the real-time emotional state of the cartoon digitizer. Further integration of these rules with the target action information outputs the final interactive result. This ensures that the cartoon digitizer's movements not only follow its unique skeletal movement patterns but also closely match its emotional state. This effectively avoids the stiffness and disjointed transitions found in other technologies, while also enhancing the expressive effect of the cartoon style. Furthermore, it ensures that the cartoon digitizer's interactive performance combines style adaptability with emotional resonance, making the overall interactive result more consistent with the expressive characteristics of the cartoon character. This addresses the shortcomings of other technologies that lack a natural, anthropomorphic feel in the interactive experience of cartoon digitizers, achieving natural and smooth modal interaction and significantly improving the overall interactive experience.

[0041] This embodiment provides a cartoon digital human modal interaction method, which can be used in the aforementioned server. Figure 3 This is a flowchart of a cartoon digital human modal interaction method according to an embodiment of the present invention, such as... Figure 3 As shown, the process includes the following steps: Step S301: Obtain multimodal data input by the user. The multimodal data includes style information, text information, and / or voice information of the target cartoon digitizer. For details, please refer to [link to relevant documentation]. Figure 2 Step S201 of the illustrated embodiment will not be described again here.

[0042] Step S302: Determine the target skeletal deformation rules of the target cartoon digital human based on the style information in the pre-built style feature library. The style feature library contains skeletal deformation rules corresponding to different styles of cartoon digital humans.

[0043] In some alternative implementations, the style feature library is constructed through the following steps: Step a1: Obtain information on cartoon characters of different styles.

[0044] For example, exclusive character feature information of various typical cartoon styles is collected, covering different style types such as Q version, Chinese style, and American cartoon.

[0045] Step a2: Extract key parameters from the information of cartoon characters of various styles and determine the initial key parameters of the corresponding styles. The key parameters are used to characterize the skeletal proportions and joint range of motion of the cartoon characters.

[0046] For example, in the embodiments of this application, under the Q-version style, based on the core visual feature of a head-to-body ratio of 1:3, the scaling factor of the head skeleton is set to 1.5 (compared to the proportion of realistic human skeletons), and the compression factor of the torso skeleton is set to 0.8. At the same time, the range of motion of joints such as shoulders, elbows, and knees is increased by 20% to adapt to the Q-version action requirements of "large arm swings and exaggerated leg bends". Under the Disney style, following the design principle of "realistic to cartoon", the rotation angle of the shoulder and neck joints is set to be limited to ±45° (to avoid excessive twisting), and the left and right swing amplitude of the hip joints is controlled within 30°. In addition, bone movement transition parameters are added (such as a 0.2-second buffer time for the elbow to bend from a straight position to 90°) to ensure the smoothness of the movement is consistent with the style of classic Disney cartoon characters. The method of determining the bone deformation rules is as follows: first, a style feature library is constructed based on official image materials (such as animation frames and character design drawings) in Q-version and Disney styles, and key parameters such as bone proportions and joint range of motion are extracted.

[0047] Step a3: Obtain motion simulation data of cartoon characters of various styles in the motion capture experiment.

[0048] For example, in this embodiment of the application, several professional cartoon animators (each specializing in chibi and Disney styles) were invited to simulate 10 core interactive actions (welcome, explanation, happiness, anger, etc.) in the OptiTrack Prime 13 motion capture device acquisition environment. 10 sets of data were collected for each type of action, for a total of 300 sets (150 sets for chibi and 150 sets for Disney), thereby obtaining motion simulation data of cartoon characters of various styles in the motion capture experiment.

[0049] Step a4: Determine the real key parameters of the corresponding cartoon character based on the motion simulation data of each style.

[0050] For example, in the embodiments of this application, parameters such as real-time joint angles (shoulder joint rotation angle, knee joint flexion angle), skeletal motion trajectory (arm swing arc radius), and motion transition time (starting time from standing to jumping) in the motion simulation data are used to obtain the real key parameters of cartoon characters of various styles.

[0051] Step a5: Optimize the initial key parameters using the real key parameters of cartoon characters of various styles to obtain the optimized key parameters of cartoon characters of various styles.

[0052] For example, in this embodiment, an improved particle swarm optimization (PSO) algorithm is used (introducing a parameter constraint mechanism to avoid joint angles and scaling factors exceeding the reasonable range of cartoon style). A multi-dimensional loss function is constructed to quantify the deviation between the initial parameters and the true reference values. The formula is: Total loss value = α × geometric deviation loss + β × fluency loss + γ × style matching loss. Among them, "geometric deviation loss" calculates the mean square error (MSE) between the joint angles and bone proportions generated by the initial parameters and the true values ​​of motion capture; "fluency loss" calculates the variance of the joint angular velocity change of adjacent frames of the action generated by the initial parameters (the larger the variance, the more choppy the action); "style matching loss" extracts the style features of the action generated by the initial parameters through a convolutional neural network (CNN) and calculates the cosine distance with the standard features of the style feature library (the larger the distance, the greater the style deviation); the weight allocation is differentiated: Q version style α=0.4, β=0.3, γ=0.3 (prioritizing geometric exaggeration), Disney style α=0.3, β=0.4, γ=0.3 (prioritizing motion fluency).

[0053] Initialize a particle swarm of 50 (each group of particles corresponds to a set of initial parameter combinations), and set the number of iterations to 100. In each iteration, calculate the total loss value of each group of particles, and update the "individual optimal solution" (the parameter corresponding to the lowest historical loss of a single particle) and the "global optimal solution" (the parameter corresponding to the lowest loss of the entire swarm). Introduce a "parameter constraint factor": if the particle parameters exceed the reasonable range of the cartoon style (such as the scaling factor of the Q version head > 2.0), add an extra penalty of 0.5 to its loss value. When the iteration reaches 100 rounds or the total loss value of the global optimal solution fluctuates < 0.01 for 10 consecutive rounds, stop the iteration and output the "optimized parameters".

[0054] Finally, through a survey of 1,000 target users (including children and cartoon enthusiasts), the parameter combinations with a user approval rate of ≥90% were selected, and the key parameters for optimizing various styles of cartoon characters were obtained.

[0055] Step a6: Determine the skeletal deformation rules of the corresponding style of cartoon digital human based on the optimized key parameters of each style of cartoon character.

[0056] For example, in this embodiment of the application, concrete quantitative parameters related to bones and joints are transformed into standardized skeletal deformation rules that can directly guide the motion presentation of cartoon digital humans. These optimized key parameters are the core quantitative basis for the exclusive expression of each cartoon style, covering core indicators such as the bone proportions, joint range of motion, and elastic amplitude of limb movements that are suitable for the style. Based on these precise parameters, specific norms and guidelines that the skeleton of the corresponding style of cartoon digital human must follow during movement and deformation are formulated, including the basic deformation range of bones, the activity threshold of joints, and the elasticity law of limb stretching / bending, etc., so that abstract parameters are transformed into underlying motion rules that can be implemented and executed. The resulting skeletal deformation rules will also become the core basis for the generation and execution of all subsequent movements of the cartoon digital human of that style, ensuring that its skeletal movement and limb deformation always conform to its own style characteristics.

[0057] Step a7: Associate the style information of each style of cartoon digital human with the skeletal deformation information to obtain the style feature library.

[0058] For example, the embodiments of this application do not limit the specific association method, as long as it is reasonable.

[0059] Step S303: Identify the emotional characteristics of the target cartoon digital human based on text and / or voice information. For details, please refer to [link to relevant documentation]. Figure 2 Step S203 of the illustrated embodiment will not be described again here.

[0060] Step S304: Input the pre-built motion mapping model based on the emotion feature information so that the motion mapping model outputs the target motion information of the target cartoon digital human. The motion mapping model is used to represent the relationship between different emotion feature information and motion information.

[0061] Specifically, step S304 includes: Step S3041: Obtain the target application scenario information of the target cartoon digital human.

[0062] For example, in the embodiments of this application, the current interaction scenario type is actively acquired or automatically identified, such as children's education, intelligent customer service, knowledge explanation or entertainment interaction.

[0063] Step S3042: Determine the action mapping model for the target application scenario based on the application scenario information. Different application scenarios correspond to different action mapping models.

[0064] For example, based on the identified target application scenario information, the specific model corresponding to the scenario is accurately retrieved from multiple pre-built action mapping models. Because the action expression norms and interaction characteristics of different application scenarios are different, the action mapping models for each scenario are trained and constructed specifically, which allows the action information output after inputting emotional features to highly match the expression needs of the scenario.

[0065] Step S3043: Input the emotion feature information into the action mapping model of the target application scenario so that the action mapping model outputs the target action information of the target cartoon digital human.

[0066] In some alternative implementations, the action mapping model is constructed through the following steps: Step b1: Obtain emotional feature information and action combination information of multiple cartoon animation emotional samples. The cartoon animation emotional samples are determined by emotional expression fragments in cartoon works.

[0067] For example, emotional feature information includes emotion category and emotion intensity, and emotions are comprehensively judged through multimodal information such as voice tone and text keywords. In this embodiment, 100,000 frames containing emotional expressions are extracted from classic cartoon works, and the emotion type (happy, angry, depressed, etc.) and intensity (1-5 levels) are labeled. The emotion type and intensity are converted into emotional feature vectors (based on a psychological emotion dimension model, such as the three dimensions of pleasure, arousal, and dominance), to obtain emotional feature information.

[0068] Step b2: Extract features from each action combination to obtain the corresponding action information.

[0069] For example, feature vectors are extracted for each action combination, including facial action units (such as the angle of the corners of the mouth and the tilt of the eyebrows), limb motion parameters (such as the amplitude of arm swing and the angle of the torso leaning forward), and action rhythm (such as the rate of change between frames).

[0070] Step b3 involves associating the emotional feature information of each cartoon animation emotion sample with the corresponding action information to obtain an associated dataset.

[0071] For example, the embodiments of this application do not specifically limit the specific association method, as long as it is reasonable.

[0072] Step b4: Train the preset deep learning hybrid model using the associated dataset until the model accuracy reaches the preset requirements, and obtain the action mapping model.

[0073] For example, in this embodiment of the application, a bidirectional LSTM-Transformer hybrid model is used to map emotions to action combinations: 1. Input layer: receives emotion feature vectors (including type and intensity) and scene labels (such as knowledge explanation, entertainment interaction); 2. Encoding layer: captures the dynamic changes in emotion intensity through LSTM (such as the process features of rising from "Happy Level 2" to "Level 4"); 3. Attention layer: the Transformer module assigns weights to scene labels (such as increasing the attention weight of "exaggerated action" by 30% in entertainment scenes) to achieve scene adaptation; 4. Output layer: outputs the feature vector of action combinations, performs cosine similarity matching with action features in the database, and selects the Top 3 candidate action combinations; the training process adopts "contrast loss + fluency constraint": contrast loss minimizes the feature deviation between predicted action and labeled action, and fluency constraint ensures that the inter-frame transition of action combinations is natural (such as the transition from "standing" to "jumping" must include the transition action of "bending knees to build up strength").

[0074] During system operation, the real-time output of emotional feature vectors (including intensity) is input into the trained model to quickly match the optimal action combination and dynamically adjust according to the following rules: 1. When the emotional intensity changes abruptly (e.g., from "calm level 1" to "excitement level 5"), an "excitement transition action" (e.g., eyes widening for 0.5 seconds) is automatically inserted to avoid abrupt action transitions; 2. Under the same emotional intensity, fine-tuning is performed based on keywords in the dialogue content (e.g., when "happy level 3" and the dialogue contains the keyword "gift", the "holding an object with both hands" action is matched first); 3. Combining user interaction feedback (e.g., if the user stares at a certain action combination for more than 2 seconds three times in a row), the matching weight of that combination is increased by 15% to achieve personalized adaptation.

[0075] Step S305: Output the interaction results of the target cartoon digital human based on the target skeleton deformation rules and target motion information. For details, please refer to [link to details]. Figure 2 Step S205 of the illustrated embodiment will not be described again here.

[0076] This embodiment provides a cartoon digital human modal interaction method, which can be used in the aforementioned server. Figure 4 This is a flowchart of a cartoon digital human modal interaction method according to an embodiment of the present invention, such as... Figure 4 As shown, the process includes the following steps: Step S401: Obtain multimodal data input by the user. The multimodal data includes style information, text information, and / or voice information of the target cartoon digital human. For details, please refer to [link to relevant documentation]. Figure 3 Step S301 of the illustrated embodiment will not be described again here.

[0077] Step S402: Based on style information, determine the target skeletal deformation rules for the target cartoon digitizer from a pre-built style feature library. The style feature library contains skeletal deformation rules corresponding to different styles of cartoon digitizers. For details, please refer to [link to details]. Figure 3 Step S302 of the illustrated embodiment will not be described again here.

[0078] Step S403: Identify the emotional characteristics of the target cartoon digital human based on text and / or voice information. For details, please refer to [link to relevant documentation]. Figure 3 Step S303 of the illustrated embodiment will not be described again here.

[0079] Step S404: Input the pre-built motion mapping model based on the emotion feature information, so that the motion mapping model outputs the target motion information of the target cartoon digital human. The motion mapping model is used to represent the correlation between different emotion feature information and motion information. For details, please refer to... Figure 3 Step S304 of the illustrated embodiment will not be described again here.

[0080] Step S405: Output the interactive results of the target cartoon digital human based on the target skeleton deformation rules and target motion information.

[0081] Specifically, step S405 above includes: Step S4051: Extract key information based on text information, and determine user intent information based on key information.

[0082] For example, in this embodiment of the application, key information includes keywords and sentence structure. Keywords (such as "how to do it", "great", "don't") and sentence structure (interrogative sentences correspond to "inquiry / knowledge seeking intent", exclamatory sentences correspond to "appreciation / feedback intent") of the user's input text are extracted to determine user intent information.

[0083] Step S4052: Input style information, emotional feature information, target application scenario information and user intent information into the pre-built large language model so that the large language model outputs dialogue information and dialogue action information.

[0084] For example, in the embodiments of this application, style, emotion, application scenario, and user intention are considered. Figure 4 The core interactive information input is a large language model specifically built for cartoon digital humans. The model relies on pre-trained multi-level specialized corpora to output dialogue information that not only fits the style and real-time emotions of the target cartoon digital human, but also adapts to the application scenario and accurately responds to the user's intentions. At the same time, it generates dialogue action information that matches the dialogue content, realizing the linkage and adaptation between dialogue content and action commands.

[0085] In some alternative implementations, the large language model is constructed through the following steps: Step c1: Construct a three-tiered specialized corpus. The first tier of the three-tiered specialized corpus contains interactive corpora corresponding to different styles. The second tier contains interactive corpora under different emotions. The third tier includes interactive corpora under different application scenarios and user intentions, as well as dialogue action information.

[0086] For example, in this embodiment of the application, complete dialogues from multiple classic cartoon animations (500,000 sentences in total) and interactive dialogues from children's educational apps (300,000 sentences) are crawled and labeled with "style tags" (such as Q-version cute style, Disney humor) to obtain a basic cartoon corpus layer (first level). For eight emotions such as "happy, depressed, and angry" and their intensities of 1-5, matching dialogue content (100,000 sentences in total) is manually written. For example, "depressed level 2" corresponds to "Don't worry, let's try again~", and "excited level 4" corresponds to "Yay! This answer is perfect!". Simultaneously, "language parameters" (speech rate coefficient, proportion of interjections, sentence length) are labeled to obtain an emotion-language association layer, i.e., the second level. Real interactive data (200,000 records) from five target scenarios, including intelligent customer service and children's education, are collected and labeled with "user intent" (such as asking questions, praising, and refusing) and "cartoon action association items" (such as the intent "praise" associated with the action "clapping") to obtain a scenario-intent corpus layer, i.e., the third level.

[0087] A "much-to-action mapping matrix" is constructed to transform user intentions into cartoon-style action combinations, improving action matching accuracy by 40%. Keywords (such as "how to do it," "great," "no") and sentence structures (interrogative sentences correspond to "inquiry / knowledge-seeking intentions," and exclamatory sentences correspond to "appreciation / feedback intentions") are extracted from user input text. A text classification model adapted to cartoon scenes is used to identify 10 core intentions (knowledge-seeking, inquiry, appreciation, refusal, etc.). Cartoon interaction-specific corpora are incorporated into the model training to ensure that intention judgments align with the needs of cartoon scenes. Fine-tuning is performed based on current scene labels (e.g., "too difficult" is judged as "helping intention" in a knowledge explanation scene, and as "giving up intention" in an entertainment scene). Text without clear semantics (such as meaningless interjections) triggers default action combinations (e.g., "tilting head + blinking").

[0088] Step c1 involves fine-tuning the initial large model based on a three-level specialized corpus to obtain a large language model.

[0089] For example, the Llama-2-7B open-source large model is chosen as the basic framework (rather than a large parameter model such as GPT-3.5) because: the 7B parameter count allows for real-time inference on ordinary GPUs (such as RTX4090), meeting the low latency requirements of interactive systems; at the same time, "model distillation" preserves general dialogue capabilities while removing redundant parameters unrelated to cartoon scenes (such as domain knowledge modules), resulting in a 30% reduction in model size. A three-stage fine-tuning strategy is designed to overcome the limitations of traditional "single-task fine-tuning" that only optimizes dialogue logic: Phase 1: Style Transfer Fine-tuning. Input the "Basic Cartoon Corpus" data and optimize the model using the "Style Loss Function" (calculating the cosine distance between the model output and the labeled style tags) to improve the cartoon style matching accuracy of the generated dialogue to 91% (e.g., avoiding the generation of "technical terms" and "complex long sentences"). Phase Two: Fine-tuning of Emotion-Language Matching. Inputting data from the "Emotion-Language Association Layer," a "Emotion-Language Parameter Mapping Loss" is constructed: for example, "Level 2 Depression" requires "speech rate coefficient 0.85 (i.e., a 15% slowdown), proportion of interjections ≥ 20%, and sentence length ≤ 15 characters." The model output is optimized through gradient descent to ensure an 88% matching accuracy between emotion and language parameters. Phase 3: Scenario-Intent Coordination Fine-tuning. Input the "Scenario-Intent Corpus" data and add "Action-Language Coordination Constraints": If the user's intent is "asking a question" and the scenario is "knowledge explanation", the model needs to generate "guided dialogue + associated 'thinking' action instruction" (such as "This question is very interesting~ Let me think about it~"), which improves the coordination rate of dialogue and action by 40%.

[0090] Step 4: Integration of Emotion-Language Dynamic Mapping Mechanism (Achieving Real-Time Adaptation) A new "Emotion-Language Mapping Module" is added to the model output layer to convert the emotional state of the cartoon digital human into language parameters in real time. Establish a mapping table of "emotional intensity - language parameters" (e.g., "low mood level 2" → speech rate 0.85, "excitement level 4" → speech rate 1.1); introduce an "action state correction factor": if the cartoon character is performing a "slow nod" action, even if the emotion is "happy level 3", the speech rate coefficient will be reduced from 1.05 to 1.0 to avoid the disjointed feeling of "slow action but rapid speech"; support personalized adjustments by users: allow developers to modify parameters through configuration files (e.g., in children's education scenarios, the upper limit of the proportion of interjections can be set to 30%).

[0091] Step S4053: Determine the rendering information of the target cartoon digital human based on emotional feature information, target skeletal deformation rules, target application scenario information, dialogue action information, and target action information.

[0092] For example, in this embodiment of the application, the input emotional feature information, target skeletal deformation rules, and target application scenario information define the basic constraints of rendering from the perspectives of emotional fit, style adaptability, and scenario matching. The dialogue action information and target action information clarify the core basis for the two types of actions that the digital human needs to present. Based on the comprehensive requirements of these elements, various visual rendering-related parameters for cartoon digital human interaction are determined, including the visual details of action presentation, the rendering range of facial expressions, and the visual adaptation standards of limb movements. This ensures that the rendering information not only follows the underlying rules of skeletal deformation in the cartoon style, but also matches real-time emotions and application scenarios, and is highly coordinated with dialogue actions and target actions.

[0093] Step S4054: Optimize the target action information and dialogue action information based on emotional feature information, target application scenario information, user intent information and dialogue information to determine the scenario action sequence.

[0094] For example, in this embodiment of the application, based on emotions, application scenarios, user intentions and dialogue information as the core basis, the target action information output by the action mapping model and the dialogue action information output by the big language model are collaboratively optimized. The connection logic of the two types of actions is coordinated, action conflicts are eliminated, and the rhythm of actions and dialogue is matched. Finally, a coherent sequence of scenario actions that fits the current complete interactive scenario is formed, so that the action presentation is highly consistent with emotions, scenarios and user intentions, and is synchronously adapted with dialogue content, ensuring the logic and fluency of the cartoon digital human's action expression.

[0095] Step S4055: Output the interaction results of the target cartoon digital human based on dialogue information, target skeleton deformation rules, target action information, and contextual action information.

[0096] For example, in the embodiments of this application, the target skeleton deformation rules are used as the underlying constraints to deeply integrate dialogue information with target action information and contextual action information, so that the cartoon digital human can synchronously present a coherent sequence of contextual actions and corresponding dialogue content that matches the emotions, scenes and user intentions according to the skeleton movement specifications that fit its own style. Finally, the complete interactive result of the cartoon digital human with coordinated speech and action, smooth movements and a high degree of fit with the overall interactive context is output.

[0097] The following specific embodiment will be used to illustrate the cartoon digital human modal interaction method provided in this application.

[0098] Example: This application provides a cartoon digital human modal interaction method, such as... Figure 5 As shown, the process of this method includes: Input stage: Users first input multimodal information including text, voice and other forms.

[0099] Preprocessing stage: Multimodal information enters the "Multimodal Information Cartoonization Processing Module" to complete the data format unification and feature extraction for cartoon scene adaptation.

[0100] The dual-branch processing stage involves the processed information flowing simultaneously to two modules: the "Large Language Model Cartoon Context Understanding Module" is responsible for parsing user intent and generating dialogue content and action commands adapted to the cartoon style; the "General Rendering Engine Module" prepares the basic parameters required for visual rendering.

[0101] Optimization and integration: The outputs of the above two modules enter the "storytelling dialogue strategy optimization module", which integrates dialogue and action information to generate a coherent action sequence that fits the interactive scenario.

[0102] Output stage: The final output is a complete cartoon digital human interactive result, which simultaneously presents action demonstrations and voice responses, completing the entire interactive loop.

[0103] Specifically, the embodiments of this application achieve a natural interactive experience that conforms to the cartoon style by deeply integrating a large language model with cartoon-style standardized motion resources.

[0104] (a) Multimodal information cartoonization processing module: 1. Style Dimension: Covers various cartoon styles such as Q-version and Disney style, with each style having its own independent skeletal deformation rules. Specifically, in the Q-version style, based on the core visual feature of a 1:3 head-to-body ratio, the head skeleton scaling factor is set to 1.5 (compared to the skeletal proportions of realistic characters), and the torso skeleton compression factor is set to 0.8. At the same time, the range of motion of joints such as shoulders, elbows, and knees is increased by 20% to accommodate the Q-version action requirements of "large arm swings and exaggerated leg bends". In the Disney style, following the design principle of "realistic to cartoon", the rotation angle of the shoulder and neck joints is limited to ±45° (to avoid excessive twisting), and the left and right swing amplitude of the hip joint is controlled within 30°. In addition, skeletal motion transition parameters are added (such as a 0.2-second buffer time for bending the elbow from a straight position to 90°) to ensure the smoothness of the action and the consistency of the style of classic Disney cartoon characters. The method for determining the skeletal deformation rules is as follows: First, a style feature library is constructed based on official Q-version and Disney-style character materials (such as animation frames and character design drawings), and key parameters such as bone proportions and joint range of motion are extracted; then, the initial parameters are fitted and optimized through 300 sets of cartoon motion capture experiments in corresponding styles (inviting animators to simulate movements). The specific fitting and optimization process includes three parts: 1.1 Experimental Data Acquisition: Three professional cartoon animators (each specializing in chibi and Disney styles) were invited to simulate 10 core interactive actions (welcome, explanation, happiness, anger, etc.) in the OptiTrack Prime 13 motion capture environment. Ten sets of data were collected for each action, for a total of 300 sets (150 sets for chibi and 150 sets for Disney). The collected data included real-time joint angles (shoulder joint rotation angle, knee joint flexion angle), skeletal motion trajectory (radius of arm swing arc), and action transition time (starting time from standing to jumping), which served as "realistic reference values". 1.2 Optimization Algorithm and Loss Function Design: An improved Particle Swarm Optimization (PSO) algorithm is adopted (introducing a parameter constraint mechanism to avoid joint angles and scaling factors exceeding the reasonable range of the cartoon style). A multi-dimensional loss function is constructed to quantify the deviation between the initial parameters and the true reference values. The formula is: Total loss value = α × geometric deviation loss + β × fluency loss + γ × style matching loss. Among them, "geometric deviation loss" calculates the mean square error (MSE) between the joint angles and bone proportions generated by the initial parameters and the true values ​​of motion capture; "fluency loss" calculates the variance of the joint angular velocity change of adjacent frames of the action generated by the initial parameters (the larger the variance, the more choppy the action); "style matching loss" extracts the style features of the action generated by the initial parameters through a convolutional neural network (CNN) and calculates the cosine distance with the standard features of the style feature library (the larger the distance, the greater the style deviation); weight allocation is differentiated: Q-style α=0.4, β=0.3, γ=0.3 (prioritizing geometric exaggeration), Disney style α=0.3, β=0.4, γ=0.3 (prioritizing motion fluency); 1.3 Iterative Optimization Execution: Initialize a particle swarm of 50 (each group of particles corresponds to a set of initial parameter combinations), and set the number of iterations to 100; calculate the total loss value of each group of particles in each iteration, and update the "individual optimal solution" (the parameter corresponding to the lowest historical loss of a single particle) and the "global optimal solution" (the parameter corresponding to the lowest loss of the entire swarm); introduce a "parameter constraint factor": if the particle parameters exceed the reasonable range of the cartoon style (e.g., the scaling factor of the Q version head > 2.0), add an extra penalty of 0.5 to its loss value; when the iteration reaches 100 rounds or the total loss value of the global optimal solution fluctuates < 0.01 for 10 consecutive rounds, stop the iteration and output the "optimized parameters"; Finally, through a survey of 1,000 target users (including children and cartoon enthusiasts), the parameter combinations with a user approval rate of ≥90% were selected, and the independent skeletal deformation rules for each style were finally determined.

[0105] 2. Emotion Dimension: Emotions are categorized by intensity (levels 1-5). Multimodal information, including voice tone and text keywords, is used to comprehensively determine emotions. For example, level 5 anger corresponds to actions such as clenching fists, frowning, and leaning forward. Here, "emotion" refers to the emotions of the cartoon avatar (not the user's emotions). Its core functions include: ① Action Combination Matching: Different emotion intensities correspond to differentiated action combinations (e.g., level 1 "slight happiness" triggers only a slight upturn of the lips, while level 5 "extreme happiness" triggers a complex action of jumping, raising both hands, and blinking eyes), ensuring the cartoon character's actions are consistent with its emotional state. This matching process is not based on pre-set fixed correspondences but is achieved automatically through a "data-driven dynamic association model," as detailed below: Step 1: Building an emotion-action association database: Three types of core data were collected to build a training library: 1. Cartoon animation emotion samples: 100,000 frames containing emotional expressions were extracted from classic cartoon works, and the emotion type (happy, angry, depressed, etc.), intensity (1-5 levels), and corresponding action combinations (e.g., "happy level 3" corresponds to "smiling + waving"); 2. Action feature quantification: Feature vectors were extracted for each action combination, including facial action units (e.g., the angle of the corners of the mouth and the tilt of the eyebrows), limb movement parameters (e.g., the amplitude of arm swing and the angle of the torso leaning forward), and action rhythm (e.g., the rate of change between frames); 3. Emotion feature quantification: The emotion type and intensity were converted into emotion feature vectors (based on a psychological emotion dimension model, such as the three dimensions of pleasure, arousal, and dominance).

[0106] Step 2: Training the dynamic matching model: A bidirectional LSTM-Transformer hybrid model is used to map emotions to action combinations: 1. Input layer: Receives emotion feature vectors (including type and intensity) and scene labels (such as knowledge explanation, entertainment interaction); 2. Encoding layer: Captures the dynamic changes in emotion intensity through LSTM (such as the process features of rising from "Happy Level 2" to "Level 4"); 3. Attention layer: The Transformer module assigns weights to scene labels (such as increasing the attention weight of "exaggerated actions" by 30% in entertainment scenes) to achieve scene adaptation; 4. Output layer: Outputs the feature vector of action combinations, performs cosine similarity matching with action features in the database, and selects the top 3 candidate action combinations; The training process adopts "contrast loss + fluency constraint": contrast loss minimizes the feature deviation between predicted actions and labeled actions, and fluency constraint ensures that the inter-frame transition of action combinations is natural (such as the transition from "standing" to "jumping" must include the transition action of "bending knees to build up strength").

[0107] Step 3: Real-time matching and dynamic adjustment: ① During system operation, the real-time output of emotional feature vectors (including intensity) is input into the trained model to quickly match the optimal action combination and dynamically adjust according to the following rules: 1. When the emotional intensity changes abruptly (e.g., from "calm level 1" to "excitement level 5"), an "excitement transition action" (e.g., eyes widening for 0.5 seconds) is automatically inserted to avoid abrupt action transitions; 2. Under the same emotional intensity, fine-tuning is performed based on keywords in the dialogue content (e.g., when "happy level 3" and the dialogue contains the keyword "gift", the "holding an object with both hands" action is matched first); 3. Combining user interaction feedback (e.g., if a user stares at a certain action combination for more than 2 seconds three times in a row), the matching weight of that combination is increased by 15% to achieve personalized adaptation; ② Adjusting the tone of voice: Synchronize emotional signals to the "Cartoon Context Understanding Module of the Large Language Model". For example, "low mood" (intensity level 2) corresponds to generating dialogue content with a 15% slower speech rate and softer wording (such as "It's okay with this problem, let's think about it slowly~"); "excitement" (intensity level 4) corresponds to generating dialogue content with a 10% faster speech rate and more interjections (such as "Wow! I'm super familiar with this knowledge point!"). ③ Coordinated Rendering Effects: Emotional state triggers adjustments to the special effects layer parameters of the rendering engine module (e.g., the outline color for the "angry" emotion changes to orange-red, and the brightness increases by 15%). This emotion is not simply set manually, but is automatically determined through multimodal information: for example, when the user's input voice contains a "cheerful tone" (tone frequency ≥ 300Hz) and the text contains keywords such as "great" or "so happy", the system automatically determines the cartoon digit's emotion as "happy" (intensity level 3-4); it also supports fine-tuning based on the scenario (e.g., in a task reminder scenario, even if the user inputs slightly negative keywords such as "a little annoyed", the system will control the cartoon digit's emotion intensity to level 1-2 to avoid conveying negative emotions and affecting the user experience).

[0108] 3. Scene Dimension: Divide the actions into scene-specific categories such as task reminders, knowledge explanations, and entertainment interactions.

[0109] (II) Large Language Model Cartoon Context Understanding Module: 1. Corpus Optimization: A large amount of animated dialogues and cartoon-style social media conversations were collected as dialogue corpus. After manual screening to remove content that did not fit the cartoon context, the corpus was integrated into a pre-trained model to optimize the dialogue logic. The construction process of this module did not directly use a general large language model, but achieved cartoon adaptation through "targeted corpus construction + multi-task fine-tuning + emotion-language dynamic mapping". The specific process is as follows: Step 1: Construction of Specialized Corpora (Breaking Through the Limitations of General-Purpose Corpora) A three-tiered specialized corpus is constructed to address the issue of insufficient adaptability to cartoon styles in general-purpose corpora. Basic Cartoon Corpus Layer: Crawl complete dialogues from classic cartoon animations (500,000 sentences in total) and interactive dialogues from children's educational apps (300,000 sentences), and label them with "style tags" (such as Q version cute style, Disney humor); Emotion-Language Association Layer: For 8 emotion categories such as "happy, depressed, angry" and intensity levels 1-5, manually generated matching dialogue content (100,000 sentences in total) are used. For example, "depressed level 2" corresponds to "Don't worry, let's try again~", and "excited level 4" corresponds to "Yay! This answer is perfect!". At the same time, "language parameters" (speech rate coefficient, proportion of interjections, sentence length) are marked. Scene-Intent Corpus Layer: Collect real interaction data (200,000 records) from 5 target scenarios, including intelligent customer service and children's education, and label "user intent" (such as asking questions, praising, and refusing) and "cartoon action association items" (such as the intent "praise" being associated with the action "clapping").

[0110] Step 2: Basic Model Selection and Initialization (Balancing Performance and Efficiency) The Llama-2-7B open-source large model is chosen as the basic framework (instead of a large parameter model such as GPT-3.5) because: the 7B parameter count can achieve real-time inference on ordinary GPUs (such as RTX4090), meeting the low latency requirements of interactive systems; at the same time, the general dialogue capabilities are retained through "model distillation", and redundant parameters unrelated to cartoon scenes (such as professional domain knowledge modules) are removed, and the model size is compressed by 30%.

[0111] Step 3: Multi-task fine-tuning (implementing cartoon-style core capabilities) Design a three-stage fine-tuning strategy to break through the limitations of traditional "single-task fine-tuning" which only optimizes dialogue logic: In the first stage, the data of the "basic cartoon corpus layer" is input, and the model is optimized with the "style loss function" (calculating the cosine distance between the model output and the labeled style label) to improve the cartoon style matching degree of the generated dialogue to 91% (such as avoiding the generation of "technical terms" and "complex long sentences"). In the second stage, the data from the "emotion-language association layer" is input to construct the "emotion-language parameter mapping loss": for example, "low mood level 2" needs to meet the following conditions: "speech rate coefficient 0.85 (i.e., slowed down by 15%), proportion of interjections ≥20%, sentence length ≤15 characters". The model output is optimized through gradient descent to ensure that the matching accuracy between emotion and language parameters reaches 88%. In the third stage, input the "scene-intent corpus layer" data and add "action-language coordination constraints": if the user's intent is "asking a question" and the scene is "knowledge explanation", the model needs to generate "guided dialogue + associated 'thinking' action instructions" (such as "This question is very interesting~ Let me think about it~"), which increases the coordination rate of dialogue and action by 40%.

[0112] Step 4: Integration of Emotion-Language Dynamic Mapping Mechanism (Achieving Real-Time Adaptation) A new "Emotion-Language Mapping Module" is added to the model output layer to convert the emotional state of the cartoon digital character into language parameters in real time: An "Emotion Intensity-Language Parameter" mapping table is established (e.g., "Depressed Level 2" → speech rate 0.85, "Excited Level 4" → speech rate 1.1); An "Action State Correction Factor" is introduced: If the cartoon character is performing a "slow nod," even if the emotion is "Happy Level 3," the speech rate coefficient will be reduced from 1.05 to 1.0 to avoid the disjointed feeling of "slow action but rapid speech"; Personalized adjustments are supported: Developers are allowed to modify parameters through configuration files (e.g., in children's education scenarios, the upper limit of the proportion of interjections can be set to 30%).

[0113] 2. Core improvements in model building: Improvement 1: Breaking through the limitations of general corpora, constructing a three-dimensional corpus of "emotion-style-scene". Existing technologies mostly use general dialogue corpora directly, resulting in low matching degree between cartoon style and emotion; this solution, through manual writing and scene collection, makes the corpus highly aligned with the interaction needs of cartoon digital humans, and the style adaptability is improved by 65% ​​compared with general corpora.

[0114] Improvement 2: Innovative multi-task fine-tuning framework, integrating "style + emotion + action" collaborative optimization. Traditional fine-tuning only focuses on "dialogue logic correctness". This solution adds "style loss" and "emotion-language matching loss" and incorporates "action-language collaborative constraints", enabling the model to simultaneously possess the three major capabilities of "cartoon style expression, emotion adaptation, and action collaboration", which improves the overall experience score by 2.1 points (out of 5) compared to single-task fine-tuning.

[0115] Improvement 3: Lightweight design combined with dynamic mapping balances efficiency and flexibility. Existing solutions using large parameter models often have inference latency exceeding 1 second; this solution uses a 7B model + distillation compression, controlling the inference latency to within 300ms; at the same time, through the "dynamic mapping module", language parameters can be adjusted in real time, avoiding the problem that traditional "static rule tables" cannot adapt to complex emotional changes.

[0116] 3. Intent-Action Mapping: Constructing an "intent-action mapping matrix" transforms user intents into cartoon-style action combinations, improving action matching accuracy by 40%. User intents are determined through text recognition, the specific process of which is as follows: Extract keywords (such as "how to do it", "great", "don't") and sentence structure (interrogative sentences correspond to "inquiry / knowledge-seeking intention", exclamatory sentences correspond to "appreciation / feedback intention") from user input text; Using a text classification model adapted to cartoon scenes, 10 core intents (seeking knowledge, seeking information, praising, refusing, etc.) are identified. The model is trained using a corpus specifically designed for cartoon interactions to ensure that intent determination aligns with the needs of cartoon scenes. Fine-tune the tags based on the current scenario (e.g., in a knowledge explanation scenario, "too difficult" is judged as "intention to ask for help", while in an entertainment scenario, it is judged as "intention to give up"). Text without clear semantics (such as meaningless interjections) triggers default action combinations (such as "tilt head + blink").

[0117] Core improvements: Focusing on the cartoon scene adaptability of text recognition, through training with dedicated corpora and scene fine-tuning, the problem of scene inconsistency in general text recognition is avoided, and the intent determination accuracy reaches 90%, meeting the simple and efficient needs of cartoon interaction.

[0118] 4. Supports dynamic adjustment of dialogue strategies based on cartoon character settings: During dialogue, the output of the large language model is adjusted in real time according to the cartoon character's current action state, facial expression changes, and the scene. For example, when the cartoon character is performing an "excited dancing" action, the dialogue content generated by the large language model should be more inclined to express positive and cheerful emotions to enhance the consistency between the dialogue and the character's actions.

[0119] (III) Through the rendering engine module: 1. Dual-layer rendering collaboration (1) Base layer: Skeleton-driven 3D model to ensure smooth movement; (2) Special effects layer: Real-time overlay of rounded celluloid coloring and dynamic outlines, with the thickness of the outlines being positively correlated with the "cute exaggeration coefficient"; 2. Material-Motion Synchronization: Material parameters are dynamically adjusted according to the action. For example, when performing a jump action, the reflectivity of the material surface increases by 30% to simulate the lively and dynamic visual effect of a cartoon character. The transparency of the special effects layer changes in sync with the action, enhancing the visual effect.

[0120] (iv) Optimization module for plot-based dialogue strategy.

[0121] 1. Large Model - Action Coordination: (1) State definition: Combine the dialogue context, cute emotions and action feedback. For example, in the knowledge explanation scenario, if the cartoon character performs the "thinking + nodding" action sequence after the user asks a question, and the dialogue content involves complex knowledge points, it is determined that the state needs further in-depth explanation; (2) Script Nodes: Supports preset cartoon plot branching strategies, such as triggering a "waving + head tilting" action sequence in a welcome scene. In entertainment and interactive scenes, triggering action combinations such as "laughing + clapping" based on user input.

[0122] Example 2: In this embodiment of the application, a knowledge explanation scenario is used as an example to illustrate the modal interaction method of cartoon digital human, such as... Figure 6 As shown, the steps are as follows: Input and preprocessing: After a user submits a question, the "Multimodal Information Cartoonization Processing Module" receives and processes the question information to prepare for subsequent analysis.

[0123] Emotion and Scene Analysis: The system first judges the emotion of the question and analyzes the intensity of the emotion (level 1-5); then it performs scene recognition to determine that the current interaction belongs to a specific scene such as knowledge explanation.

[0124] Intent and instruction generation: The "large language model cartoon context understanding module" performs intent analysis on the questions and transforms the user's intent into specific action combination requirements and dialogue requirements.

[0125] Parallel generation of actions and speech: The action selection library matches and outputs appropriate action combinations based on the knowledge explanation scenario and corresponding emotions. The speech generation module combines the language style of the cartoon character to generate corresponding answer speech.

[0126] Rendering and optimization integration: The rendering engine module is based on the underlying skeleton to drive the motion and adds special effects rendering; the "storytelling dialogue strategy optimization module" integrates motion and voice, and performs final optimization in combination with the plot branching strategy.

[0127] Final output: The cartoon character performs actions and plays voice simultaneously, completing a full interactive response to the user's question.

[0128] This embodiment also provides a cartoon digital human modal interaction device, which is used to implement the above embodiments and preferred embodiments, and will not be repeated as already described. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0129] This embodiment provides a cartoon digital human modal interaction device, such as Figure 7 As shown, it includes: The acquisition module 701 is used to acquire multimodal data input by the user, including style information, text information and / or voice information of the target cartoon digital human; The first determining module 702 is used to determine the target skeleton deformation rules of the target cartoon digital human based on style information in a pre-built style feature library. The style feature library contains skeleton deformation rules corresponding to different styles of cartoon digital humans. Recognition module 703 is used to recognize the emotional characteristics of the target cartoon digital human based on text information and / or voice information; The second determining module 704 is used to input a pre-built action mapping model based on emotional feature information, so that the action mapping model outputs the target action information of the target cartoon digital human. The action mapping model is used to represent the correlation between different emotional feature information and action information. Output module 705 is used to output the interactive results of the target cartoon digital human based on the target skeleton deformation rules and target motion information.

[0130] In some alternative implementations, the style feature library is constructed through the following steps: Obtain information on cartoon characters of different styles; Key parameters are extracted from the information of cartoon characters of various styles to determine the initial key parameters of the corresponding styles. The key parameters are used to characterize the skeletal proportions and joint range of motion of the cartoon characters. Acquire motion simulation data of cartoon characters of various styles in motion capture experiments; Based on motion simulation data of cartoon characters of various styles, determine the real key parameters of the corresponding style of cartoon characters; The initial key parameters were optimized using the real key parameters of cartoon characters of various styles, resulting in optimized key parameters for each style of cartoon character. Determine the skeletal deformation rules of cartoon digital humans of corresponding styles based on the optimized key parameters of cartoon characters of various styles. By associating the style information of cartoon digital humans of various styles with skeletal deformation information, a style feature library is obtained.

[0131] In some alternative implementations, the action mapping model is constructed through the following steps: The emotional characteristics and action combination information of multiple cartoon animation emotional samples are obtained. The cartoon animation emotional samples are determined by emotional expression segments in cartoon works. Feature extraction is performed on each action combination to obtain the corresponding action information; The emotional features of each cartoon animation emotion sample are associated with the corresponding action information to obtain an associated dataset; The pre-defined deep learning hybrid model is trained using the associated dataset until the model accuracy reaches the pre-defined requirements, thus obtaining the action mapping model.

[0132] In some alternative implementations, the second determining module 704 includes: The acquisition submodule is used to acquire the target application scenario information of the target cartoon digital human; The first determination submodule is used to determine the action mapping model of the target application scenario based on the application scenario information. Different application scenarios correspond to different action mapping models. The second determining submodule is used to input emotional feature information into the action mapping model of the target application scenario, so that the action mapping model outputs the target action information of the target cartoon digital human.

[0133] In some alternative implementations, the output module 705 includes: The extraction module is used to extract key information based on text information and determine user intent information based on the key information. The third determination submodule is used to input style information, emotional feature information, target application scenario information and user intent information into the pre-built large language model so that the large language model can output dialogue information and dialogue action information. The fourth determination submodule is used to determine the rendering information of the target cartoon digital human based on emotional feature information, target skeletal deformation rules, target application scenario information, dialogue action information, and target action information. The fifth determination submodule is used to optimize the target action information and dialogue action information based on emotional feature information, target application scenario information, user intent information and dialogue information to determine the scenario action sequence; The output submodule outputs the interaction results of the target cartoon digital human based on dialogue information, target skeleton deformation rules, target action information, and contextual action information.

[0134] In some alternative implementations, the large language model is constructed through the following steps: A three-tiered specialized corpus was constructed. The first tier of the corpus contains interactive corpora corresponding to different styles. The second tier contains interactive corpora under different emotions. The third tier includes interactive corpora under different application scenarios and user intentions, as well as dialogue action information. The initial large model was fine-tuned based on a three-level specialized corpus to obtain a large language model.

[0135] The cartoon digital human modal interaction device provided in this embodiment of the invention can execute the cartoon digital human modal interaction method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method execution. Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0136] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.

[0137] The following is a detailed reference. Figure 8This diagram illustrates a suitable structural schematic for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 801, which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) 802 or a program loaded from memory 808 into random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device. The processor 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0138] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0139] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a memory 808, or installed from a ROM 802. When the computer program is executed by the processor 801, it performs the functions defined in the cartoon digital human modal interaction method of the embodiments of the present invention.

[0140] Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0141] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the cartoon digital human modal interaction method shown in the above embodiments is implemented.

[0142] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0143] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for modal interaction of a cartoon digital human, characterized in that, The method includes: Acquire multimodal data input by the user, wherein the multimodal data includes style information, text information and / or voice information of the target cartoon digital human; Based on the style information, the target skeletal deformation rules of the target cartoon digital human are determined in a pre-built style feature library, which contains skeletal deformation rules corresponding to different styles of cartoon digital humans. Based on the text and / or voice information, identify the emotional characteristics of the target cartoon digital human; Based on the emotional feature information, a pre-constructed motion mapping model is input so that the motion mapping model outputs the target motion information of the target cartoon digital human. The motion mapping model is used to represent the correlation between different emotional feature information and motion information. Based on the target skeleton deformation rules and target motion information, the interactive results of the target cartoon digital human are output.

2. The method according to claim 1, characterized in that, The style feature library is constructed through the following steps: Obtain information on cartoon characters of different styles; Key parameters are extracted from the information of cartoon characters of various styles to determine the initial key parameters of the corresponding styles. The key parameters are used to characterize the skeletal proportions and joint range of motion of the cartoon characters. Acquire motion simulation data of cartoon characters of various styles in motion capture experiments; Based on the motion simulation data of cartoon characters of various styles, the true key parameters of the corresponding cartoon characters of various styles are determined; The initial key parameters are optimized using the actual key parameters of cartoon characters of various styles to obtain the optimized key parameters of cartoon characters of various styles. Determine the skeletal deformation rules of cartoon digital humans of corresponding styles based on the optimized key parameters of cartoon characters of various styles. By associating the style information of cartoon digital humans of various styles with skeletal deformation information, a style feature library is obtained.

3. The method according to claim 1 or 2, characterized in that, The action mapping model is constructed through the following steps: The emotional feature information and action combination information of multiple cartoon animation emotional samples are obtained, wherein the cartoon animation emotional samples are determined by emotional expression fragments in cartoon works; Feature extraction is performed on each action combination to obtain the corresponding action information; The emotional features of each cartoon animation emotion sample are associated with the corresponding action information to obtain an associated dataset; The preset deep learning hybrid model is trained using the associated dataset until the model accuracy reaches the preset requirement, thus obtaining the action mapping model.

4. The method according to claim 1, characterized in that, The step of inputting the emotional feature information into a pre-constructed motion mapping model, so that the motion mapping model outputs the target motion information of the target cartoon digital human, includes: Obtain target application scenario information for the target cartoon digital human; Based on the application scenario information, an action mapping model for the target application scenario is determined. Different application scenarios correspond to different action mapping models. The emotional feature information is input into the action mapping model of the target application scenario, so that the action mapping model outputs the target action information of the target cartoon digital human.

5. The method according to claim 4, characterized in that, The steps for outputting the interactive results of the target cartoon digital human based on the target skeletal deformation rules and target motion information include: Key information is extracted from the text information, and user intent information is determined based on the key information. The style information, emotional feature information, target application scenario information, and user intent information are input into a pre-built large language model so that the large language model outputs dialogue information and dialogue action information. The rendering information of the target cartoon digital human is determined based on the emotional feature information, the target skeletal deformation rules, the target application scenario information, the dialogue action information, and the target action information. Based on the emotional feature information, target application scenario information, user intent information, and dialogue information, the target action information and dialogue action information are optimized to determine the scenario action sequence; Based on the dialogue information, the target skeleton deformation rules, the target action information, and the contextual action information, the interaction results of the target cartoon digital human are output.

6. The method according to claim 5, characterized in that, The large language model is constructed through the following steps: A three-tiered specialized corpus is constructed. The first tier of the three-tiered specialized corpus contains interactive corpus corresponding to different styles. The second tier contains interactive corpus under different emotions. The third tier includes interactive corpus under different application scenarios and user intentions, as well as dialogue action information. The initial large model was fine-tuned based on a three-level specialized corpus to obtain the large language model.

7. A cartoon digital human modal interactive device, characterized in that, The device includes: The acquisition module is used to acquire multimodal data input by the user, wherein the multimodal data includes style information, text information and / or voice information of the target cartoon digital human; The first determining module is used to determine the target skeletal deformation rules of the target cartoon digital human based on the style information in a pre-built style feature library, wherein the style feature library contains skeletal deformation rules corresponding to different styles of cartoon digital humans. The recognition module is used to recognize the emotional characteristics of the target cartoon digital human based on the text information and / or voice information; The second determining module is used to input a pre-constructed action mapping model based on the emotional feature information, so that the action mapping model outputs the target action information of the target cartoon digital human. The action mapping model is used to represent the correlation between different emotional feature information and action information. The output module is used to output the interactive results of the target cartoon digital human based on the target skeleton deformation rules and target motion information.

8. An electronic device, characterized in that, include: A memory and a processor are interconnected, the memory stores computer instructions, and the processor executes the computer instructions to perform the cartoon digital human modal interaction method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the cartoon digital human modal interaction method according to any one of claims 1 to 6.

10. A computer program product, characterized in that, Includes computer instructions for causing a computer to execute the cartoon digital human modal interaction method according to any one of claims 1 to 6.