Characteristic tag-based cartoon character generation method and system
By adopting the combination method of feature recognition model, emotion recognition model and label matching model in the animation character generation technology, the problem of isolated tags and lack of fine expression of emotion in the prior art is solved, and the high accuracy and diversity of animation character generation is achieved.
Patent Information
- Application Number
- CN202510539384.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The existing anime character generation technology has problems of inconsistency caused by label isolation processing, and lacks refined control over the emotional expression of characters, resulting in the generated character image inconsistent with the expected emotions.
A cartoon character generation method based on feature tags is adopted, and a pre-trained feature recognition model, emotion recognition model and label matching model is used to extract character feature vectors and emotional enhancement information, and emotional correlation analysis and optimization are performed to generate emotionally consistent anime characters.
It significantly improves the accuracy and expressiveness of generating roles, realizes emotional consistency and diversity of emotional expressions between feature labels, and meets users' diverse needs for role images.
Smart Images

Figure CN120070675A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of virtual character generation, and in particular to a method and system for generating an animation character based on feature tags. Background Art
[0002] Anime character generation refers to the automatic or semi-automatic creation of anime characters with specific appearance, personality and characteristics through computer technology. With the development of artificial intelligence and computer graphics, this technology has evolved from early manual drawing to today's intelligent generation stage.
[0003] In the existing animation character generation technology, template matching or rule-driven methods are usually used to build character images. Such technologies often rely on a preset character component library and fixed matching rules, such as directly matching corresponding visual elements such as clothing and hairstyles through text keywords entered by users.
[0004] However, the existing anime character generation technology has the following significant defects: First, the traditional label matching mechanism often processes various feature labels in isolation, ignoring the emotional correlation between labels, and easily generates characters with inconsistent visual expressions. For example, the personality label of "elegant" may be inconsistent with the action label of "fighting posture"; secondly, most systems lack refined control over the emotional expression of the characters, resulting in the generated character image not matching the expected emotions. For example, an "angry" character may only be expressed through a simple red hue, lacking details such as micro-expressions and action tension; the above problems not only severely limit the expressiveness and personalization of the generated characters, but also make it difficult to accurately generate the anime characters required by users. Summary of the invention
[0005] In order to solve the above defects, the present application provides a method and system for generating an animated character based on feature tags.
[0006] The above-mentioned invention objective of the present application is achieved through the following technical solutions: A method for generating an animation character based on feature labels, comprising the steps of: In response to a role generation request issued by a user terminal, obtaining generation requirement information from the user terminal; Inputting the generation requirement information into a pre-trained feature recognition model, so that the feature recognition model generates a role feature vector based on the generation requirement information; Inputting the generation demand information into a pre-trained emotion recognition model so that the emotion recognition model generates emotion enhancement information; Inputting the character feature vector and the emotion enhancement information into the pre-trained label matching model, so that the label matching model matches a number of feature labels based on the pre-stored feature label system and performs emotion enhancement optimization; Generate corresponding anime characters based on the enhanced feature tags output by the tag matching model.
[0007] By adopting the above technical solution, with an emotion-enhanced tag matching model, it realizes fine control of the anime character generation process and cross-tag collaborative optimization, and has the effect of significantly improving the accuracy of the generated characters, as well as the expressiveness and personalization of the characters; specifically, by responding to the character generation request sent by the client, obtaining the generation requirement information, and inputting it into the pre-trained feature recognition model and emotion recognition model, generating a character feature vector and emotion enhancement information respectively, and then through the tag matching model, based on the pre-stored feature tag system, performing tag matching and emotion enhancement optimization, finally generating the corresponding anime character; this application extracts multi-modal features through the feature recognition model and constructs a character feature vector, combines the emotion enhancement information generated by the emotion recognition model, and uses the tag matching model to perform emotion correlation analysis and optimization on the feature tags, solves the incoordination problem caused by the isolated processing of tags in the traditional technology, and realizes the emotion consistency between the feature tags; at the same time, by introducing the emotion enhancement information, it avoids the limitation of expressing emotions only through a single visual element, can comprehensively adjust the comprehensive features of the character, accurately convey the expected emotion, and has the effect of significantly improving the accuracy of the generated character, as well as the expressiveness and personalization of the character, and meeting the diverse needs of users for the character image.
[0008] In a preferred example of this application, it can be further configured as: the feature recognition model includes an input encoding layer, a feature decoupling layer, and a vector fusion layer, and the step of inputting the generation requirement information into the pre-trained feature recognition model to make the feature recognition model generate a character feature vector based on the generation requirement information includes the steps: The input encoding layer extracts multi-modal information based on the input generation requirement information and constructs a multi-modal association matrix; The feature decoupling layer extracts independent decoupled features based on the multi-modal association matrix, and the independent decoupled features include style features, structural features, semantic features, and extended features; The vector fusion layer fuses the independent decoupled features to generate a character feature vector.
[0009] By adopting the above technical solution, the input encoding layer extracts multi-modal information from the generation requirement information, constructs a multi-modal association matrix, and fully explores the potential associations of different modal information such as text, image, and speech; the feature decoupling layer extracts independent decoupled features based on the multi-modal association matrix, including style features, structural features, semantic features, and extended features, to achieve refined decomposition and classification of multi-modal information; then the vector fusion layer fuses the independent decoupled features to generate a role feature vector with rich semantics and multi-dimensional characteristics; through the deep fusion and decoupling processing of multi-modal information, this application realizes the efficient parsing and feature extraction of the generation requirement information, and can provide feature input for subsequent emotion enhancement and label matching, significantly improving the accuracy and expressiveness of role generation.
[0010] In a preferred example, this application can be further configured as follows: the emotion recognition model includes a feature extraction layer, an emotion separation layer, and an enhancement generation layer. The step of inputting the generation requirement information into the pre-trained emotion recognition model to make the emotion recognition model generate emotion enhancement information includes the steps: The feature extraction layer performs multi-modal parsing on the generation requirement information and extracts preliminary emotion features; The emotion separation layer performs emotion dimension separation and scenario adaptation optimization based on the preliminary emotion features to extract contextual emotion features; The enhancement generation layer generates emotion enhancement information based on the contextual emotion features.
[0011] By adopting the above technical solution, the feature extraction layer performs multi-modal parsing on the generation requirement information to extract preliminary emotion features to capture the emotional tendency in the user's requirements; the emotion separation layer performs emotion dimension separation and scenario adaptation optimization on the preliminary emotion features to extract contextual emotion features, realizing refined decomposition and scenario adaptation of emotion features; the enhancement generation layer generates emotion enhancement information based on the contextual emotion features to further refine and strengthen the emotional expression; through multi-modal emotion parsing, emotion dimension separation, and scenario adaptation optimization, this application realizes the in-depth mining and accurate expression of emotion features in the generation requirement information, and can provide more emotionally contagious and scenario-adaptive enhancement information for role generation, significantly improving the emotional expressiveness and personalization degree of the generated role.
[0012] In a preferred example, this application can be further configured as follows: the step of the emotion separation layer performing emotion dimension separation and scenario adaptation optimization based on the preliminary emotion features to extract contextual emotion features includes the steps: The emotion separation layer extracts emotion dimension vectors based on the preliminary emotion features, and the emotion dimension vectors include valence, arousal, and control degree; The emotional separation layer identifies the scenario requirement information in the generated requirement information, and performs weight assignment and contextualized modification extraction on the emotional dimension vector based on the scenario requirement information, so as to generate contextualized emotional features.
[0013] By adopting the above technical solution, the emotional separation layer separates the emotional dimensions of the preliminary emotional features, extracts the emotional dimension vector including valence, arousal, and control degree, and comprehensively analyzes the core attributes of the emotional features; at the same time, it identifies and generates the scenario requirement information in the requirement information, combines the scenario characteristics to perform weight assignment and contextualized modification extraction on the emotional dimension vector, and generates contextualized emotional features adapted to specific scenarios; through the refined separation of emotional dimensions and the dynamic adaptation of scenario requirements in this application, the deep decoupling and contextualized optimization of emotional features are realized, and the rationality and appeal of emotional performance in the character generation process can be significantly improved, making the generated character more in line with the emotional needs expected by users.
[0014] In a preferred example, this application can be further configured as: the label matching model includes a label matching layer and an enhancement optimization layer, and the step of inputting the character feature vector and the emotional enhancement information into the pre-trained label matching model to enable the label matching model to match a number of feature labels based on the pre-stored feature label system and perform emotional enhancement optimization includes the steps of: The label matching layer matches and filters out a candidate label set from the pre-stored feature label system based on the character feature vector; The enhancement optimization layer performs cross-label collaborative optimization based on emotional enhancement on the candidate label set based on the emotional enhancement information, so as to generate enhanced feature labels.
[0015] By adopting the above technical solution, the label matching layer matches and filters out a candidate label set from the pre-stored feature label system based on the character feature vector, realizing the efficient classification and preliminary labeling of character features; the enhancement optimization layer performs cross-label collaborative optimization based on emotional enhancement on the candidate label set based on the emotional enhancement information, further improving the matching degree between the label and the character feature and the richness of emotional expression, so as to generate enhanced feature labels; through the joint modeling of character features and emotional information in this application, the accuracy of label matching and the deep optimization of emotional enhancement are realized, and the integrity of the label system and the emotional expressiveness of the character in the character generation process can be significantly improved, making the generated character more in line with the emotional characteristics and personalized scenarios required by users.
[0016] In a preferred example, this application can be further configured as: the candidate label set includes personality labels, clothing labels, action labels, and background labels, and the step of the enhancement optimization layer performing cross-label collaborative optimization based on emotional enhancement on the candidate label set to generate enhanced feature labels includes the steps of: The enhancement and optimization layer generates a benchmark emotion matrix based on the personality tags, and sets emotion deviation thresholds for the clothing tags, action tags, and background tags respectively; The enhancement and optimization layer performs corresponding optimizations on the clothing tags, action tags, and background tags respectively based on the set emotion deviation thresholds; The enhancement and optimization layer generates enhanced feature tags based on the personality tags and the optimized clothing tags, action tags, and background tags.
[0017] By adopting the above technical solution, the enhancement and optimization layer generates a benchmark emotion matrix based on the personality tags, sets emotion deviation thresholds for the clothing tags, action tags, and background tags respectively, and uses the emotion deviation thresholds to perform corresponding optimizations on the clothing tags, action tags, and background tags to ensure that the emotional expressions of each tag are consistent with the core emotional characteristics of the personality tags; on this basis, enhanced feature tags are generated based on the personality tags and the optimized clothing tags, action tags, and background tags to achieve the emotional collaborative optimization of multi-dimensional tags; through the emotional benchmark guidance of the personality tags and the emotional deviation optimization of multi-tags, the present application realizes the enhancement of the emotional consistency and multi-dimensional collaborative optimization of the role feature tags, can significantly improve the emotional adaptability and overall coordination among the tags in the role generation process, and make the generated role more in line with the emotional characteristics and personalized scenarios required by the user.
[0018] In a preferred example of the present application, it can be further configured as follows: after the step that the enhancement and optimization layer generates enhanced feature tags based on the personality tags and the optimized clothing tags, action tags, and background tags, the following steps are executed: Detect the enhanced feature tags based on the pre-set tag conflict rules; When it is detected that there is a tag conflict, handle the tag conflict based on the conflict handling strategy.
[0019] By adopting the above technical solution, after the enhancement and optimization layer generates enhanced feature tags based on the personality tags and the optimized clothing tags, action tags, and background tags, detect the enhanced feature tags based on the pre-set tag conflict rules to identify possible contradictions or inconsistencies between tags; when a tag conflict is detected, adjust or optimize the conflicting tags based on the conflict handling strategy to ensure that the final generated set of feature tags reaches the optimal state in terms of emotional expression and logical consistency; through the introduction of the tag conflict detection and conflict handling mechanism, the present application realizes the intelligent verification and optimization of the enhanced feature tags, can effectively avoid the problem of the role generation result being mismatched or unnatural caused by tag conflicts, and make the generated role more in line with the emotional needs and scenario adaptability expected by the user.
[0020] In a preferred example, the present application can be further configured as follows: The step of handling the tag conflict based on the conflict handling strategy when detecting the existence of a tag conflict includes the steps: When detecting the existence of a tag conflict, identify its conflict type, where the conflict type includes hard conflict and soft conflict; If the conflict type is a hard conflict, calculate the deviation degree between the enhanced feature tag associated with the hard conflict and the benchmark sentiment matrix, identify the abnormal tag and its tag type, and match the corresponding preset tag replacement strategy based on the tag type; If the conflict type is a soft conflict, identify the tag type of the enhanced feature tag associated with the soft conflict, and match the corresponding preset reconciliation strategy based on the tag type.
[0021] By adopting the above technical solutions, when detecting a tag conflict, identify its conflict type, classify the conflict into two types: hard conflict and soft conflict, and adopt different processing strategies for different conflict types; for a hard conflict, by calculating the deviation degree between the enhanced feature tag and the benchmark sentiment matrix, identify the abnormal tag and its type, and match the preset tag replacement strategy based on the tag type to ensure that the replacement of the conflict tag conforms to the sentiment benchmark; for a soft conflict, by identifying the type of the associated tag and matching the preset reconciliation strategy, optimize and adjust the conflict tag to eliminate disharmony; through the accurate identification and replacement processing of hard conflicts and the reconciliation and optimization of soft conflicts, the present application realizes the intelligent processing of tag conflicts and the maintenance of sentiment consistency, can effectively avoid the problem of inconsistent or unnatural character generation results caused by tag conflicts, and makes the generated characters more in line with the user's expected emotional needs and scenario adaptability.
[0022] The above second inventive object of the present application is achieved by the following technical solutions: An anime character generation system based on feature tags, including: A requirement acquisition module, used to obtain generation requirement information from the user side in response to a character generation request sent by the user side; A first input module, used to input the generation requirement information into a pre-trained feature recognition model, so that the feature recognition model generates a character feature vector based on the generation requirement information; A second input module, used to input the generation requirement information into a pre-trained emotion recognition model, so that the emotion recognition model generates emotion enhancement information; A third input module, used to input the character feature vector and the emotion enhancement information into a pre-trained tag matching model, so that the tag matching model matches a number of feature tags based on a pre-stored feature tag system and performs emotion enhancement optimization; A character generation module, used to generate a corresponding anime character based on the enhanced feature tags output by the tag matching model.
[0023] By adopting the above technical solution, a requirement acquisition module is configured to obtain generation requirement information from the user terminal in response to a role generation request sent by the user terminal; a first input module is configured to input the generation requirement information into a pre-trained feature recognition model, so that the feature recognition model generates a role feature vector based on the generation requirement information; a second input module is configured to input the generation requirement information into a pre-trained emotion recognition model, so that the emotion recognition model generates emotion enhancement information; a third input module is configured to input the role feature vector and the emotion enhancement information into a pre-trained label matching model, so that the label matching model matches a plurality of feature labels based on a pre-stored feature label system and performs emotion enhancement optimization; a role generation module is configured to generate a corresponding anime role based on the enhanced feature labels output by the label matching model.
[0024] In summary, the present application includes at least one of the following beneficial technical effects: 1. The present application extracts multi-modal features through a feature recognition model and constructs a role feature vector. Combining the emotion enhancement information generated by the emotion recognition model, the label matching model is used to perform emotional relevance analysis and optimization on the feature labels, solving the incoordination problem caused by the isolated processing of labels in the traditional technology and realizing the emotional consistency between the feature labels. At the same time, by introducing the emotion enhancement information, the limitation of expressing emotions only through a single visual element is avoided, the comprehensive features of the role can be comprehensively adjusted, the expected emotion can be accurately conveyed, and the accuracy, expressiveness and personalization degree of the generated role are significantly improved, meeting the diverse needs of users for the role image. 2. The present application realizes the intelligent processing of label conflicts and the maintenance of emotional consistency through the accurate identification and replacement of hard conflicts and the reconciliation and optimization of soft conflicts, effectively avoiding the problem of mismatched or unnatural role generation results caused by label conflicts, and making the generated role more in line with the expected emotional needs and scenario adaptability of users. Description of the Drawings
[0025] Figure 1 is a flowchart of an embodiment of a method for generating an anime role based on feature labels according to the present application; Figure 2 is an implementation flowchart of step S20 in an embodiment of a method for generating an anime role based on feature labels according to the present application; Figure 3 is an implementation flowchart of step S30 in an embodiment of a method for generating an anime role based on feature labels according to the present application; Figure 4 is an implementation flowchart of step S42 in an embodiment of a method for generating an anime role based on feature labels according to the present application; Figure 5It is a flowchart of the implementation of step S44 in an embodiment of a method for generating anime characters based on feature tags in this application. Detailed implementation
[0026] The following is a further detailed description of this application in combination with the attached Figures 1 - 5 drawings.
[0027] In one embodiment, as Figure 1 shown, this application discloses a method for generating anime characters based on feature tags, which specifically includes the following steps: S10: In response to a character generation request sent by the client, obtain generation requirement information from the client; In this embodiment, the character generation request is an electrical signal sent by the client to generate an anime character; the generation requirement information is information sent by the client for describing the character generation requirements, including the basic attributes of the character (such as gender, age, personality, etc.), emotional needs (such as "happy", "angry"), scene requirements (such as "battle scene", "daily scene"), etc.; Specifically, receive the character generation request sent by the client (such as a web page, APP, etc.), and obtain the generation requirement information of the user through the interaction interface, where the generation requirement information can be a text description (such as "generate a happy female warrior character") or other forms (such as emotional scoring, scene selection, etc.). For example, the user inputs "I want a happy warrior character in a battle scene, wearing red armor, and the actions should be full of strength".
[0028] S20: Input the generation requirement information into a pre-trained feature recognition model, so that the feature recognition model generates a character feature vector based on the generation requirement information; In this embodiment, the feature recognition model is a pre-trained deep learning model for extracting the basic features of the character from the generation requirement information and generating a character feature vector. Among them, the character feature vector is a multi-dimensional feature representation of the character, including attributes such as personality, clothing, actions, and background; Specifically, input the generation requirement information into the feature recognition model, so that the feature recognition model parses and extracts the basic features of the character and generates a character feature vector. Among them, the character feature vector is a multi-dimensional feature representation of the character, including attributes such as personality, clothing, actions, and background. For example, the feature recognition model extracts the following feature vectors from "a happy warrior character, wearing red armor, and the actions should be full of strength": Personality: happy, brave; Clothing: red armor; Actions: full of strength; Background: battle scene.
[0029] S30: Input the generation requirement information into a pre-trained emotion recognition model, so that the emotion recognition model generates emotion enhancement information; In this embodiment, the emotion recognition model is a pre-trained deep learning model, which is used to extract emotion features from the generated demand information and generate emotion enhancement information. The emotion enhancement information is the refinement and supplementary information of the character's emotion features; Specifically, the generated demand information is input into the emotion recognition model, so that the emotion recognition model analyzes and extracts emotion features and generates emotion enhancement information for refining and supplementing the character's emotion features.
[0030] S40: Input the character feature vector and the emotion enhancement information into the pre-trained label matching model, so that the label matching model matches a number of feature labels based on the pre-stored feature label system and performs emotion enhancement optimization; In this embodiment, the label matching model is a pre-trained deep learning model, which is used to match a number of feature labels from the pre-stored feature label system based on the character feature vector and the emotion enhancement information, and generate enhanced feature labels through emotion enhancement optimization. The feature label system includes personality labels, clothing labels, action labels, background labels, etc.; the enhanced feature labels are the set of feature labels optimized by the label matching model, including the attributes of the character's personality, clothing, action, background, etc., and perform emotion adaptation and optimization on the labels in combination with the emotion enhancement information; Specifically, the character feature vector and the emotion enhancement information are input into the label matching model, so that the label matching model matches a number of feature labels based on the pre-stored feature label system (such as personality labels, clothing labels, action labels, background labels, etc.) and generates enhanced feature labels through emotion enhancement optimization. The emotion enhancement optimization ensures that the labels match the emotion enhancement information and enhances the emotional expression of the character.
[0031] For example, the label matching model matches the following labels according to the character feature vector and the emotion enhancement information: Personality label: brave, optimistic; Clothing label: red armor; Action label: swing sword, run; Background label: battlefield; After emotion enhancement optimization, the enhanced feature labels are generated: Personality label: brave (+0.8 valence); Clothing label: red armor (high arousal modification); Action label: swing sword (full of strength, high arousal); Background label: battlefield (battle scene).
[0032] S50: Generate the corresponding anime character based on the enhanced feature labels output by the label matching model.
[0033] In this embodiment, the anime character is a virtual character generated based on the enhanced feature labels, including attributes such as personality characteristics, clothing styles, action performances, background scenes, etc., and can meet the user's generation requirements; Specifically, an anime character is generated according to the enhanced feature tags output by the tag matching model. The generated anime character includes attributes such as personality traits, clothing styles, action performances, background scenes, etc., and conforms to the user's needs and emotion enhancement information.
[0034] For example, based on the enhanced feature tags: personality tag: brave; clothing tag: red armor; action tag: swinging a sword; background tag: battlefield, an anime character is generated, and the specific manifestation is: personality: brave, optimistic, and with hair color, hairstyle and facial form that match this personality; clothing: red armor with gold decorations; action: waving a huge sword, running posture full of strength; background: on the battlefield, the background is a burning castle and gunpowder smoke.
[0035] In one embodiment, the feature recognition model includes an input encoding layer, a feature decoupling layer, and a vector fusion layer, as Figure 2 shown, step S20 includes the steps: S21: The input encoding layer extracts multimodal information based on the input generation requirement information and constructs a multimodal association matrix; S22: The feature decoupling layer extracts independent decoupled features based on the multimodal association matrix. The independent decoupled features include style features, structural features, semantic features, and extended features; S23: The vector fusion layer fuses the independent decoupled features to generate a character feature vector.
[0036] In this embodiment, the input encoding layer is the first layer of the feature recognition model, which is used for multimodal parsing and encoding of the generation requirement information. The input encoding layer converts the input text, image or other forms of requirement information into a unified numerical representation and constructs a multimodal association matrix to capture the correlation between different modalities; the multimodal information is information from multiple data sources, such as text descriptions, image features, voice information, etc.; the multimodal association matrix is a mathematical representation used to describe the correlation and interaction between different modality information. Through the multimodal association matrix, the semantic relationship between different modality information can be captured, so as to understand the generation requirement information more comprehensively; the feature decoupling layer is the second layer of the feature recognition model, which is used to extract independent decoupled features from the multimodal association matrix; the independent decoupled features are features of different dimensions separated from the multimodal information, and these features are independent of each other and have clear semantic meanings. The independent decoupled features include style features, structural features, semantic features, and extended features; the vector fusion layer is the third layer of the feature recognition model, which is used to fuse the independent decoupled features to generate a unified character feature vector; the character feature vector is the output result of the feature recognition model, which is a multi-dimensional vector used to represent the comprehensive features of the character. The character feature vector includes the style, structure, semantics, and other extended information of the character, and provides input for subsequent emotion recognition and tag matching.
[0037] Specifically, the input encoding layer performs multimodal parsing on the generated demand information, extracts information in text, image, or other modalities, converts this information into a numerical representation, and constructs a multimodal correlation matrix to capture the correlation between different modality information. For example, there may be a semantic correlation between the "warrior" mentioned in the text and the weapon or armor features that may be included in the image; the feature decoupling layer extracts independently decoupled features from the multimodal correlation matrix. These features include style features (such as the visual style of the character), structural features (such as the body proportion of the character), semantic features (such as the character description of the character), and extended features (such as additional scene or prop information). Through decoupling, different dimensions of features can be understood more clearly; the vector fusion layer fuses the independently decoupled features to generate a unified character feature vector. This character feature vector is a numerical representation of the multi-dimensional features of the character and contains the style, structure, semantics, and other extended information of the character.
[0038] In one embodiment, the emotion recognition model includes a feature extraction layer, an emotion separation layer, and an enhancement generation layer. As Figure 3 shown, step S30 includes the steps of: S31: The feature extraction layer performs multimodal parsing based on the generated demand information and extracts preliminary emotion features; S32: The emotion separation layer performs emotion dimension separation and scenario adaptation optimization based on the preliminary emotion features to extract contextual emotion features; S33: The enhancement generation layer generates emotion enhancement information based on the contextual emotion features.
[0039] In this embodiment, the feature extraction layer is the first layer of the emotion recognition model, which is used to perform multimodal parsing on the generated demand information and extract preliminary emotion features. The feature extraction layer extracts emotion-related features from multimodal information such as text, images, and speech as the basis for subsequent emotion separation; the emotion separation layer is the second layer of the emotion recognition model, which is used to perform emotion dimension separation and scenario adaptation optimization on the preliminary emotion features. The emotion separation layer decomposes the emotion features into core dimensions such as valence, arousal, and dominance, and adapts and optimizes the emotion features in combination with the scenario requirements to generate contextual emotion features; the enhancement generation layer is the third layer of the emotion recognition model, which is used to generate emotion enhancement information based on the contextual emotion features. The enhancement generation layer generates more expressive and adaptable emotion enhancement information by refining and supplementing the emotion features to provide support for subsequent label matching and role generation; the preliminary emotion features are the original features related to emotions extracted from the generated demand information, which contain the preliminary emotion tendencies of multimodal information but have not undergone dimension separation and scenario adaptation optimization; the contextual emotion features are emotion features generated after emotion dimension separation and scenario adaptation optimization based on the preliminary emotion features, which are more in line with specific scenario requirements and have higher emotion expressiveness and adaptability; the emotion enhancement information is the emotion refinement result generated based on the contextual emotion features, which contains the specific values of emotion dimensions and emotion modification features related to the scenario.
[0040] Specifically, the feature extraction layer performs multimodal parsing on the generated demand information and extracts preliminary features related to emotions; the emotion separation layer performs emotion dimension separation on the preliminary emotion features, decomposes them into core dimensions such as valence, arousal, and dominance, and adapts and optimizes the emotion features in combination with the scenario requirements (such as "battle scene") in the generated demand information to generate contextual emotion features; the enhancement generation layer further refines and supplements the emotion information based on the contextual emotion features to generate emotion enhancement information. The emotion enhancement information contains the specific values of valence, arousal, and dominance, as well as emotion modification features related to the scenario, providing emotion input for the subsequent label matching model.
[0041] In one embodiment, step S32 includes the steps of: S321: The emotion separation layer extracts an emotion dimension vector based on the preliminary emotion features, and the emotion dimension vector includes valence, arousal, and dominance; S322: The emotion separation layer identifies the scenario requirement information in the generated demand information, and performs weight assignment and contextual modification extraction on the emotion dimension vector based on the scenario requirement information, thereby generating contextual emotion features.
[0042] In this embodiment, the emotional dimension vector is the core dimension extracted from the preliminary emotional features to describe the emotional features, including valence, arousal, and dominance. Valence represents the positive or negative degree of emotion, usually ranging from [-1, 1], where negative values indicate negative emotions and positive values indicate positive emotions. Arousal represents the intensity or activation degree of emotion, usually ranging from [0, 1], and the higher the value, the stronger the emotion. Dominance represents the dominance or sense of control of emotion, usually ranging from [0, 1], and the higher the value, the more dominant the emotion. The scenario requirement information is the description related to a specific scenario in the generated requirement information, such as "battle scenario", "daily scenario", etc. The scenario requirement information is used to guide the adaptation and optimization of emotional features to make the emotional features more in line with the requirements of a specific scenario. The contextualized emotional feature is an emotional feature generated after weight assignment and contextualized modification extraction based on the emotional dimension vector and combined with the scenario requirement information. It is more in line with the specific scenario requirements and has higher emotional expressiveness and adaptability. Weight assignment is to assign different weights to dimensions such as valence, arousal, and dominance in the emotional dimension vector according to the scenario requirement information to highlight the emotional features in a specific scenario. For example, in a "battle scenario", the weights of arousal and dominance may be increased. Contextualized modification extraction is to extract modification features related to the scenario requirement from the emotional dimension vector, such as "aggressiveness", "explosiveness", "gentleness", etc., to further refine and supplement the emotional features.
[0043] Specifically, the emotion separation layer decouples the preliminary emotional features and extracts the emotional dimension vector, including valence, arousal, and dominance. The emotional dimension vector is the core description of the emotional features and can reflect the basic attributes of emotions. The emotion separation layer identifies the scenario requirement information (such as "battle scenario") in the generated requirement information, assigns weights to the emotional dimension vector according to the scenario requirement, and highlights the emotional features in a specific scenario. At the same time, it extracts contextualized modification features to further refine and supplement the emotional features and generates contextualized emotional features.
[0044] In one embodiment, the label matching model includes a label matching layer and an enhancement and optimization layer. Step S40 includes the steps: S41: The label matching layer matches and filters out a candidate label set from the pre-stored feature label system based on the role feature vector; S42: The enhancement and optimization layer performs cross-label collaborative optimization based on emotion enhancement on the candidate label set based on the emotion enhancement information, thereby generating enhanced feature labels.
[0045] In this embodiment, the label matching layer is the first layer of the label matching model, which is used to match and filter out a candidate label set from a pre-stored feature label system based on the role feature vector. The candidate label set includes character labels, clothing labels, action labels, background labels, etc. The enhancement and optimization layer is the second layer of the label matching model, which is used to perform cross-label collaborative optimization on the candidate label set based on the emotion enhancement information, enhance the adaptability of the labels to the role features and emotion features, and generate enhanced feature labels. The role feature vector is the output result of the feature recognition model, which contains multi-dimensional feature information such as the character, clothing, action, and background of the role, and is used to describe the comprehensive features of the role. The emotion enhancement information is the output result of the emotion recognition model, which contains specific values of emotion dimensions such as valence, arousal, and control degree, as well as emotion modification features related to the scene, and is used to refine and supplement the role emotion features. The candidate label set is a preliminary label set filtered from the pre-stored feature label system based on the role feature vector, including character labels, clothing labels, action labels, background labels, etc. The enhanced feature labels are the label set optimized by the enhancement and optimization layer, which combines the role feature vector and the emotion enhancement information, and has higher emotion adaptability and scene adaptability. The pre-stored feature label system is a pre-defined label library, which contains labels of multiple dimensions. For example: character labels: brave, optimistic, angry, melancholy; clothing labels: red armor, blue robe, black cloak; action labels: swing sword, run, jump; background labels: battlefield, forest, city.
[0046] Specifically, the label matching layer analyzes the role feature vector, extracts features such as the character, clothing, action, and background of the role, and matches and filters out a candidate label set related to the role features from the pre-stored feature label system. The candidate label set is a preliminary label set and has not been optimized by emotion enhancement yet. The enhancement and optimization layer combines emotion enhancement information (such as emotion dimension values such as valence, arousal, and control degree and contextual modification features), performs cross-label collaborative optimization on the candidate label set, and enhances the adaptability of the labels to the role features and emotion features through cross-label collaborative optimization to generate enhanced feature labels.
[0047] In one embodiment, the candidate label set includes character labels, clothing labels, action labels, and background labels. As Figure 4 shown, step S42 includes the steps: S421: The enhancement and optimization layer generates a benchmark emotion matrix based on the character label, and sets emotion deviation thresholds for the clothing label, action label, and background label respectively; S422: The enhancement and optimization layer performs corresponding optimization on the clothing label, action label, and background label respectively based on the set emotion deviation thresholds; S423: The enhancement and optimization layer generates enhanced feature labels based on the character label and the optimized clothing label, action label, and background label.
[0048] In this embodiment, the candidate label set is a preliminary label set selected from a pre-stored feature label system based on the role feature vector, including personality labels, clothing labels, action labels, and background labels; the personality labels are labels describing the personality characteristics of the role, such as "brave", "angry", "optimistic", etc. The personality label is the core of the role's emotional characteristics and is usually used as a benchmark for emotional optimization; the clothing label is a label describing the clothing style of the role, such as "red armor", "blue robe", "black cloak", etc. The emotional optimization of the clothing label needs to match the emotional characteristics of the personality label; the action label is a label describing the action performance of the role, such as "swinging a sword", "running", "jumping", etc. The emotional optimization of the action label needs to reflect the emotional intensity and action characteristics of the role; the background label is a label describing the scene where the role is located, such as "battlefield", "forest", "city", etc. The emotional optimization of the background label needs to match the emotional characteristics of the role and the scene requirements; the benchmark emotion matrix is an emotion feature matrix generated based on the personality label, used to describe the emotional dimensions (valence, arousal, control) of the personality label and their mutual relationships. The benchmark emotion matrix provides a reference for the emotional optimization of other labels; the emotion deviation threshold is the range of emotional feature adjustment set for the clothing label, action label, and background label, used to control the matching degree of the emotional characteristics of these labels with the emotional characteristics of the personality label. The larger the deviation threshold, the wider the range of emotional feature adjustment; the enhanced feature label is a set of labels optimized by the enhanced optimization layer, including the optimization results of the personality label, clothing label, action label, and background label. The enhanced feature label is more in line with the role's emotional characteristics and scene requirements.
[0049] Specifically, the enhanced optimization layer generates a benchmark emotion matrix based on the personality label, describes the emotional dimensions (valence, arousal, control) of the personality label, and sets emotion deviation thresholds for the clothing label, action label, and background label respectively, used to control the matching range of the emotional characteristics of these labels with the emotional characteristics of the personality label; the enhanced optimization layer optimizes the clothing label, action label, and background label according to the benchmark emotion matrix of the personality label and the set emotion deviation thresholds. During the optimization process, ensure that the emotional characteristics of these labels are consistent with the emotional characteristics of the personality label, and at the same time reflect the adaptability of the role and the scene; the enhanced optimization layer integrates the personality label with the optimized clothing label, action label, and background label to generate an enhanced feature label. The enhanced feature label contains the optimization results of the role's personality, clothing, action, and background, and has higher emotional adaptability and scene adaptability.
[0050] In one embodiment, after step S423, the following steps are performed: S43: Detect the enhanced feature label based on the pre-set label conflict rule; S44: When a tag conflict is detected, handle the tag conflict based on the conflict handling strategy.
[0051] In this embodiment, a tag conflict refers to a situation where there are contradictions or inconsistencies in the emotional features, semantic features, or scene adaptability among different tags in the enhanced feature tags. For example, if the personality tag is "angry", but the clothing tag is "white robe" (usually associated with calmness or elegance), it may lead to inconsistent emotional features; the tag conflict rule is a predefined rule set used to detect whether there are conflicts in the enhanced feature tags, and the tag conflict rules include: Emotional dimension conflict: such as significant deviations in valence, arousal, and control; Semantic conflict: such as the semantic meanings of tags being contradictory to each other (such as "angry" and "gentle"); Scene adaptation conflict: such as the tag not matching the scene requirements (such as "peace dove" appearing in a "battle scene"); The conflict handling strategy is the processing method adopted for the detected tag conflict, used to eliminate or alleviate the conflict, and the conflict handling strategies include: Replacing tags: replacing the conflicting tags with tags that are more consistent with the character features and emotional features; Adjusting emotional features: adjusting the emotional dimension values of the conflicting tags to make them consistent with other tags; Priority processing: retaining important tags according to the priority of the tags (such as personality tags taking precedence over clothing tags) and adjusting secondary tags; The enhanced feature tags are a set of tags optimized based on the character feature vector and emotional enhancement information, including personality tags, clothing tags, action tags, and background tags. After conflict detection and processing, the enhanced feature tags will be more coordinated and consistent.
[0052] Specifically, the enhancement and optimization layer uses the preset tag conflict rules to detect the enhanced feature tags, identify whether there are conflicts between tags, and the detection content includes emotional dimension conflict, semantic conflict, and scene adaptation conflict; if a conflict is detected, record the conflict type and the tags involved, and the enhancement and optimization layer adjusts or replaces the conflicting tags according to the conflict handling strategy to eliminate or alleviate the conflict, and the handling strategies include replacing tags, adjusting emotional features, or priority processing.
[0053] In one embodiment, as Figure 5 shown, step S44 includes the steps: S441: When a tag conflict is detected, identify its conflict type, and the conflict type includes hard conflict and soft conflict; S442: If the conflict type is a hard conflict, calculate the deviation degree of the enhanced feature tags associated with the hard conflict from the benchmark emotion matrix, identify the abnormal tags and their tag types, and match the corresponding preset tag replacement strategy based on the tag type; S443: If the conflict type is a soft conflict, identify the tag types of the enhanced feature tags associated with the soft conflict, and match the corresponding preset reconciliation strategy based on the tag type.
[0054] In this embodiment, a hard conflict refers to a situation where there are significant contradictions or inconsistencies between enhanced feature labels, which cannot be resolved by simple adjustment. A hard conflict needs to be resolved by replacing labels; a soft conflict refers to a situation where there are minor contradictions or inconsistencies between enhanced feature labels, which can be resolved by adjusting emotional features or semantic modifications. For example, the arousal level of the clothing label "blue robe" is slightly lower than that of the personality label "angry". This kind of conflict can be alleviated by adjusting the emotional dimension value; the deviation degree is the degree of difference between the emotional dimension value of the enhanced feature label and the emotional dimension value of the benchmark emotional matrix. The deviation degree is used to quantify the emotional consistency between the label and the personality label. The higher the deviation degree, the more significant the conflict; the label replacement strategy is a pre-set processing method for hard conflicts, which is used to replace the conflicting label with a label that is more consistent with the character features and emotional features. The label replacement strategy is usually based on the label type (such as clothing label, action label) and scene requirements; the reconciliation strategy is a pre-set processing method for soft conflicts, which is used to alleviate the conflict by adjusting the emotional dimension value or semantic modification of the label to make it more consistent with the emotional features of other labels; the label type is the category of enhanced feature labels, including personality labels, clothing labels, action labels, and background labels. Different types of labels may apply different strategies in conflict processing.
[0055] Specifically, after detecting a label conflict, the enhancement and optimization layer first identifies the type of conflict. The conflict types are divided into hard conflicts and soft conflicts; for hard conflicts, the enhancement and optimization layer calculates the emotional dimension deviation degree between the conflicting label and the benchmark emotional matrix, quantifies the severity of the conflict, and identifies the abnormal label and its type. Based on the label type, it matches the pre-set label replacement strategy and replaces the abnormal label with a more appropriate label; for soft conflicts, the enhancement and optimization layer identifies the type of the conflicting label and matches the pre-set reconciliation strategy based on the label type, and alleviates the conflict by adjusting the emotional dimension value or semantic modification to make it more consistent with the emotional features of other labels.
[0056] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0057] In one embodiment, a feature label-based anime character generation system is provided. The feature label-based anime character generation system corresponds one-to-one with the feature label-based anime character generation method in the above embodiment.
[0058] A feature label-based anime character generation system includes: A requirement acquisition module, which is used to obtain generation requirement information from the user terminal in response to a character generation request sent by the user terminal; The first input module is used to input the generated requirement information into a pre-trained feature recognition model, so that the feature recognition model generates a character feature vector based on the generated requirement information; The second input module is used to input the generated requirement information into a pre-trained emotion recognition model, so that the emotion recognition model generates emotion enhancement information; The third input module is used to input the character feature vector and the emotion enhancement information into a pre-trained label matching model, so that the label matching model matches a number of feature labels based on a pre-stored feature label system and performs emotion enhancement optimization; The character generation module is used to generate a corresponding anime character based on the enhanced feature labels output by the label matching model.
[0059] For the specific limitations of an anime character generation system based on feature labels, reference can be made to the limitations of an anime character generation method based on feature labels in the foregoing text, which will not be elaborated here. Each module in the above-mentioned anime character generation system based on feature labels can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0060] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for generating an animation character based on feature tags, characterized in that: Includes steps: In response to a role generation request issued by a user terminal, obtaining generation requirement information from the user terminal; Inputting the generation requirement information into a pre-trained feature recognition model, so that the feature recognition model generates a role feature vector based on the generation requirement information; Inputting the generation demand information into a pre-trained emotion recognition model so that the emotion recognition model generates emotion enhancement information; Inputting the character feature vector and the emotion enhancement information into the pre-trained label matching model, so that the label matching model matches a number of feature labels based on the pre-stored feature label system and performs emotion enhancement optimization; Generate corresponding anime characters based on the enhanced feature labels output by the label matching model.
2. The method for generating an animated character based on feature tags according to claim 1, characterized in that: The feature recognition model includes an input coding layer, a feature decoupling layer and a vector fusion layer. The step of inputting the generation requirement information into the pre-trained feature recognition model so that the feature recognition model generates a role feature vector based on the generation requirement information includes the following steps: The input encoding layer extracts multimodal information based on the input generation requirement information and constructs a multimodal association matrix; The feature decoupling layer extracts independent decoupled features based on the multimodal association matrix, wherein the independent decoupled features include style features, structural features, semantic features, and extended features; The vector fusion layer fuses the independent decoupled features to generate a character feature vector.
3. The method for generating an animated character based on feature tags according to claim 1, characterized in that: The emotion recognition model includes a feature extraction layer, an emotion separation layer, and an enhancement generation layer. The step of inputting the generation requirement information into the pre-trained emotion recognition model so that the emotion recognition model generates emotion enhancement information includes the following steps: The feature extraction layer performs multimodal analysis based on the generated demand information and extracts preliminary emotional features; The emotion separation layer performs emotion dimension separation and scenario adaptation optimization based on preliminary emotion features to extract contextualized emotion features. The enhancement generation layer generates emotion enhancement information based on contextualized emotion features.
4. The method for generating an animated character based on feature tags according to claim 3, characterized in that: The emotion separation layer performs emotion dimension separation and scenario adaptation optimization based on the preliminary emotion features to extract the contextualized emotion features, including the following steps: The emotion separation layer extracts an emotion dimension vector based on the preliminary emotion features, wherein the emotion dimension vector includes valence, arousal, and control; The emotion separation layer identifies the scene demand information in the generated demand information, and based on the scene demand information, performs weight assignment and contextual modification extraction on the emotion dimension vector, thereby generating contextualized emotion features.
5. The method for generating an animated character based on feature tags according to claim 1, characterized in that: The label matching model includes a label matching layer and an enhancement optimization layer. The step of inputting the character feature vector and the emotion enhancement information into the pre-trained label matching model so that the label matching model matches a plurality of feature labels based on a pre-stored feature label system and performs emotion enhancement optimization includes the following steps: The tag matching layer matches and selects candidate tag sets from the pre-stored feature tag system based on the role feature vector; The enhanced optimization layer performs sentiment-enhanced cross-label collaborative optimization on the candidate label set based on sentiment enhancement information to generate enhanced feature labels.
6. The method for generating an animated character based on feature tags according to claim 5, characterized in that: The candidate tag set includes personality tags, clothing tags, action tags and background tags. The enhancement optimization layer performs emotion-enhanced cross-tag collaborative optimization on the candidate tag set based on emotion enhancement information to generate enhanced feature tags, including the following steps: The enhanced optimization layer generates a baseline sentiment matrix based on the personality label and sets sentiment deviation thresholds for clothing labels, action labels, and background labels respectively; The enhanced optimization layer optimizes clothing labels, action labels, and background labels based on the set emotional deviation threshold. The enhancement optimization layer generates enhanced feature labels based on the personality labels and the optimized clothing labels, action labels, and background labels.
7. The method for generating an animated character based on feature tags according to claim 6, characterized in that: After the step of generating an enhanced feature tag based on the personality tag and the optimized clothing tag, action tag and background tag, the enhanced optimization layer performs the following steps: Detecting enhanced feature tags based on preset tag conflict rules; When a label conflict is detected, the label conflict is handled based on a conflict handling strategy.
8. The method for generating an animated character based on feature tags according to claim 7, characterized in that: The step of handling the label conflict based on the conflict handling strategy when a label conflict is detected comprises the following steps: When a tag conflict is detected, identifying the conflict type, wherein the conflict type includes a hard conflict and a soft conflict; If the conflict type is a hard conflict, the deviation between the enhanced feature label associated with the hard conflict and the benchmark sentiment matrix is calculated, and the abnormal label and its label type are identified, and the corresponding preset label replacement strategy is matched based on the label type; If the conflict type is a soft conflict, the tag type of the enhanced feature tag associated with the soft conflict is identified, and a corresponding preset reconciliation strategy is matched based on the tag type.
9. A feature tag-based animation character generation system, characterized in that: include: A requirement acquisition module, used for obtaining generation requirement information from the user terminal in response to a role generation request issued by the user terminal; A first input module is used to input the generation requirement information into a pre-trained feature recognition model, so that the feature recognition model generates a role feature vector based on the generation requirement information; The second input module is used to input the generation demand information into the pre-trained emotion recognition model, so that the emotion recognition model generates emotion enhancement information; The third input module is used to input the character feature vector and the emotion enhancement information into the pre-trained label matching model, so that the label matching model matches a number of feature labels based on the pre-stored feature label system and performs emotion enhancement optimization; The character generation module is used to generate corresponding animation characters based on the enhanced feature labels output by the label matching model.
Citation Information
Patent Citations
Internet video script role emotion recognition method based on big data
CN116821333A
Method for generating role expression animation based on artificial intelligence semantic analysis
CN117611715A
Method and system for generating synthesis voice using style tag represented by natural language
US20240105160A1