Anime Character Generation Method and System Based on Feature Tags

Through feature recognition and emotion recognition models, character feature vectors and emotional enhancement information are generated, and emotional correlation analysis is performed in combination with label matching model. The inconsistency problem caused by label isolation processing in anime character generation is solved, the consistency of character features and emotions is achieved, and the accuracy and expressiveness of generation are improved.

CN120070675BActive Publication Date: 2025-07-04GUANGDONG OPEN UNIV (GUANGDONG POLYTECHNIC VOCATIONAL COLLEGE)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510539384.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-04
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

In the existing anime character generation technology, the tag matching mechanism isolated the various feature tags, ignoring the emotional correlation between tags, resulting in inconsistent visual performance of the generated characters, and lacking refined control of the emotional expression of the characters, making it difficult to generate a character image that meets user needs.

Method used

The character feature vector is generated through the feature recognition model, and the emotional enhancement information is generated by combining the emotion recognition model. The label matching model is used for emotional correlation analysis and optimization, and enhancement feature tags are generated to achieve cross-label collaborative optimization to ensure the consistency of character features and emotions.

Benefits of technology

It significantly improves the accuracy and expressiveness of anime character generation, meets users' diverse needs for character images, avoids the limitations of expressing emotions in a single visual element, and the generated characters are more in line with the user's expected emotions and scene adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070675B_ABST
    Figure CN120070675B_ABST
Patent Text Reader

Abstract

The present application relates to a method and system for generating anime characters based on feature tags, which includes the steps of: obtaining generation requirement information from a user terminal and inputting it into a feature recognition model to generate a character feature vector and inputting it into an emotion recognition model to generate emotion enhancement information; inputting the character feature vector and the emotion enhancement information into a tag matching model to match a number of feature tags and perform emotion enhancement optimization; generating an anime character based on the enhanced feature tags; The present application extracts multi-modal features through a feature recognition model and constructs a character feature vector, combines the emotion enhancement information generated by the emotion recognition model, and uses the tag matching model to perform emotional relevance analysis and optimization on the feature tags, achieving emotional consistency between the feature tags; at the same time, by introducing emotion enhancement information, it avoids the limitation of expressing emotions only through a single visual element, and has the effects of improving the accuracy of the generated characters and the expressiveness of the characters, and meeting the diverse needs of users for character images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of virtual character generation, and particularly to an anime character generation method and system based on feature tags. Background Art

[0002] Anime character generation refers to automatically or semi-automatically creating anime character images with specific appearances, personalities, and characteristics through computer technology. With the development of artificial intelligence and computer graphics, this technology has evolved from early manual drawing to the current intelligent generation stage.

[0003] In existing anime character generation technologies, methods based on template matching or rule-driven are usually adopted to construct character images. Such technologies often rely on preset character part libraries and fixed matching rules. For example, visual elements such as clothing and hairstyles are directly matched through text keywords input by users.

[0004] However, existing anime character generation technologies have the following significant defects: First, traditional label matching mechanisms often process various feature tags in isolation, ignoring the emotional relevance between tags, and easily generating characters with inconsistent visual presentations. For example, a "graceful" personality label may conflict with an "attack posture" action label. Second, most systems lack refined control over character emotional expression, resulting in generated character images not matching the expected emotions. For example, an "angry" character may only be represented by a simple red color tone, lacking details such as micro-expressions and action tensions. The above problems not only severely limit the expressiveness and personalization of the generated characters but also make it difficult to accurately generate the anime characters required by users. Summary of the Invention

[0005] To address the above defects, this application provides an anime character generation method and system based on feature tags.

[0006] The first invention objective of this application is achieved through the following technical solutions:

[0007] An anime character generation method based on feature tags, comprising the steps of:

[0008] Responding to a character generation request issued by the client, obtaining generation requirement information from the client;

[0009] Inputting the generation requirement information into a pre-trained feature recognition model, enabling the feature recognition model to generate a character feature vector based on the generation requirement information;

[0010] Inputting the generation requirement information into a pre-trained emotion recognition model, enabling the emotion recognition model to generate emotion enhancement information;

[0011] Input the character feature vector and the emotion enhancement information into the pre-trained label matching model, so that the label matching model matches a number of feature labels based on the pre-stored feature label system and performs emotion enhancement optimization;

[0012] Generate the corresponding anime character based on the enhanced feature labels output by the label matching model.

[0013] By adopting the above technical solution, with the emotion-enhanced label matching model, it realizes the fine control of the anime character generation process and the cross-label collaborative optimization, and has the effect of significantly improving the accuracy of the generated character, as well as the expressiveness and personalization degree of the character; specifically, by responding to the character generation request sent by the client, obtaining the generation requirement information, and inputting it into the pre-trained feature recognition model and emotion recognition model, generating the character feature vector and emotion enhancement information respectively, and then through the label matching model based on the pre-stored feature label system for label matching and emotion enhancement optimization, finally generating the corresponding anime character; this application extracts multi-modal features through the feature recognition model and constructs the character feature vector, combines the emotion enhancement information generated by the emotion recognition model, and uses the label matching model to perform emotion correlation analysis and optimization on the feature labels, solves the incoordination problem caused by the isolated processing of labels in the traditional technology, and realizes the emotion consistency between the feature labels; at the same time, by introducing the emotion enhancement information, it avoids the limitation of expressing emotions only through a single visual element, can comprehensively adjust the comprehensive features of the character, accurately convey the expected emotion, has the effect of significantly improving the accuracy of the generated character, as well as the expressiveness and personalization degree of the character, and meets the diverse needs of users for the character image.

[0014] This application can be further configured in a preferred example as follows: the feature recognition model includes an input encoding layer, a feature decoupling layer, and a vector fusion layer. The step of inputting the generation requirement information into the pre-trained feature recognition model to make the feature recognition model generate a character feature vector based on the generation requirement information includes the steps:

[0015] The input encoding layer extracts multi-modal information based on the input generation requirement information and constructs a multi-modal association matrix;

[0016] The feature decoupling layer extracts independent decoupled features based on the multi-modal association matrix, and the independent decoupled features include style features, structural features, semantic features, and extended features;

[0017] The vector fusion layer fuses the independent decoupled features to generate a character feature vector.

[0018] By adopting the above technical solution, the input encoding layer extracts multi-modal information from the generation requirement information, constructs a multi-modal association matrix, and fully explores the potential associations of different modal information such as text, images, and voices; the feature decoupling layer extracts independent decoupled features based on the multi-modal association matrix, including style features, structural features, semantic features, and extended features, to achieve refined decomposition and classification of multi-modal information; then the vector fusion layer fuses the independent decoupled features to generate a role feature vector with rich semantics and multi-dimensional characteristics; through the deep fusion and decoupling processing of multi-modal information, this application realizes the efficient parsing and feature extraction of the generation requirement information, and can provide feature input for subsequent emotion enhancement and label matching, significantly improving the accuracy and expressiveness of role generation.

[0019] In a preferred example, this application can be further configured as follows: the emotion recognition model includes a feature extraction layer, an emotion separation layer, and an enhancement generation layer. The step of inputting the generation requirement information into the pre-trained emotion recognition model to enable the emotion recognition model to generate emotion enhancement information includes the steps:

[0020] The feature extraction layer performs multi-modal parsing on the generation requirement information and extracts preliminary emotion features;

[0021] The emotion separation layer performs emotion dimension separation and scenario adaptation optimization based on the preliminary emotion features to extract contextual emotion features;

[0022] The enhancement generation layer generates emotion enhancement information based on the contextual emotion features.

[0023] By adopting the above technical solution, the feature extraction layer performs multi-modal parsing on the generation requirement information and extracts preliminary emotion features to capture the emotional tendency in the user's requirements; the emotion separation layer performs emotion dimension separation and scenario adaptation optimization on the preliminary emotion features to extract contextual emotion features, realizing refined decomposition and scenario adaptation of emotion features; the enhancement generation layer generates emotion enhancement information based on the contextual emotion features to further refine and strengthen the emotional expression; through multi-modal emotion parsing, emotion dimension separation, and scenario adaptation optimization, this application realizes the in-depth mining and accurate expression of emotion features in the generation requirement information, and can provide more emotionally contagious and scenario-adaptive enhancement information for role generation, significantly improving the emotional expressiveness and personalization degree of the generated role.

[0024] In a preferred example, this application can be further configured as follows: the step of the emotion separation layer performing emotion dimension separation and scenario adaptation optimization based on the preliminary emotion features to extract contextual emotion features includes the steps:

[0025] The emotion separation layer extracts an emotion dimension vector based on the preliminary emotion features, and the emotion dimension vector includes valence, arousal, and controllability;

[0026] The emotion separation layer identifies and generates the scenario requirement information in the demand information, and performs weight assignment and contextualized modification extraction on the emotion dimension vector based on the scenario requirement information, so as to generate contextualized emotion features.

[0027] By adopting the above technical solution, the emotion separation layer separates the emotion dimensions of the preliminary emotion features, extracts the emotion dimension vector including valence, arousal, and controllability, and comprehensively analyzes the core attributes of the emotion features; at the same time, it identifies and generates the scenario requirement information in the demand information, combines the scenario characteristics to perform weight assignment and contextualized modification extraction on the emotion dimension vector, and generates contextualized emotion features adapted to a specific scenario; through the refined separation of emotion dimensions and the dynamic adaptation of scenario requirements in this application, the deep decoupling and contextual optimization of emotion features are realized, and the rationality and appeal of emotion expression in the character generation process can be significantly improved, making the generated character more in line with the emotional needs expected by the user.

[0028] In a preferred example of this application, it can be further configured that: the label matching model includes a label matching layer and an enhancement optimization layer, and the step of inputting the character feature vector and the emotion enhancement information into the pre-trained label matching model to make the label matching model match several feature labels based on the pre-stored feature label system and perform emotion enhancement optimization includes the steps:

[0029] The label matching layer matches and filters out a candidate label set from the pre-stored feature label system based on the character feature vector;

[0030] The enhancement optimization layer performs cross-label collaborative optimization based on emotion enhancement on the candidate label set based on the emotion enhancement information, so as to generate enhanced feature labels.

[0031] By adopting the above technical solution, the label matching layer matches and filters out a candidate label set from the pre-stored feature label system based on the character feature vector, realizing the efficient classification and preliminary labeling of character features; the enhancement optimization layer performs cross-label collaborative optimization based on emotion enhancement on the candidate label set based on the emotion enhancement information, further improving the matching degree between the label and the character features and the richness of emotion expression, so as to generate enhanced feature labels; through the joint modeling of character features and emotion information in this application, the accuracy of label matching and the deep optimization of emotion enhancement are realized, and the integrity of the label system and the character emotion expressiveness in the character generation process can be significantly improved, making the generated character more in line with the emotion characteristics and personalized scenarios required by the user.

[0032] In a preferred example, the present application can be further configured as follows: the candidate tag set includes personality tags, clothing tags, action tags, and background tags. The step of the enhancement and optimization layer performing sentiment-enhanced cross-tag collaborative optimization on the candidate tag set based on sentiment enhancement information to generate enhanced feature tags includes the steps:

[0033] The enhancement and optimization layer generates a benchmark sentiment matrix based on the personality tags, and sets sentiment deviation thresholds for the clothing tags, action tags, and background tags respectively;

[0034] The enhancement and optimization layer performs corresponding optimizations on the clothing tags, action tags, and background tags respectively based on the set sentiment deviation thresholds;

[0035] The enhancement and optimization layer generates enhanced feature tags based on the personality tags and the optimized clothing tags, action tags, and background tags.

[0036] By adopting the above technical solution, the enhancement and optimization layer generates a benchmark sentiment matrix based on the personality tags, and sets sentiment deviation thresholds for the clothing tags, action tags, and background tags respectively. The sentiment deviation thresholds are used to perform corresponding optimizations on the clothing tags, action tags, and background tags to ensure that the sentiment expressions of each tag are consistent with the core sentiment features of the personality tags; on this basis, enhanced feature tags are generated based on the personality tags and the optimized clothing tags, action tags, and background tags to achieve sentiment collaborative optimization of multi-dimensional tags. Through the sentiment benchmark guidance of the personality tags and the sentiment deviation optimization of multi-tags, the present application realizes the enhancement of the sentiment consistency and multi-dimensional collaborative optimization of the role feature tags, which can significantly improve the sentiment adaptability and overall coordination among the tags in the role generation process, making the generated role more in line with the sentiment features and personalized scenarios required by the user.

[0037] In a preferred example, the present application can be further configured as follows: after the step in which the enhancement and optimization layer generates enhanced feature tags based on the personality tags and the optimized clothing tags, action tags, and background tags, the following steps are executed:

[0038] Detect the enhanced feature tags based on pre-set tag conflict rules;

[0039] When it is detected that there is a tag conflict, handle the tag conflict based on the conflict handling strategy.

[0040] By adopting the above technical solution, after generating enhanced feature tags based on personality tags and optimized clothing tags, action tags, and background tags in the enhancement and optimization layer, the enhanced feature tags are detected based on pre-set tag conflict rules to identify possible contradictions or inconsistencies between tags; when tag conflicts are detected, the conflicting tags are adjusted or optimized based on the conflict handling strategy to ensure that the finally generated set of feature tags reaches an optimal state in terms of emotional expression and logical consistency; through the introduction of the tag conflict detection and conflict handling mechanism in this application, intelligent verification and optimization of the enhanced feature tags are achieved, which can effectively avoid problems such as mismatched or unnatural character generation results caused by tag conflicts, and make the generated characters more in line with the emotional needs and scenario adaptability expected by users.

[0041] In a preferred example, this application can be further configured as: the step of handling tag conflicts based on the conflict handling strategy when tag conflicts are detected includes the steps of:

[0042] When tag conflicts are detected, identify their conflict types, where the conflict types include hard conflicts and soft conflicts;

[0043] If the conflict type is a hard conflict, calculate the deviation degree of the enhanced feature tags associated with the hard conflict from the benchmark emotion matrix, identify the abnormal tags and their tag types, and match the corresponding pre-set tag replacement strategy based on the tag types;

[0044] If the conflict type is a soft conflict, identify the tag types of the enhanced feature tags associated with the soft conflict, and match the corresponding pre-set reconciliation strategy based on the tag types.

[0045] By adopting the above technical solution, when detecting tag conflicts, identify their conflict types, classify the conflicts into two types: hard conflicts and soft conflicts, and adopt different handling strategies for different conflict types; for hard conflicts, by calculating the deviation degree of the enhanced feature tags from the benchmark emotion matrix, identify the abnormal tags and their types, and match the pre-set tag replacement strategy based on the tag types to ensure that the replacement of the conflicting tags conforms to the emotion benchmark; for soft conflicts, by identifying the types of the associated tags and matching the pre-set reconciliation strategy, optimize and adjust the conflicting tags to eliminate inconsistencies; through the accurate identification and replacement processing of hard conflicts and the reconciliation and optimization of soft conflicts in this application, intelligent handling of tag conflicts and maintenance of emotional consistency are achieved, which can effectively avoid problems such as mismatched or unnatural character generation results caused by tag conflicts, and make the generated characters more in line with the emotional needs and scenario adaptability expected by users.

[0046] The second above-mentioned inventive object of this application is achieved through the following technical solution:

[0047] Anime character generation system based on feature tags, comprising:

[0048] A requirement acquisition module, configured to obtain generation requirement information from the client in response to a character generation request sent by the client;

[0049] A first input module, configured to input the generation requirement information into a pre-trained feature recognition model, so that the feature recognition model generates a character feature vector based on the generation requirement information;

[0050] A second input module, configured to input the generation requirement information into a pre-trained emotion recognition model, so that the emotion recognition model generates emotion enhancement information;

[0051] A third input module, configured to input the character feature vector and the emotion enhancement information into a pre-trained label matching model, so that the label matching model matches a number of feature tags based on a pre-stored feature tag system and performs emotion enhancement optimization;

[0052] A character generation module, configured to generate a corresponding anime character based on the enhanced feature tags output by the label matching model.

[0053] By adopting the above technical solution, a requirement acquisition module, configured to obtain generation requirement information from the client in response to a character generation request sent by the client; a first input module, configured to input the generation requirement information into a pre-trained feature recognition model, so that the feature recognition model generates a character feature vector based on the generation requirement information; a second input module, configured to input the generation requirement information into a pre-trained emotion recognition model, so that the emotion recognition model generates emotion enhancement information; a third input module, configured to input the character feature vector and the emotion enhancement information into a pre-trained label matching model, so that the label matching model matches a number of feature tags based on a pre-stored feature tag system and performs emotion enhancement optimization; a character generation module, configured to generate a corresponding anime character based on the enhanced feature tags output by the label matching model.

[0054] In summary, the present application includes at least one of the following beneficial technical effects:

[0055] 1. The present application extracts multi-modal features through a feature recognition model and constructs a character feature vector. Combining the emotion enhancement information generated by the emotion recognition model, the label matching model is used to perform emotion correlation analysis and optimization on the feature tags, solving the incoordination problem caused by the isolated processing of tags in the traditional technology and realizing the emotion consistency between feature tags; at the same time, by introducing emotion enhancement information, the limitation of expressing emotions only through a single visual element is avoided, the comprehensive features of the character can be adjusted comprehensively, the expected emotion can be accurately conveyed, and the accuracy, expressiveness and personalization degree of the generated character are significantly improved, meeting the diverse needs of users for the character image.

[0056] 2. This application realizes the intelligent processing of label conflicts and the maintenance of emotional consistency through the precise identification and replacement of hard conflicts and the reconciliation and optimization of soft conflicts, which can effectively avoid the problems of mismatched or unnatural character generation results caused by label conflicts, and make the generated characters more in line with the emotional needs and scene adaptability expected by users. Description of the Drawings

[0057] Figure 1 is a flowchart of an embodiment of a method for generating anime characters based on feature tags in this application;

[0058] Figure 2 is an implementation flowchart of step S20 in an embodiment of a method for generating anime characters based on feature tags in this application;

[0059] Figure 3 is an implementation flowchart of step S30 in an embodiment of a method for generating anime characters based on feature tags in this application;

[0060] Figure 4 is an implementation flowchart of step S42 in an embodiment of a method for generating anime characters based on feature tags in this application;

[0061] Figure 5 is an implementation flowchart of step S44 in an embodiment of a method for generating anime characters based on feature tags in this application. Detailed Description of the Embodiment

[0062] The following is a further detailed description of this application in conjunction with the attached Figures 1-5 drawings.

[0063] In one embodiment, as Figure 1 shown, this application discloses a method for generating anime characters based on feature tags, which specifically includes the following steps:

[0064] S10: In response to a character generation request sent by the user terminal, obtain generation requirement information from the user terminal;

[0065] In this embodiment, the character generation request is an electrical signal sent by the user terminal to generate an anime character; the generation requirement information is the information sent by the user terminal for describing the character generation requirements, including the basic attributes of the character (such as gender, age, personality, etc.), emotional needs (such as "happy", "angry"), scene needs (such as "battle scene", "daily scene"), etc.;

[0066] Specifically, receive a character generation request sent by a user terminal (such as a web page, APP, etc.), and obtain the user's generation requirement information through an interaction interface. The generation requirement information can be a text description (such as "generate a happy female warrior character") or other forms (such as emotional scores, scenario selections, etc.). For example, the user inputs "I want a happy warrior character in a combat scenario, wearing red armor, and with powerful movements."

[0067] S20: Input the generation requirement information into a pre-trained feature recognition model, so that the feature recognition model generates a character feature vector based on the generation requirement information;

[0068] In this embodiment, the feature recognition model is a pre-trained deep learning model, which is used to extract the basic features of a character from the generation requirement information and generate a character feature vector. Among them, the character feature vector is a multi-dimensional feature representation of the character, including attributes such as personality, clothing, movement, and background;

[0069] Specifically, input the generation requirement information into the feature recognition model, so that the feature recognition model uses parsing technology to extract the basic features of the character and generate a character feature vector. Among them, the character feature vector is a multi-dimensional feature representation of the character, including attributes such as personality, clothing, movement, and background. For example, the feature recognition model extracts the following feature vectors from "a happy warrior character, wearing red armor, and with powerful movements": Personality: happy, brave; Clothing: red armor; Movement: powerful; Background: combat scenario.

[0070] S30: Input the generation requirement information into a pre-trained emotion recognition model, so that the emotion recognition model generates emotion enhancement information;

[0071] In this embodiment, the emotion recognition model is a pre-trained deep learning model, which is used to extract emotion features from the generation requirement information and generate emotion enhancement information. Among them, the emotion enhancement information is refined and supplementary information for the character's emotion features;

[0072] Specifically, input the generation requirement information into the emotion recognition model, so that the emotion recognition model parses and extracts emotion features and generates emotion enhancement information for refining and supplementing the character's emotion features.

[0073] S40: Input the character feature vector and the emotion enhancement information into a pre-trained label matching model, so that the label matching model matches several feature labels based on a pre-stored feature label system and performs emotion enhancement optimization;

[0074] In this embodiment, the label matching model is a pre-trained deep learning model, which is used to match a number of feature labels from a pre-stored feature label system based on the role feature vector and emotion enhancement information, and generate enhanced feature labels through emotion enhancement optimization. Among them, the feature label system includes personality labels, clothing labels, action labels, background labels, etc.; the enhanced feature labels are a set of feature labels optimized by the label matching model, including attributes such as the personality, clothing, actions, and background of the role, and combining emotion enhancement information to perform emotion adaptation and optimization on the labels;

[0075] Specifically, input the role feature vector and emotion enhancement information into the label matching model, so that the label matching model matches a number of feature labels based on the pre-stored feature label system (such as personality labels, clothing labels, action labels, background labels, etc.), and generates enhanced feature labels through emotion enhancement optimization. The emotion enhancement optimization ensures that the labels match the emotion enhancement information and enhances the emotional expression of the role.

[0076] For example, the label matching model matches the following labels according to the role feature vector and emotion enhancement information: Personality label: Brave, Optimistic; Clothing label: Red armor; Action label: Swing sword, Run; Background label: Battlefield;

[0077] After emotion enhancement optimization, generate enhanced feature labels: Personality label: Brave (+0.8 valence); Clothing label: Red armor (high arousal modification); Action label: Swing sword (full of power, high arousal); Background label: Battlefield (battle scene).

[0078] S50: Generate corresponding anime characters based on the enhanced feature labels output by the label matching model.

[0079] In this embodiment, the anime character is a virtual character generated based on the enhanced feature labels, including attributes such as personality characteristics, clothing styles, action performances, and background scenes, and can meet the generation needs of users;

[0080] Specifically, generate corresponding anime characters according to the enhanced feature labels output by the label matching model. Among them, the generated anime characters include attributes such as personality characteristics, clothing styles, action performances, and background scenes, and meet the needs of users and emotion enhancement information.

[0081] For example, based on the enhanced feature labels: Personality label: Brave; Clothing label: Red armor; Action label: Swing sword (; Background label: Battlefield, generate an anime character, specifically shown as: Personality: Brave, Optimistic, and with hair color, hairstyle, and facial form that match this personality; Clothing: Red armor with gold decorations; Action: Wave a huge sword, running posture full of power; Background: On the battlefield, the background is a burning castle and gunpowder smoke.

[0082] In one embodiment, the feature recognition model includes an input encoding layer, a feature decoupling layer, and a vector fusion layer. As Figure 2 shown, step S20 includes the steps of:

[0083] S21: The input encoding layer extracts multimodal information based on the input generation requirement information and constructs a multimodal association matrix;

[0084] S22: The feature decoupling layer extracts independent decoupled features based on the multimodal association matrix. The independent decoupled features include style features, structural features, semantic features, and expansion features;

[0085] S23: The vector fusion layer fuses the independent decoupled features to generate a role feature vector.

[0086] In this embodiment, the input encoding layer is the first layer of the feature recognition model and is used for multimodal parsing and encoding of the generation requirement information. The input encoding layer converts the input text, image, or other forms of requirement information into a unified numerical representation and constructs a multimodal association matrix to capture the correlation between different modalities; the multimodal information is information from multiple data sources, such as text descriptions, image features, voice information, etc.; the multimodal association matrix is a mathematical representation used to describe the correlation and interaction between different modality information. Through the multimodal association matrix, the semantic relationship between different modality information can be captured, so as to more comprehensively understand the generation requirement information; the feature decoupling layer is the second layer of the feature recognition model and is used to extract independent decoupled features from the multimodal association matrix; the independent decoupled features are features of different dimensions separated from the multimodal information. These features are independent of each other and have clear semantic meanings. The independent decoupled features include style features, structural features, semantic features, and expansion features; the vector fusion layer is the third layer of the feature recognition model and is used to fuse the independent decoupled features to generate a unified role feature vector; the role feature vector is the output result of the feature recognition model and is a multi-dimensional vector used to represent the comprehensive features of the role. The role feature vector includes the style, structure, semantics, and other expansion information of the role, providing input for subsequent emotion recognition and label matching.

[0087] Specifically, the input encoding layer performs multimodal parsing on the generation requirement information, extracts information in text, image or other modalities, converts this information into a numerical representation, and constructs a multimodal association matrix to capture the correlation between different modality information. For example, there may be a semantic association between the "warrior" mentioned in the text and the weapon or armor features that may be included in the image; the feature decoupling layer extracts independently decoupled features from the multimodal association matrix. These features include style features (such as the visual style of the character), structural features (such as the body proportions of the character), semantic features (such as the character description of the character), and extended features (such as additional scene or prop information). Through decoupling, different dimensional features can be understood more clearly; the vector fusion layer fuses the independently decoupled features to generate a unified character feature vector. This character feature vector is a numerical representation of the multi-dimensional features of the character, including the style, structure, semantics, and other extended information of the character.

[0088] In one embodiment, the emotion recognition model includes a feature extraction layer, an emotion separation layer, and an enhancement generation layer. As Figure 3 shown, step S30 includes the steps of:

[0089] S31: The feature extraction layer performs multimodal parsing based on the generation requirement information and extracts preliminary emotion features;

[0090] S32: The emotion separation layer performs emotion dimension separation and scenario adaptation optimization based on the preliminary emotion features to extract contextual emotion features;

[0091] S33: The enhancement generation layer generates emotion enhancement information based on the contextual emotion features.

[0092] In this embodiment, the feature extraction layer is the first layer of the emotion recognition model, which is used to perform multi-modal parsing on the generated demand information and extract preliminary emotion features. The feature extraction layer extracts emotion-related features from multi-modal information such as text, images, and speech, serving as the basis for subsequent emotion separation; the emotion separation layer is the second layer of the emotion recognition model, which is used to perform emotion dimension separation and scenario adaptation optimization on the preliminary emotion features. The emotion separation layer decomposes the emotion features into core dimensions such as valence, arousal, and control, and adapts and optimizes the emotion features in combination with the scenario requirements, generating context-aware emotion features; the enhancement generation layer is the third layer of the emotion recognition model, which is used to generate emotion enhancement information based on the context-aware emotion features. The enhancement generation layer generates more expressive and adaptable emotion enhancement information by refining and supplementing the emotion features, providing support for subsequent label matching and role generation; the preliminary emotion features are the original features related to emotions extracted from the generated demand information, including the preliminary emotion tendencies of multi-modal information, but have not undergone dimension separation and scenario adaptation optimization; the context-aware emotion features are emotion features generated based on the preliminary emotion features after emotion dimension separation and scenario adaptation optimization, which are more in line with specific scenario requirements and have higher emotion expressiveness and adaptability; the emotion enhancement information is the emotion refinement result generated based on the context-aware emotion features, including the specific values of emotion dimensions and emotion modification features related to the scenario.

[0093] Specifically, the feature extraction layer performs multi-modal parsing on the generated demand information and extracts preliminary features related to emotions; the emotion separation layer performs emotion dimension separation on the preliminary emotion features, decomposing them into core dimensions such as valence, arousal, and control, and performs scenario adaptation optimization on the emotion features in combination with the scenario requirements (such as "battle scenario") in the generated demand information, generating context-aware emotion features; the enhancement generation layer further refines and supplements the emotion information based on the context-aware emotion features, generating emotion enhancement information. The emotion enhancement information includes the specific values of valence, arousal, and control, as well as emotion modification features related to the scenario, providing emotion input for the subsequent label matching model.

[0094] In one embodiment, step S32 includes the following steps:

[0095] S321: The emotion separation layer extracts an emotion dimension vector based on the preliminary emotion features, and the emotion dimension vector includes valence, arousal, and control;

[0096] S322: The emotion separation layer identifies the scenario requirement information in the generated demand information, and performs weight assignment and context-aware modification extraction on the emotion dimension vector based on the scenario requirement information, thereby generating context-aware emotion features.

[0097] In this embodiment, the emotional dimension vector is the core dimension extracted from the preliminary emotional features to describe the emotional features, including valence, arousal, and dominance. Valence represents the positive or negative degree of emotion, usually ranging from [-1, 1], where negative values indicate negative emotions and positive values indicate positive emotions. Arousal represents the intensity or activation degree of emotion, usually ranging from [0, 1], and the higher the value, the stronger the emotion. Dominance: represents the dominance or sense of control of emotion, usually ranging from [0, 1], and the higher the value, the more dominant the emotion. The scenario requirement information is the description related to a specific scenario in the generated requirement information, such as "battle scenario", "daily scenario", etc. The scenario requirement information is used to guide the adaptation and optimization of emotional features to make the emotional features more in line with the requirements of a specific scenario. The contextualized emotional feature is an emotional feature generated after weight assignment and contextualized modification extraction based on the emotional dimension vector and combined with the scenario requirement information. It is more in line with the specific scenario requirements and has higher emotional expressiveness and adaptability. Weight assignment is to assign different weights to dimensions such as valence, arousal, and dominance in the emotional dimension vector according to the scenario requirement information to highlight the emotional features in a specific scenario. For example, in a "battle scenario", the weights of arousal and dominance may be increased. Contextualized modification extraction is to extract modification features related to the scenario requirement from the emotional dimension vector, such as "aggressiveness", "explosiveness", "gentleness", etc., to further refine and supplement the emotional features.

[0098] Specifically, the emotion separation layer decouples the preliminary emotional features and extracts the emotional dimension vector, including valence, arousal, and dominance. The emotional dimension vector is the core description of the emotional features and can reflect the basic attributes of emotions. The emotion separation layer identifies the scenario requirement information (such as "battle scenario") in the generated requirement information, assigns weights to the emotional dimension vector according to the scenario requirement, and highlights the emotional features in a specific scenario. At the same time, it extracts contextualized modification features to further refine and supplement the emotional features and generates contextualized emotional features.

[0099] In one embodiment, the label matching model includes a label matching layer and an enhancement and optimization layer. Step S40 includes the steps of:

[0100] S41: The label matching layer matches and filters out a candidate label set from the pre-stored feature label system based on the role feature vector;

[0101] S42: The enhancement and optimization layer performs cross-label collaborative optimization based on emotion enhancement information on the candidate label set to generate enhanced feature labels.

[0102] In this embodiment, the label matching layer is the first layer of the label matching model, which is used to match and screen out a candidate label set from a pre-stored feature label system based on the role feature vector. The candidate label set includes character labels, clothing labels, action labels, background labels, etc. The enhancement and optimization layer is the second layer of the label matching model, which is used to perform cross-label collaborative optimization on the candidate label set based on the emotion enhancement information, enhance the adaptability of the labels to the role features and emotion features, and generate enhanced feature labels. The role feature vector is the output result of the feature recognition model, which contains multi-dimensional feature information such as the character, clothing, action, and background of the role, and is used to describe the comprehensive features of the role. The emotion enhancement information is the output result of the emotion recognition model, which contains specific values of emotion dimensions such as valence, arousal, and control degree, as well as emotion modification features related to the scene, and is used to refine and supplement the role emotion features. The candidate label set is a preliminary label set screened from the pre-stored feature label system based on the role feature vector, and includes character labels, clothing labels, action labels, background labels, etc. The enhanced feature label is a label set optimized by the enhancement and optimization layer, which combines the role feature vector and the emotion enhancement information, and has higher emotion adaptability and scene adaptability. The pre-stored feature label system is a pre-defined label library, which contains labels of multiple dimensions, such as: character labels: brave, optimistic, angry, melancholy; clothing labels: red armor, blue robe, black cloak; action labels: swing sword, run, jump; background labels: battlefield, forest, city.

[0103] Specifically, the label matching layer analyzes the role feature vector, extracts features such as the character, clothing, action, and background of the role, and matches and screens out a candidate label set related to the role features from the pre-stored feature label system. The candidate label set is a preliminary label set and has not been optimized by emotion enhancement. The enhancement and optimization layer combines emotion enhancement information (such as emotion dimension values such as valence, arousal, and control degree and contextual modification features), performs cross-label collaborative optimization on the candidate label set, and enhances the adaptability of the labels to the role features and emotion features through cross-label collaborative optimization, generating enhanced feature labels.

[0104] In one embodiment, the candidate label set includes character labels, clothing labels, action labels, and background labels. As Figure 4 shown, step S42 includes the steps:

[0105] S421: The enhancement and optimization layer generates a benchmark emotion matrix based on the character label, and sets emotion deviation thresholds for the clothing label, action label, and background label respectively;

[0106] S422: The enhancement and optimization layer performs corresponding optimization on the clothing label, action label, and background label respectively based on the set emotion deviation thresholds;

[0107] S423: The enhancement and optimization layer generates enhanced feature tags based on the personality tags and the optimized clothing tags, action tags, and background tags.

[0108] In this embodiment, the candidate tag set is a preliminary tag set selected from a pre-stored feature tag system based on the character feature vector, including personality tags, clothing tags, action tags, and background tags; the personality tags are tags describing the character traits of the character, such as "brave", "angry", "optimistic", etc. The personality tag is the core of the character's emotional characteristics and is usually used as a benchmark for emotional optimization; the clothing tag is a tag describing the clothing style of the character, such as "red armor", "blue robe", "black cloak", etc. The emotional optimization of the clothing tag needs to match the emotional characteristics of the personality tag; the action tag is a tag describing the action performance of the character, such as "swinging a sword", "running", "jumping", etc. The emotional optimization of the action tag needs to reflect the emotional intensity and action characteristics of the character; the background tag is a tag describing the scene where the character is located, such as "battlefield", "forest", "city", etc. The emotional optimization of the background tag needs to match the emotional characteristics of the character and the scene requirements; the benchmark emotional matrix is an emotional feature matrix generated based on the personality tag, used to describe the emotional dimensions (valence, arousal, control) of the personality tag and their mutual relationships. The benchmark emotional matrix provides a reference for the emotional optimization of other tags; the emotional deviation threshold is the emotional feature adjustment range set for the clothing tag, action tag, and background tag, used to control the matching degree of the emotional characteristics of these tags with the emotional characteristics of the personality tag. The larger the deviation threshold, the wider the adjustment range of the emotional characteristics; the enhanced feature tags are the tag set optimized by the enhancement and optimization layer, including the optimization results of the personality tag, clothing tag, action tag, and background tag. The enhanced feature tags are more in line with the character's emotional characteristics and scene requirements.

[0109] Specifically, the enhancement and optimization layer generates a benchmark emotional matrix based on the personality tag, describes the emotional dimensions (valence, arousal, control) of the personality tag, and sets emotional deviation thresholds for the clothing tag, action tag, and background tag respectively, used to control the matching range of the emotional characteristics of these tags with the emotional characteristics of the personality tag; the enhancement and optimization layer optimizes the clothing tag, action tag, and background tag according to the benchmark emotional matrix of the personality tag and the set emotional deviation thresholds. During the optimization process, ensure that the emotional characteristics of these tags are consistent with the emotional characteristics of the personality tag, and at the same time reflect the suitability of the character and the scene; the enhancement and optimization layer integrates the personality tag with the optimized clothing tag, action tag, and background tag to generate enhanced feature tags. The enhanced feature tags contain the optimization results of the character's personality, clothing, action, and background, and have higher emotional suitability and scene suitability.

[0110] In one embodiment, after step S423, the following steps are performed:

[0111] S43: Detect the enhanced feature tags based on the preset tag conflict rules;

[0112] S44: When a tag conflict is detected, handle the tag conflict based on the conflict handling strategy.

[0113] In this embodiment, a tag conflict refers to a situation where there are contradictions or inconsistencies in the emotional features, semantic features, or scene adaptability between different tags in the enhanced feature tags. For example, if the personality tag is "angry", but the clothing tag is "white robe" (usually associated with calmness or elegance), it may lead to inconsistent emotional features; the tag conflict rules are a predefined set of rules used to detect whether there are conflicts in the enhanced feature tags, and the tag conflict rules include: Emotional dimension conflict: such as significant deviations in valence, arousal, and control; Semantic conflict: such as the semantic meanings of tags being contradictory to each other (such as "angry" and "gentle"); Scene adaptability conflict: such as tags not matching the scene requirements (such as "peace dove" appearing in a "battle scene"); The conflict handling strategy is the method adopted for the detected tag conflict to eliminate or alleviate the conflict, and the conflict handling strategy includes: Replacing tags: replacing the conflicting tags with tags that are more consistent with the character features and emotional features; Adjusting emotional features: adjusting the emotional dimension values of the conflicting tags to make them consistent with other tags; Priority processing: retaining important tags according to the priority of tags (such as personality tags taking precedence over clothing tags) and adjusting secondary tags; The enhanced feature tags are a set of tags optimized based on the character feature vector and emotional enhancement information, including personality tags, clothing tags, action tags, and background tags. After conflict detection and handling, the enhanced feature tags will be more coordinated and consistent.

[0114] Specifically, the enhancement and optimization layer uses the preset tag conflict rules to detect the enhanced feature tags, identify whether there are conflicts between the tags, and the detection content includes emotional dimension conflicts, semantic conflicts, and scene adaptability conflicts; if a conflict is detected, record the conflict type and the tags involved. The enhancement and optimization layer adjusts or replaces the conflicting tags according to the conflict handling strategy to eliminate or alleviate the conflict, and the handling strategy includes replacing tags, adjusting emotional features, or priority processing.

[0115] In one embodiment, as Figure 5 shown, step S44 includes the steps:

[0116] S441: When a tag conflict is detected, identify its conflict type, and the conflict type includes hard conflicts and soft conflicts;

[0117] S442: If the conflict type is a hard conflict, calculate the deviation degree of the enhanced feature tags associated with the hard conflict from the benchmark emotion matrix, identify the abnormal tags and their tag types, and match the corresponding preset tag replacement strategy based on the tag types;

[0118] S443: If the conflict type is a soft conflict, identify the label type of the enhanced feature label associated with the soft conflict, and match the corresponding preset reconciliation strategy based on the label type.

[0119] In this embodiment, a hard conflict is a situation where there are significant contradictions or inconsistencies between enhanced feature labels, which cannot be resolved by simple adjustment. A hard conflict needs to be resolved by replacing the label; a soft conflict is a situation where there are minor contradictions or inconsistencies between enhanced feature labels, which can be resolved by adjusting the emotional features or semantic modifications. For example, the arousal level of the clothing label "blue robe" is slightly lower than that of the personality label "angry", and this conflict can be alleviated by adjusting the emotional dimension value; the deviation degree is the degree of difference between the emotional dimension value of the enhanced feature label and the emotional dimension value of the benchmark emotional matrix. The deviation degree is used to quantify the emotional consistency between the label and the personality label. The higher the deviation degree, the more significant the conflict; the label replacement strategy is a processing method preset for hard conflicts, which is used to replace the conflicting label with a label that is more consistent with the character features and emotional features. The label replacement strategy is usually based on the label type (such as clothing label, action label) and the scenario requirements; the reconciliation strategy is a processing method preset for soft conflicts, which is used to alleviate the conflict by adjusting the emotional dimension value or semantic modification of the label to make it more consistent with the emotional features of other labels; the label type is the category of the enhanced feature label, including personality labels, clothing labels, action labels, and background labels. Different types of labels may apply different strategies in conflict processing.

[0120] Specifically, after detecting a label conflict, the enhancement and optimization layer first identifies the type of conflict. The conflict type is divided into hard conflicts and soft conflicts; for hard conflicts, the enhancement and optimization layer calculates the emotional dimension deviation degree between the conflicting label and the benchmark emotional matrix, quantifies the severity of the conflict, and identifies the abnormal label and its type, and matches the preset label replacement strategy based on the label type, and replaces the abnormal label with a more appropriate label; for soft conflicts, the enhancement and optimization layer identifies the type of the conflicting label, and matches the preset reconciliation strategy based on the label type, and alleviates the conflict by adjusting the emotional dimension value or semantic modification to make it more consistent with the emotional features of other labels.

[0121] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0122] In one embodiment, a feature label-based anime character generation system is provided. This feature label-based anime character generation system corresponds one-to-one with the feature label-based anime character generation method in the above embodiment.

[0123] Anime character generation system based on feature tags, comprising:

[0124] A requirement acquisition module, configured to obtain generation requirement information from the user terminal in response to a character generation request sent by the user terminal;

[0125] A first input module, configured to input the generation requirement information into a pre-trained feature recognition model, so that the feature recognition model generates a character feature vector based on the generation requirement information;

[0126] A second input module, configured to input the generation requirement information into a pre-trained emotion recognition model, so that the emotion recognition model generates emotion enhancement information;

[0127] A third input module, configured to input the character feature vector and the emotion enhancement information into a pre-trained label matching model, so that the label matching model matches a number of feature tags based on a pre-stored feature tag system and performs emotion enhancement optimization;

[0128] A character generation module, configured to generate a corresponding anime character based on the enhanced feature tags output by the label matching model.

[0129] For the specific limitations of an anime character generation system based on feature tags, reference can be made to the limitations of an anime character generation method based on feature tags in the foregoing text, which will not be elaborated here. Each module in the above-mentioned anime character generation system based on feature tags can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0130] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for generating anime characters based on feature tags, characterized in that: Including the steps: In response to a role generation request sent by the client, obtain generation requirement information from the client; Input the generation requirement information into a pre-trained feature recognition model, so that the feature recognition model generates a role feature vector based on the generation requirement information; Input the generation requirement information into a pre-trained emotion recognition model, so that the emotion recognition model generates emotion enhancement information; Input the role feature vector and the emotion enhancement information into a pre-trained label matching model, so that the label matching model matches several feature labels based on a pre-stored feature label system and performs emotion enhancement optimization; Generate a corresponding anime role based on the enhanced feature labels output by the label matching model; The emotion recognition model includes a feature extraction layer, an emotion separation layer, and an enhancement generation layer. The step of inputting the generation requirement information into the pre-trained emotion recognition model so that the emotion recognition model generates emotion enhancement information includes the steps: The feature extraction layer performs multimodal parsing based on the generation requirement information and extracts preliminary emotion features; The emotion separation layer performs emotion dimension separation and scenario adaptation optimization based on the preliminary emotion features to extract contextual emotion features; The enhancement generation layer generates emotion enhancement information based on the contextual emotion features; The step of the emotion separation layer performing emotion dimension separation and scenario adaptation optimization based on the preliminary emotion features to extract contextual emotion features includes the steps: The emotion separation layer extracts an emotion dimension vector based on the preliminary emotion features, and the emotion dimension vector includes valence, arousal, and control; The emotion separation layer identifies the scenario requirement information in the generation requirement information, and performs weight assignment and contextual modification extraction on the emotion dimension vector based on the scenario requirement information, so as to generate contextual emotion features; The label matching model includes a label matching layer and an enhancement optimization layer. The step of inputting the role feature vector and the emotion enhancement information into the pre-trained label matching model so that the label matching model matches several feature labels based on a pre-stored feature label system and performs emotion enhancement optimization includes the steps: The label matching layer matches and filters a candidate label set from the pre-stored feature label system based on the role feature vector; The enhancement optimization layer performs cross-label collaborative optimization based on emotion enhancement on the candidate label set, so as to generate enhanced feature labels.

2. The method for generating an anime character based on feature tags according to claim 1, wherein: The feature recognition model includes an input encoding layer, a feature decoupling layer, and a vector fusion layer. The step of inputting the generation requirement information into the pre-trained feature recognition model so that the feature recognition model generates a role feature vector based on the generation requirement information includes the steps: The input encoding layer extracts multimodal information based on the input generation requirement information and constructs a multimodal association matrix; The feature decoupling layer extracts independent decoupled features based on the multimodal association matrix, and the independent decoupled features include style features, structural features, semantic features, and extended features; The vector fusion layer fuses the independent decoupled features to generate a role feature vector.

3. A method for generating an anime character based on feature tags according to claim 1, characterized in that: The candidate tag set includes personality tags, clothing tags, action tags, and background tags. The enhancement and optimization layer performs sentiment-enhanced cross-tag collaborative optimization on the candidate tag set based on sentiment enhancement information to generate enhanced feature tags. The steps include: The enhancement and optimization layer generates a benchmark sentiment matrix based on personality tags, and sets sentiment deviation thresholds for clothing tags, action tags, and background tags respectively; The enhancement and optimization layer performs corresponding optimization on clothing tags, action tags, and background tags respectively based on the set sentiment deviation thresholds; The enhancement and optimization layer generates enhanced feature tags based on personality tags and the optimized clothing tags, action tags, and background tags.

4. The method for generating an anime character based on feature tags according to claim 3, wherein: After the step where the enhancement and optimization layer generates enhanced feature tags based on personality tags and the optimized clothing tags, action tags, and background tags, the following steps are executed: Detect the enhanced feature tags based on pre-set tag conflict rules; When tag conflicts are detected, handle the tag conflicts based on conflict handling strategies.

5. A method for generating an anime character based on feature tags according to claim 4, characterized in that: The step of handling tag conflicts based on conflict handling strategies when tag conflicts are detected includes the steps: When tag conflicts are detected, identify their conflict types, and the conflict types include hard conflicts and soft conflicts; If the conflict type is a hard conflict, calculate the deviation degree between the enhanced feature tags associated with the hard conflict and the benchmark sentiment matrix, identify the abnormal tags and their tag types, and match the corresponding pre-set tag replacement strategies based on the tag types; If the conflict type is a soft conflict, identify the tag types of the enhanced feature tags associated with the soft conflict, and match the corresponding pre-set reconciliation strategies based on the tag types.

6. An anime character generation system based on feature tags, which is used for the steps of an anime character generation method based on feature tags as described in any one of claims 1-5, and is characterized in that, Including: A requirement acquisition module for obtaining generation requirement information from the user terminal in response to a role generation request issued by the user terminal; A first input module for inputting the generation requirement information into a pre-trained feature recognition model, so that the feature recognition model generates a role feature vector based on the generation requirement information; A second input module for inputting the generation requirement information into a pre-trained sentiment recognition model, so that the sentiment recognition model generates sentiment enhancement information; A third input module for inputting the role feature vector and the sentiment enhancement information into a pre-trained tag matching model, so that the tag matching model matches a number of feature tags based on a pre-stored feature tag system and performs sentiment enhancement optimization; A role generation module for generating corresponding anime characters based on the enhanced feature tags output by the tag matching model.

Citation Information

Patent Citations

  • Internet video script role emotion recognition method based on big data

    CN116821333A

  • Method and system for generating synthesis voice using style tag represented by natural language

    US20240105160A1