Virtual character generation method and device based on AI, terminal and storage medium
Through multimodal fusion model and knowledge graph technology, multimodal features are integrated and semantic enhancement and logical reasoning are carried out, the intelligent bottleneck of the virtual character generation system is solved, and the generation quality and user experience are improved.
Patent Information
- Application Number
- CN202510535441.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing virtual role generation system is difficult to effectively handle multimodal inputs, and lacks deep semantic understanding and logical reasoning capabilities, resulting in deviations from user intentions, attribute contradictions and worldviews, and it is difficult to adapt to the rule constraints of different application scenarios.
Multimodal fusion model and knowledge graph technology are used to integrate multimodal features through pre-trained multimodal fusion model, and use knowledge graphs to perform semantic enhancement and logical reasoning to generate enhanced role features.
It improves the quality and user experience of virtual character generation, ensures the consistency and rationality of the generated characters in abilities and background settings, adapts to the rules and constraints of different scenarios, and enhances the realism and scene adaptability of character generation.
Smart Images

Figure CN120472058A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of artificial intelligence, and in particular to an AI-based virtual character generation method, device, terminal, and storage medium. Background Art
[0002] As a core component of digital content creation, avatar generation technology plays a crucial role in pan-entertainment industries, including game development, film and television production, and virtual reality. With the rise of the metaverse concept and the widespread use of digital humans, this technology has gradually expanded into a wider range of fields, including education, healthcare, and social interaction.
[0003] In the field of virtual character generation, existing technologies typically use parametric templates or prefabricated component libraries to build virtual characters, manually setting attribute ranges and association rules to achieve character generation. These systems are effective for simple character creation, but their limitations are becoming increasingly prominent as applications such as games and virtual reality demand ever-increasing character personalization and diversity.
[0004] First, existing virtual character generation systems have difficulty effectively processing multimodal input (such as text descriptions combined with reference images), and each modal information is often processed in isolation, resulting in a deviation between the generated results and the user's intention; second, existing virtual character generation systems lack deep semantic understanding and logical reasoning capabilities, and the generated virtual characters may have contradictory attributes or inconsistent worldviews, such as mistakenly adding a "fire attack" ability to the "underwater creature" character; third, the knowledge representation method of existing virtual character generation systems is rigid, and it is difficult to dynamically adapt to the rule constraints of different application scenarios, resulting in the generated characters lacking realism and sense of immersion.
[0005] The above-mentioned defects seriously restrict the intelligence level and user experience of the virtual character generation system. There is an urgent need for a new solution that can deeply integrate multimodal understanding, knowledge reasoning and generative AI technology. Summary of the Invention
[0006] In order to solve the above-mentioned defects, the present application provides an AI-based virtual character generation method, device, terminal and storage medium.
[0007] The above-mentioned invention objective of this application is achieved through the following technical solutions:
[0008] A method for generating a virtual character based on AI, comprising the steps of:
[0009] In response to a character generation instruction input by a user terminal, corresponding multimodal features are obtained;
[0010] Input the multimodal features into a pre-trained multimodal fusion model so that the multimodal fusion model outputs the generated character features;
[0011] Perform semantic enhancement and logical reasoning on the generated character features through knowledge graph to obtain enhanced character features;
[0012] Generating initial virtual character data based on the enhanced character features and sending it to the user end for adjustment;
[0013] When the optimized virtual character data fed back by the user is received, a corresponding virtual character is generated.
[0014] By adopting the above technical solution, multimodal features are obtained in response to the character generation instructions input by the user end, and the pre-trained multimodal fusion model is used to integrate the features to achieve collaborative analysis of multimodal input, which has the effect of eliminating the problem that the generation results of the traditional system deviate from the user's intention due to the isolated processing of different modal information; further, the generated character features are semantically enhanced and logically reasoned through the knowledge graph to solve the problem of inconsistent character attributes or inconsistent world views caused by the lack of deep semantic understanding in the existing technology, so that the generated virtual characters maintain rationality and consistency in ability and background settings, and at the same time overcome the defect that the traditional system is difficult to adapt to the rules and constraints of different scenarios due to the rigid knowledge representation, and significantly improve the realism and scene adaptability of the generated characters; this application effectively solves the bottleneck of the existing virtual character generation system in the intelligence level by combining multimodal feature extraction, knowledge graph reasoning and generative AI technology, and has the effect of improving the quality of virtual character generation and user experience.
[0015] In a preferred example, the present application may be further configured as follows: the step of obtaining corresponding multimodal features in response to a character generation instruction input by the user terminal includes the following steps:
[0016] Identifying the type of the character generation instruction, wherein the instruction type includes text instructions, image instructions, video instructions, and voice instructions;
[0017] Multimodal features are extracted from the character generation instructions based on the instruction type to obtain multimodal features.
[0018] By adopting the above-mentioned technical solution, the character generation instructions are identified by instruction type, including text instructions, image instructions, video instructions and voice instructions, and corresponding multimodal feature extraction processing is performed based on different instruction types to achieve compatibility and efficient analysis of multimodal input methods; this application accurately captures the virtual character creation intentions expressed by users through different media through intelligent instruction recognition and feature extraction, and realizes automatic feature encoding of multimodal input data, which has the effect of improving the accuracy of virtual character generation and lowering the user usage threshold, thereby expanding the application scenarios and user groups of the system.
[0019] In a preferred example, the present application can be further configured as follows: the multimodal fusion model includes a feature encoding layer, a cross-modal interaction layer, and a feature decoding layer; the step of inputting the multimodal features into a pre-trained multimodal fusion model so that the multimodal fusion model outputs generated character features includes the following steps:
[0020] The feature encoding layer performs modal encoding on multimodal features to extract multimodal semantic representations;
[0021] The cross-modal interaction layer performs multimodal alignment and cross-modal fusion on the multimodal semantic representations to generate a semantic space vector of a unified role;
[0022] The feature decoding layer generates and outputs the generated role features based on the semantic space vector of the unified role.
[0023] By adopting the above technical solution, a multimodal fusion model including a feature encoding layer, a cross-modal interaction layer and a feature decoding layer is constructed to achieve in-depth understanding and fusion of multi-source input data; the feature encoding layer performs modal encoding on the multimodal features to ensure that the semantic information of each modality such as text and image is fully extracted; the cross-modal interaction layer generates a semantic space vector of a unified role through a multimodal alignment and cross-modal fusion mechanism; the feature decoding layer generates and outputs generated role features based on the semantic space vector of the unified role; this application realizes the conversion from the multimodal features of the original input to the role features through a hierarchical and progressive multimodal fusion model processing architecture, which has the effect of improving the accuracy of multimodal data fusion and the consistency of generated role features, and can more accurately capture the complex semantic associations in user input.
[0024] In a preferred example, the present application may be further configured as follows: the step of performing semantic enhancement and logical reasoning on the generated character features through the knowledge graph to obtain enhanced character features includes the following steps:
[0025] Construct a knowledge graph of associated role attributes, where nodes are discretized role attributes and their characteristic values, and edges are the associations and numerical constraints between role attributes;
[0026] Map the generated role features to the embedding space of the knowledge graph and calculate the similarity distribution between them and each node;
[0027] Based on the knowledge graph topology and semantic association path, the generated role features are attribute-extended and semantically associated to generate semantically enhanced extended role features;
[0028] Constraint verification is performed based on predefined inference rules to generate enhanced role features.
[0029] By adopting the above technical solution, a knowledge graph of associated role attributes is constructed and a semantic enhancement and reasoning mechanism based on the graph is designed to realize intelligent optimization and verification of virtual character features; specifically, the discrete role attributes and their association relationships are structured and stored in the knowledge graph, and the generated role features are mapped to the graph embedding space and the similarity distribution is calculated to achieve accurate docking of role attributes with domain knowledge; based on the knowledge graph topology and semantic path, the generated role features are subjected to attribute extension and semantic association reasoning to generate semantically enhanced extended role features; constraint verification is performed through predefined reasoning rules to ensure the logical consistency between role attributes and generate enhanced role features; this application realizes the integration and application of domain knowledge in the virtual character generation process through a knowledge graph-driven feature enhancement method, which has the effect of significantly improving the semantic richness and logical rationality of the generated character, so that the final generated virtual character not only has stronger background setting integrity, but also avoids attribute conflicts, and meets the requirements of different application scenarios for character authenticity and consistency.
[0030] In a preferred example, the present application may be further configured as follows: the step of performing attribute extension and semantic association on the generated role features based on the knowledge graph and the semantic association path to generate semantically enhanced extended role features includes the following steps:
[0031] Identify abstract attribute nodes in the knowledge graph and fill them with instantiations;
[0032] Based on the semantic association path, the generated role features are propagated through multiple hops with a predefined number of hops to obtain an extended feature subgraph;
[0033] Based on the knowledge graph and semantic association path, a semantic similarity matrix between node attributes is constructed, and cross-domain associations are established to obtain a feature association graph;
[0034] Semantically enhanced extended role features are generated based on the extended feature subgraph and feature association graph.
[0035] By adopting the above technical solution, the semantic extension and association mining of generated character features are realized by instantiating and filling abstract attribute nodes in the knowledge graph and the multi-hop feature propagation mechanism; the abstract concept nodes in the knowledge graph are converted into specific instances to provide a richer entity expression for the character features; semantic diffusion is carried out in the knowledge graph through multi-hop feature propagation with a predefined number of hops to construct an extended feature subgraph containing core features and their associated attributes; at the same time, cross-domain associations are established based on the semantic similarity matrix to form a feature association graph covering the multi-dimensional attributes of the character; this application realizes the expansion from core features to peripheral attributes through a hierarchical knowledge graph reasoning method, which has the effect of significantly enhancing the depth and breadth of the character semantic expression, so that the generated virtual character not only has basic attribute features, but also can automatically obtain related derivative features, thereby improving the integrity and scalability of the generated virtual character.
[0036] In a preferred example, the present application can be further configured as follows: the step of identifying abstract attribute nodes in the knowledge graph and performing instantiation and filling includes the following steps:
[0037] Identify all nodes in the knowledge graph, filter out conceptual nodes that are not bound to specific feature values, and identify parent nodes with incomplete attributes that have inheritance relationships;
[0038] Identify the basic features of the role in the generated role features, extract the scene context features, and generate a set of instantiation constraints;
[0039] The filtered conceptual nodes and the identified parent nodes are instantiated and populated based on the instantiation constraint set.
[0040] By adopting the above technical solution, conceptual nodes that need to be instantiated and incomplete attribute parent nodes with inheritance relationships in the knowledge graph structure are identified and screened; then the basic role features and scene context information in the generated role features are identified and extracted to construct a set of instantiation constraints; based on the set of instantiation constraints, the screened conceptual nodes and the identified parent nodes are instantiated and filled; this application realizes the conversion from abstract concepts to specific role attributes through a constraint-based instantiation filling mechanism, which has the effect of significantly improving the personalization of role generation and scene adaptability, so that the generated virtual characters not only get rid of the limitations of traditional template generation, but also can automatically adjust the attribute details according to the specific application scenario, to ensure that each generated character has a unique feature combination and attribute configuration that meets the scene requirements.
[0041] The second object of the present invention is achieved through the following technical solutions:
[0042] An AI-based virtual character generation device, comprising:
[0043] A feature acquisition module, configured to acquire corresponding multimodal features in response to a character generation instruction input by a user terminal;
[0044] An input module is used to input multimodal features into a pre-trained multimodal fusion model so that the multimodal fusion model outputs generated character features;
[0045] Enhanced feature generation module, used to perform semantic enhancement and logical reasoning on generated character features through knowledge graph to obtain enhanced character features;
[0046] An output module, for generating initial virtual character data based on the enhanced character features and sending the data to the user end for adjustment;
[0047] The character generation module is used to generate a corresponding virtual character when receiving the optimized virtual character data fed back by the user end.
[0048] By adopting the above technical solution, the feature acquisition module is used to obtain corresponding multimodal features in response to the character generation instruction input by the user terminal; the input module is used to input the multimodal features into a pre-trained multimodal fusion model, so that the multimodal fusion model outputs the generated character features; the enhanced feature generation module is used to perform semantic enhancement and logical reasoning on the generated character features through the knowledge graph to obtain enhanced character features; the output module is used to generate initial virtual character data based on the enhanced character features and send it to the user terminal for adjustment; the character generation module is used to generate the corresponding virtual character when receiving the optimized virtual character data fed back by the user terminal.
[0049] In a preferred example, the present application can be further configured as follows: the feature acquisition module includes:
[0050] A type identification submodule is used to identify the instruction type of the character generation instruction, wherein the instruction type includes text instruction, image instruction, video instruction and voice instruction;
[0051] The feature extraction submodule is used to extract multimodal features of the character generation instructions based on the instruction type, thereby obtaining multimodal features.
[0052] By adopting the above technical solution, the type identification submodule is used to identify the instruction type of the character generation instruction, and the instruction type includes text instructions, image instructions, video instructions and voice instructions; the feature extraction submodule is used to extract multimodal features of the character generation instruction based on the instruction type, thereby obtaining multimodal features.
[0053] In summary, this application includes at least one of the following beneficial technical effects:
[0054] 1. This application effectively addresses the bottleneck of the existing virtual character generation system in terms of intelligence level by combining multimodal feature extraction, knowledge graph reasoning and generative AI technology, and has the effect of improving the quality of virtual character generation and user experience;
[0055] 2. This application uses intelligent command recognition and feature extraction to accurately capture the user's creative intent for virtual characters expressed through different media, achieving automated feature encoding of multimodal input data. This improves the accuracy of virtual character generation and lowers the user barrier to entry, expanding the system's application scenarios and user base.
[0056] 3. This application uses a hierarchical and progressive multimodal fusion model processing architecture to achieve the conversion from the original input multimodal features to character features. This has the effect of improving the accuracy of multimodal data fusion and the consistency of generated character features, and can more accurately capture the complex semantic associations in user input;
[0057] 4. This application uses a knowledge graph-driven feature enhancement method to integrate and apply domain knowledge in the virtual character generation process, significantly improving the semantic richness and logical rationality of the generated characters. This ensures that the final generated virtual characters not only have stronger background setting integrity but also avoid attribute conflicts, meeting the requirements of different application scenarios for character authenticity and consistency.
[0058] 5. This application uses a hierarchical knowledge graph reasoning method to expand from core features to peripheral attributes, significantly enhancing the depth and breadth of character semantic expression. This allows the generated virtual character to not only possess basic attribute features but also automatically acquire related derivative features, thereby improving the integrity and scalability of the generated virtual character.
[0059] 6. This application realizes the transformation from abstract concepts to specific character attributes through a constraint-based instantiation filling mechanism, which has the effect of significantly improving the personalization and scene adaptability of character generation. The generated virtual characters not only break away from the limitations of traditional template generation, but also automatically adjust attribute details according to specific application scenarios, ensuring that each generated character has a unique feature combination and attribute configuration that meets the scene requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 This is a flow chart of an embodiment of an AI-based virtual character generation method of the present application;
[0061] Figure 2 This is a flowchart for implementing step S20 in an embodiment of an AI-based virtual character generation method of the present application;
[0062] Figure 3This is a flowchart of an implementation of step S30 in an embodiment of an AI-based virtual character generation method of the present application;
[0063] Figure 4 This is a flowchart for implementing step S33 in an embodiment of an AI-based virtual character generation method of the present application;
[0064] Figure 5 This is an implementation flowchart of step S331 in an embodiment of an AI-based virtual character generation method of the present application. DETAILED DESCRIPTION
[0065] The following is combined with Figure 1-5 This application is described in further detail.
[0066] In one embodiment, if Figure 1 As shown, the present application discloses an AI-based virtual character generation method, which specifically includes the following steps:
[0067] S10: Responding to a character generation instruction input by the user terminal, obtaining corresponding multimodal features;
[0068] In this embodiment, the character generation instruction is an electronic signal instruction including character generation requirements issued by the user end to generate a virtual character; the multimodal feature is a feature representation extracted from input data in different forms (modalities) (such as text, image, video, voice, etc.). For example, text descriptions can be converted into semantic vectors, images can be used to extract visual features, and voice can be converted into text or acoustic features.
[0069] Specifically, it receives creative instructions input by users through text, images, voice, etc. (such as "generate an elf warrior who is good at fire magic" or upload a character sketch), preprocesses the input data, and extracts corresponding multimodal features (such as semantic vectors of text, visual features of images, etc.).
[0070] S20: inputting the multimodal features into a pre-trained multimodal fusion model, so that the multimodal fusion model outputs generated character features;
[0071] In this embodiment, the multimodal fusion model is a pre-trained deep learning model that can integrate features from different modalities and output a unified feature representation; the generated character features are structured character data features output by the multimodal fusion model, which may include basic attributes (race, occupation, age, etc.), appearance features (hairstyle, clothing, weapons, etc.), ability settings (skills, special abilities, etc.), behavioral features (action style, combat method, etc.), etc.;
[0072] Specifically, the multimodal fusion model integrates multimodal features, eliminates information conflicts (such as the deviation between text description and image content), and generates a unified character feature representation, such as "elf race + warrior profession + fire magic".
[0073] S30: Perform semantic enhancement and logical reasoning on the generated character features through the knowledge graph to obtain enhanced character features;
[0074] In this embodiment, the knowledge graph is a structured role knowledge representation database consisting of nodes (entities or attributes) and edges (relationships or constraints); semantic enhancement is the process of semantically expanding and correcting the generated role feature representation through the knowledge graph; logical reasoning is the formal process of deriving new conclusions from known premises on the knowledge graph, including attribute consistency verification (ensuring that role attributes do not violate domain constraints), semantic association deduction (discovering implicit role feature associations), and conflict resolution (detecting and resolving mutually exclusive attribute combinations); enhanced role features are role features with enhanced attributes obtained through semantic enhancement of generated role features based on the knowledge graph and logical reasoning;
[0075] Specifically, semantic enhancement and logical reasoning are performed on the generated character features through the knowledge graph. The semantic enhancement process is to expand the character attributes based on the knowledge graph, such as adding "agility", "natural affinity" and other related characteristics to "elf warrior"; the logical reasoning process includes attribute consistency verification (ensuring that the character attributes do not violate domain constraints), semantic association deduction (discovering implicit character feature associations) and conflict resolution (detecting and resolving mutually exclusive attribute combinations).
[0076] S40: generating initial virtual character data based on the enhanced character features and sending the data to the user terminal for adjustment;
[0077] In this embodiment, the initial virtual character data is an initial structured digital representation describing the virtual character generated based on the enhanced character features;
[0078] Specifically, the character features will be enhanced to generate specific initial virtual character data, including generating a character model (such as 3D appearance, skill list), which will be provided to users for preview and modification.
[0079] S50: When receiving the optimized virtual character data fed back by the user terminal, generating a corresponding virtual character.
[0080] In this embodiment, the optimized virtual character data is an optimized structured digital representation describing the virtual character, which is obtained by the user confirming, adjusting and optimizing the initial virtual character data;
[0081] Specifically, the corresponding final virtual character is generated based on the optimized virtual character data (such as modified skills or appearance) fed back by the user after adjustment.
[0082] In one embodiment, step S10 includes the steps of:
[0083] S11: Identifying the instruction type of the character generation instruction, where the instruction type includes text instruction, image instruction, video instruction, and voice instruction;
[0084] S12: Perform multimodal feature extraction on the character generation instruction based on the instruction type, thereby obtaining multimodal features.
[0085] In this embodiment, instruction type recognition is the process of classifying the modal category of the user input instruction, and its output space is a discrete set of modal labels; text instructions are character generation requirements expressed in the form of natural language character sequences, and their technical characteristics include: degree of structure (from free text to semi-structured templates), semantic density (the amount of semantic information contained in a unit character), context dependence (the degree of dependence on background knowledge), etc.; image instructions are input forms that convey the visual features of the character through a two-dimensional pixel array, and their key attributes include: modal specificity (including visual features such as color, texture, and shape), semantic abstraction, etc. (the degree of representation from concrete to abstract), style indicativeness (implicit artistic style information), etc.; video instructions are dynamic visual inputs consisting of a sequence of time-sequential image frames. Their characteristics that distinguish them from image instructions include: temporal information (dynamic features such as movements and expressions), inter-frame correlation (semantic continuity between adjacent frames), and multimodality (possibly including a synchronized audio track). Voice instructions are oral character descriptions conveyed through acoustic signals. Core features include: voice-text bimodality (including both phoneme and prosodic information), non-text elements (suprasegmental features such as intonation and stress), and real-time interactivity (supporting streaming processing).
[0086] Specifically, the character generation instructions are identified by instruction type, including text instructions, image instructions, video instructions and voice instructions, and corresponding multimodal feature extraction processing is performed based on different instruction types to achieve compatibility and efficient analysis of multimodal input methods.
[0087] In one embodiment, the multimodal fusion model includes a feature encoding layer, a cross-modal interaction layer, and a feature decoding layer. Figure 2 As shown, step S20 includes the steps of:
[0088] S21: The feature encoding layer performs modal encoding on the multimodal features to extract multimodal semantic representations;
[0089] S22: The cross-modal interaction layer performs multimodal alignment and cross-modal fusion on the multimodal semantic representations to generate a semantic space vector of a unified role;
[0090] S23: The feature decoding layer generates and outputs the generated role features based on the semantic space vector of the unified role.
[0091] In this embodiment, the feature encoding layer is the first layer of the multimodal fusion model, which is used to perform modal encoding on multimodal features to extract multimodal semantic representations; the multimodal semantic representation is a set of feature vectors obtained after processing by the feature encoding layer that retains the original semantic information; the cross-modal interaction layer is the second layer of the multimodal fusion model, which is used to perform multimodal alignment and cross-modal fusion on the multimodal semantic representations to generate a semantic space vector of a unified role; multimodal alignment is the process of aligning different modal features in the semantic space, which is achieved by minimizing contrast loss; cross-modal fusion is the process of integrating the aligned multimodal features into a unified representation, and typical methods include vector concatenation, weighted averaging, and tensor multiplication; the semantic space vector is the vector that represents the complete semantics of the role after fusion, and its characteristics include decoupling (different dimensions correspond to different semantic factors), combination (support vector arithmetic to achieve semantic editing), and continuity (adjacent vectors correspond to similar roles); the feature decoding layer is the third layer of the multimodal fusion model, which is used to generate and output the generated role features based on the semantic space vector of the unified role;
[0092] Specifically, a multimodal fusion model is constructed, which includes a feature encoding layer, a cross-modal interaction layer, and a feature decoding layer, to achieve in-depth understanding and fusion of multi-source input data; the feature encoding layer performs modal encoding on multimodal features to ensure that the semantic information of each modality such as text and image is fully extracted; the cross-modal interaction layer generates a semantic space vector of a unified role through multimodal alignment and cross-modal fusion mechanisms; the feature decoding layer generates and outputs the generated role features based on the semantic space vector of the unified role.
[0093] In one embodiment, if Figure 3 As shown, step S30 includes the steps of:
[0094] S31: Construct a knowledge graph of associated role attributes, where nodes are discretized role attributes and their characteristic values, and edges are the associations and value constraints between role attributes;
[0095] S32: Mapping the generated role features to the embedding space of the knowledge graph and calculating the similarity distribution between them and each node;
[0096] S33: Based on the knowledge graph topology and semantic association path, the generated role features are attribute-extended and semantically associated to generate semantically enhanced extended role features;
[0097] S34: Perform constraint verification based on predefined inference rules to generate enhanced role features.
[0098] In this embodiment, the character attribute knowledge graph is a database network that stores all relevant attributes of virtual characters, wherein the nodes are discretized character attributes, including basic attribute nodes (such as "race corresponds to elves", "profession corresponds to mages"), composite attribute nodes (such as "elf mage suit"), numerical attribute nodes (such as "attack power = 85"), etc.; the edges are the association relationships and numerical constraint relationships between character attributes, including inclusion relationships ("weapons correspond to long swords"), mutual exclusion relationships ("underwater breathing and flame manipulation are mutually exclusive"), strength relationships ("metal armor corresponds to defense +30%"), etc.; embedding space mapping is the process of converting the character features described in text into digital coordinates of the knowledge graph; similarity distribution is the process of calculating and generating a matching score table between the character and all attributes in the knowledge base; the topological structure is the connection method of attributes in the knowledge graph, which may include: star structure, that is, the core attribute (such as "profession") radiates to connect multiple sub-attributes, chain structure, Structure, i.e., continuous association (e.g., "Warrior → Heavy Armor → Slow Movement"), and network structure, i.e., the cross-influence of multiple attributes (e.g., the combination of "Race + Occupation + Equipment"). Semantic association paths are thought chains connecting related attributes in the knowledge graph, including strengthening paths (e.g., "Elf → Agility → Increased Dodge"), weakening paths (e.g., "Heavy Armor → Speed Decrease → Reduced Dodge"), and transformation paths (e.g., "Fire Magic → Water → Steam Damage"). Attribute expansion involves adding relevant details based on existing features, for example: Necessary attributes: automatically adding "Navigation Skills" when selecting "Pirate", recommended attributes: selecting "Mage" suggests "Elemental Affinity", environmental attributes: adding "Cold Resistance" based on the "Snow" scene, etc.; inference rules are pre-set logical checking principles, for example: common sense rules: "Fish characters cannot equip flame weapons", balance rules: "High attack power requires low defense", and style rules: "Cyberpunk style prohibits magic spells", etc.
[0099] Specifically, a knowledge graph of associated role attributes is constructed and a graph-based semantic enhancement and reasoning mechanism is designed to achieve intelligent optimization and verification of virtual role features. Specifically, the discrete role attributes and their association relationships are structured and stored in the knowledge graph. By mapping the generated role features to the graph embedding space and calculating the similarity distribution, the role attributes and domain knowledge are accurately connected. Based on the knowledge graph topology and semantic path, the generated role features are extended with attributes and semantic association reasoning to generate semantically enhanced extended role features. Constraint verification is performed through predefined reasoning rules to ensure the logical consistency between role attributes and generate enhanced role features.
[0100] In one embodiment, if Figure 4 As shown, step S33 includes the steps of:
[0101] S331: Identify abstract attribute nodes in the knowledge graph and perform instantiation and filling;
[0102] S332: performing multi-hop feature propagation of a predefined number of hops on the generated role features based on the semantic association path to obtain an extended feature subgraph;
[0103] S333: Construct a semantic similarity matrix between node attributes based on the knowledge graph and semantic association paths, establish cross-domain associations, and obtain a feature association graph;
[0104] S334: Generate semantically enhanced extended role features based on the extended feature subgraph and the feature association graph.
[0105] In this embodiment, an abstract attribute node is a conceptual node in the knowledge graph that is not bound to a specific value, representing a class of instantiable attribute templates, such as: a profession node (melee profession, and no specific profession is specified), a skill node (elemental magic, and no element type is specified), an equipment node (heavy armor, and no material is specified), etc.; instantiation filling is the process of assigning specific feature values to abstract nodes; multi-hop feature propagation is the gradual expansion of features along the relationship chain in the knowledge graph, and its propagation mechanism includes: single-hop propagation, that is, direct association of attributes (such as "samurai → katana"), double-hop propagation, that is, indirect association (such as "katana → forging process → quenching technology"), and restricted propagation, that is, presetting the maximum number of hops, etc.; an extended feature subgraph is a local knowledge graph fragment obtained through propagation; a semantic similarity matrix is a matrix that quantifies the strength of association between attributes; a cross-domain association is a bridge relationship association connecting different attribute fields; a feature association graph is an enhanced feature association graph network covering the multi-dimensional attribute relationships of a role;
[0106] Specifically, the semantic extension and association mining of generated role features are realized by instantiating and filling abstract attribute nodes in the knowledge graph and the multi-hop feature propagation mechanism; the abstract concept nodes in the knowledge graph are converted into specific instances to provide a richer entity expression for role features; semantic diffusion is carried out in the knowledge graph through multi-hop feature propagation with a predefined number of hops to construct an extended feature subgraph containing core features and their associated attributes; at the same time, cross-domain associations are established based on the semantic similarity matrix to form a feature association graph covering the multi-dimensional attributes of the role.
[0107] In one embodiment, if Figure 5 As shown, step S331 includes the following steps:
[0108] S3311: Identify all nodes in the knowledge graph, filter out conceptual nodes that are not bound to specific feature values, and identify parent nodes with incomplete attributes that have inheritance relationships;
[0109] S3312: Identify basic character features in the generated character features, extract scene context features, and generate a set of instantiation constraints;
[0110] S3313: Instantiate and populate the screened conceptual nodes and the identified parent nodes based on the instantiation constraint condition set.
[0111] In this embodiment, a conceptual node is a node that represents an abstract category or template in the knowledge graph and may have the following characteristics: no specific value is bound, that is, only a type definition but no instance data (such as "weapon type" instead of "long sword"), inheritability, that is, it can be used as the parent class of other nodes (such as "magic creature" is the parent class of "elf"), and a to-be-instantiated tag, that is, it has a specific label; the parent node of incomplete attributes is the upper node whose subclass features are not fully defined in the inheritance relationship; the basic character features are the core character attributes directly specified by the user; the scene context features are implicit in the generated environment. Constraints include: environment type (such as "underwater level" and "space station" scene restrictions), narrative background (such as "post-apocalyptic wasteland" and "fairy world" style requirements), technical constraints (such as the rendering capability restrictions of the target platform), etc.; the instantiated constraint set is a set of rules that control attribute filling, including type constraints, which limit the value types that can be filled (such as "weapon type" can only be filled with cold weapons), range constraints, which are the upper and lower limits of numerical attributes (such as "age∈[10,200]"), and dependency constraints, which are necessary conditions between attributes (such as "flight ability → wings required").
[0112] Specifically, the conceptual nodes that need to be instantiated and the incomplete attribute parent nodes with inheritance relationships in the knowledge graph structure are identified and screened; then the basic role features and scene context information in the generated role features are identified and extracted to construct a set of instantiation constraints; based on the set of instantiation constraints, the screened conceptual nodes and the identified parent nodes are instantiated and filled.
[0113] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0114] In one embodiment, an AI-based virtual character generation device is provided, which corresponds one-to-one to the AI-based virtual character generation method in the above embodiment. The AI-based virtual character generation device includes:
[0115] A feature acquisition module, configured to acquire corresponding multimodal features in response to a character generation instruction input by a user terminal;
[0116] An input module is used to input multimodal features into a pre-trained multimodal fusion model so that the multimodal fusion model outputs generated character features;
[0117] Enhanced feature generation module, used to perform semantic enhancement and logical reasoning on generated character features through knowledge graph to obtain enhanced character features;
[0118] An output module, for generating initial virtual character data based on the enhanced character features and sending the data to the user end for adjustment;
[0119] A character generation module is used to generate a corresponding virtual character when receiving the optimized virtual character data fed back by the user terminal;
[0120] Optionally, the feature acquisition module includes:
[0121] A type identification submodule is used to identify the instruction type of the character generation instruction, wherein the instruction type includes text instruction, image instruction, video instruction and voice instruction;
[0122] A feature extraction submodule is used to extract multimodal features of the character generation instructions based on the instruction type, thereby obtaining multimodal features;
[0123] Optionally, also include:
[0124] The feature encoding layer module is used to perform modal encoding on multimodal features to extract multimodal semantic representations;
[0125] The cross-modal interaction layer module is used to perform multimodal alignment and cross-modal fusion on multimodal semantic representations to generate a semantic space vector of a unified role;
[0126] The feature decoding layer module is used to generate and output the generated role features based on the semantic space vector of the unified role;
[0127] Optionally, the enhanced feature generation module includes:
[0128] The knowledge graph construction submodule is used to construct a knowledge graph of associated role attributes, where nodes are discretized role attributes and their characteristic values, and edges are the associations and numerical constraints between role attributes;
[0129] The embedding space mapping submodule is used to map the generated character features to the embedding space of the knowledge graph and calculate the similarity distribution between them and each node;
[0130] The extended feature generation submodule is used to perform attribute extension and semantic association on the generated role features based on the knowledge graph topology structure and semantic association path to generate semantically enhanced extended role features;
[0131] The constraint verification submodule is used to perform constraint verification based on predefined inference rules and generate enhanced role features;
[0132] Optionally, the extended feature generation submodule includes:
[0133] The instantiation filling submodule is used to identify abstract attribute nodes in the knowledge graph and perform instantiation filling;
[0134] The multi-hop feature propagation submodule is used to perform multi-hop feature propagation of a predefined number of hops on the generated role features based on the semantic association path to obtain an extended feature subgraph;
[0135] The cross-domain association establishment submodule is used to build a semantic similarity matrix between node attributes based on the knowledge graph and semantic association path, and establish cross-domain associations to obtain a feature association graph;
[0136] The feature generation submodule is used to generate semantically enhanced extended role features based on the extended feature subgraph and feature association graph;
[0137] Optionally, instantiate the populated submodule specifically for:
[0138] Identify all nodes in the knowledge graph, filter out conceptual nodes that are not bound to specific feature values, and identify parent nodes with incomplete attributes that have inheritance relationships;
[0139] Identify the basic features of the role in the generated role features, extract the scene context features, and generate a set of instantiation constraints;
[0140] The filtered conceptual nodes and the identified parent nodes are instantiated and populated based on the instantiation constraint set.
[0141] For the specific definition of an AI-based virtual character generation device, please refer to the definition of an AI-based virtual character generation method above, and will not be repeated here. The various modules in the above-mentioned AI-based virtual character generation device can be implemented in whole or in part through software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0142] In one embodiment, a smart terminal is provided, comprising a memory and a processor, wherein the memory stores at least one instruction, at least one program, code set, or instruction set, and the at least one instruction, at least one program, code set, or instruction set is loaded and executed by the processor to implement the steps of an AI-based virtual character generation method.
[0143] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, a method for generating a virtual character based on AI is implemented.
[0144] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0145] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0146] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A method for generating a virtual character based on AI, characterized by: Including steps: In response to a character generation instruction input by a user terminal, corresponding multimodal features are obtained; Input the multimodal features into a pre-trained multimodal fusion model so that the multimodal fusion model outputs the generated character features; Perform semantic enhancement and logical reasoning on the generated character features through knowledge graph to obtain enhanced character features; Generating initial virtual character data based on the enhanced character features and sending it to the user end for adjustment; When the optimized virtual character data fed back by the user is received, a corresponding virtual character is generated.
2. The AI-based virtual character generation method according to claim 1, characterized in that: The step of obtaining corresponding multimodal features in response to a character generation instruction input by a user terminal includes the following steps: Identifying the type of the character generation instruction, wherein the instruction type includes text instructions, image instructions, video instructions, and voice instructions; Multimodal features are extracted from the character generation instructions based on the instruction type to obtain multimodal features.
3. The AI-based virtual character generation method according to claim 1, characterized in that: The multimodal fusion model includes a feature encoding layer, a cross-modal interaction layer, and a feature decoding layer. The step of inputting the multimodal features into the pre-trained multimodal fusion model so that the multimodal fusion model outputs the generated character features includes the following steps: The feature encoding layer performs modal encoding on multimodal features to extract multimodal semantic representations; The cross-modal interaction layer performs multimodal alignment and cross-modal fusion on the multimodal semantic representations to generate a semantic space vector of a unified role; The feature decoding layer generates and outputs the generated role features based on the semantic space vector of the unified role.
4. The AI-based virtual character generation method according to claim 1, characterized in that: The step of performing semantic enhancement and logical reasoning on the generated character features through the knowledge graph to obtain enhanced character features includes the following steps: Construct a knowledge graph of associated role attributes, where nodes are discretized role attributes and their characteristic values, and edges are the associations and numerical constraints between role attributes; Map the generated role features to the embedding space of the knowledge graph and calculate the similarity distribution between them and each node; Based on the knowledge graph topology and semantic association path, the generated role features are attribute-extended and semantically associated to generate semantically enhanced extended role features; Constraint verification is performed based on predefined inference rules to generate enhanced role features.
5. The AI-based virtual character generation method according to claim 4, characterized in that: The step of performing attribute extension and semantic association on the generated role features based on the knowledge graph and the semantic association path to generate semantically enhanced extended role features includes the following steps: Identify abstract attribute nodes in the knowledge graph and fill them with instantiations; Based on the semantic association path, the generated role features are propagated through multiple hops with a predefined number of hops to obtain an extended feature subgraph; Based on the knowledge graph and semantic association path, a semantic similarity matrix between node attributes is constructed, and cross-domain associations are established to obtain a feature association graph; Semantically enhanced extended role features are generated based on the extended feature subgraph and feature association graph.
6. The AI-based virtual character generation method according to claim 5, characterized in that: The step of identifying abstract attribute nodes in the knowledge graph and performing instantiation and filling includes the following steps: Identify all nodes in the knowledge graph, filter out conceptual nodes that are not bound to specific feature values, and identify parent nodes with incomplete attributes that have inheritance relationships; Identify the basic features of the role in the generated role features, extract the scene context features, and generate a set of instantiation constraints; The filtered conceptual nodes and the identified parent nodes are instantiated and populated based on the instantiation constraint set.
7. An AI-based virtual character generation device, characterized by: include: A feature acquisition module, configured to acquire corresponding multimodal features in response to a character generation instruction input by a user terminal; An input module is used to input multimodal features into a pre-trained multimodal fusion model so that the multimodal fusion model outputs generated character features; Enhanced feature generation module, used to perform semantic enhancement and logical reasoning on generated character features through knowledge graph to obtain enhanced character features; An output module, for generating initial virtual character data based on the enhanced character features and sending the data to the user end for adjustment; The character generation module is used to generate a corresponding virtual character when receiving the optimized virtual character data fed back by the user end.
8. The AI-based virtual character generation device according to claim 7, characterized in that: The feature acquisition module includes: A type identification submodule is used to identify the instruction type of the character generation instruction, wherein the instruction type includes text instruction, image instruction, video instruction and voice instruction; The feature extraction submodule is used to extract multimodal features of the character generation instructions based on the instruction type, thereby obtaining multimodal features.
9. An intelligent terminal, characterized in that: The method comprises a memory and a processor, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the steps of the AI-based virtual character generation method as described in any one of claims 1 to 6.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the AI-based virtual character generation method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Data synthesis method and device, electronic equipment and storage medium
CN121581012A
Data synthesis method and apparatus, electronic device, and storage medium
CN121581012B