A knowledge enhancement-based multi-modal digital human semantic reasoning system

CN122509342APending Publication Date: 2026-08-04XINJIANG MEITE INTELLIGENT SAFETY ENG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XINJIANG MEITE INTELLIGENT SAFETY ENG CO LTD
Filing Date
2026-05-18
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

现有技术依赖单一的语义匹配逻辑,未建立有效的知识缺失识别与补全体系,当待理解交互数据存在语义不完整、要素缺失等情况时,无法主动检测并基于关联知识进行语义补全,导致数字人对不完整交互信息的理解能力不足,难以应对实际交互中用户表达不充分、信息碎片化的场景,降低交互的智能化水平与用户体验

Benefits of technology

[0053]本发明通过构建包含标准知识语义因子的多模态知识增强语义关联架构,将分散异构的多模态知识拆解为标准化的知识语义因子,实现知识的结构化沉淀与关联化管理,从根源上解决多模态交互数据语义混杂、标准缺失的问题,为全流程语义推理提供稳定可靠的知识支撑,提升语义理解的一致性与规范性。通过生成匹配比对的递进式处理流程,实现待理解交互数据到标准化语义因子的转化与匹配,能够识别待理解语义中的知识缺失元素,有效弥补传统语义推理中知识补全不足的缺陷,提升语义推理的完整性与准确性。通过第一语义推理系数与第二语义推理系数的双维度量化计算,将知识缺失程度与关联知识支撑能力转化为可量化的推理指标,实现语义推理的精细化、数据化评估,避免传统定性推理的主观性偏差,通过关联标准知识语义因子的调用与融合,实现知识对语义推理的深度增强,能够基于关联知识补全缺失语义,提升多模态数字人对不完整交互数据的理解能力,能够适配多模态交互全场景的语义推理需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122509342A_ABST
    Figure CN122509342A_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal digital person semantic reasoning systems based on knowledge enhancement, it is related to semantic reasoning technical field, its technical solution key points include: the multi-modal knowledge enhancement semantic association architecture of standard knowledge semantic factor is included in construction, wherein, the standard knowledge semantic factor is associated with factor type, knowledge semantic factor element and factor reasoning mark determined based on knowledge fusion sequence;The type to be understood of multi-modal digital person interactive data to be understood is acquired, and corresponding knowledge semantic factor to be understood is generated according to the type to be understood;Based on the factor type of knowledge semantic factor to be understood;Second semantic reasoning coefficient is obtained based on the associated standard knowledge semantic factor;According to first semantic reasoning coefficient and second semantic reasoning coefficient, output multi-modal digital person semantic reasoning result, effect is to improve the understanding ability of multi-modal digital person to incomplete interactive data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of semantic reasoning technology, and more specifically, to a knowledge-enhanced multimodal digital human semantic reasoning system. Background Technology

[0002] With the rapid development of artificial intelligence technology, the application scope of multimodal digital humans in interactive scenarios is constantly expanding, from basic customer service interactions to complex intelligent service scenarios, placing higher demands on the semantic reasoning capabilities of digital humans. Existing technologies rely on a single semantic matching logic and have not established an effective system for identifying and completing missing knowledge. When the interactive data to be understood contains semantic incompleteness or missing elements, it cannot proactively detect and complete the semantics based on related knowledge. This results in insufficient understanding of incomplete interactive information by digital humans, making it difficult to cope with scenarios where users' expressions are insufficient and information is fragmented in actual interactions, thus reducing the level of intelligence and user experience. Furthermore, the lack of quantification of the degree of knowledge deficiency and the ability to support related knowledge makes it impossible to characterize the rationality and reliability of semantic reasoning through quantitative indicators. Simultaneously, the lack of a related knowledge retrieval mechanism based on factor reasoning tags leads to a lack of scientific quantitative support for the semantic reasoning process, resulting in insufficient persuasiveness and accuracy of the reasoning results, and hindering the optimization and iteration of the reasoning process. Summary of the Invention

[0003] In view of the shortcomings of existing technologies, the purpose of this invention is to provide a knowledge-enhanced multimodal digital human semantic reasoning system.

[0004] To achieve the above objectives, the present invention provides the following technical solution:

[0005] A knowledge-enhanced multimodal digital human semantic reasoning system includes:

[0006] Construction module: Constructs a multimodal knowledge-enhanced semantic association architecture containing standard knowledge semantic factors, wherein the standard knowledge semantic factors are associated with factor types, knowledge semantic factor elements, and factor inference tags determined based on the knowledge fusion order;

[0007] Generation module: Obtain the type of multimodal digital human interaction data to be understood, and generate corresponding semantic factors of knowledge to be understood based on the type of data to be understood;

[0008] Matching module: Based on the type of the semantic factors to be understood, the target standard knowledge semantic factors are matched from the multimodal knowledge-enhanced semantic association architecture.

[0009] Comparison module: The uncomprehended semantic factors of the knowledge semantic factors to be understood are compared with the knowledge semantic factors of the target standard knowledge semantic factors to obtain the knowledge missing elements. The uncomprehended semantic factors with knowledge missing elements are marked as knowledge enhancement anomalous semantic factors.

[0010] Processing module: The first semantic inference coefficient is obtained based on the knowledge-deficient elements of the knowledge-enhanced anomalous semantic factors; the associated standard knowledge semantic factors are obtained based on the factor inference tags of the knowledge-enhanced anomalous semantic factors; and the second semantic inference coefficient is obtained based on the associated standard knowledge semantic factors.

[0011] Output module: Outputs the semantic reasoning results of the multimodal digital human based on the first semantic reasoning coefficient and the second semantic reasoning coefficient.

[0012] Preferably, a multimodal knowledge-enhanced semantic association architecture containing standard knowledge semantic factors is constructed, specifically including the following steps:

[0013] A standard template for digital human semantic understanding is set, which includes understanding types and digital human semantic understanding elements corresponding to the understanding types;

[0014] Set up a knowledge semantic factor template library containing knowledge semantic factor templates, where each knowledge semantic factor template corresponds to a template understanding type;

[0015] Construct a multimodal knowledge-enhanced semantic association architecture based on the matching results between the understanding type and the template understanding type.

[0016] Preferably, generating corresponding semantic factors of knowledge to be understood based on the type to be understood specifically includes the following steps:

[0017] Basic semantic units are obtained by processing multimodal digital human interaction data to understand them;

[0018] Based on the type to be understood, each basic semantic unit is assigned a corresponding semantic association attribute. The basic semantic units with semantic association attributes are arranged in an orderly manner according to the interaction logic order to form an initial semantic combination.

[0019] After performing semantic integrity verification on the initial semantic combination, retain the semantically clear and logically coherent unit combination;

[0020] The primary and secondary relationships of unit combinations are distinguished according to their semantic importance, and the unit combinations are then regularized and integrated based on these relationships to generate semantic factors of the knowledge to be understood.

[0021] Preferably, the basic semantic units are obtained by processing the multimodal digital human interaction data to be understood, specifically including the following steps:

[0022] Semantic feature decomposition is performed on the multimodal digital human interaction data to be understood to obtain the interaction semantic dimensions of the types to be understood;

[0023] Features from each interactive semantic dimension are aggregated to form basic semantic units.

[0024] Preferably, based on the type of the semantic factors to be understood, the target standard knowledge semantic factors are matched from the multimodal knowledge-enhanced semantic association architecture, specifically including the following steps:

[0025] The standard set of semantic factors of knowledge is obtained by processing the types of factors to be understood in the semantic factors of knowledge to be understood.

[0026] The semantic relevance of each standard knowledge semantic factor in the standard knowledge semantic factor set is ranked to obtain the ranking result. Based on the ranking result, the factor with the highest semantic fit is selected as the initial standard knowledge semantic factor.

[0027] Verify the semantic compatibility between the initial selected standard knowledge semantic factors and the types of factors to be understood;

[0028] Based on semantic adaptability, the corresponding initial standard knowledge semantic factors are determined as target standard knowledge semantic factors.

[0029] Preferably, the standard knowledge semantic factor set is obtained by processing the type of the knowledge semantic factor to be understood, specifically including the following steps:

[0030] Based on the type of semantic factors to be understood, hierarchical semantic processing is performed within the multimodal knowledge-enhanced semantic association architecture to obtain the processing results;

[0031] Based on the processing results, a set of standard knowledge semantic factors that match the type of factors to be understood is obtained.

[0032] Preferably, the missing knowledge elements are obtained by comparing the elements of the semantic factors of the knowledge to be understood with the elements of the semantic factors of the target standard knowledge factor, specifically including the following steps:

[0033] By analyzing the relationships between the various factor elements to be understood, a relationship system of the factor elements to be understood can be obtained;

[0034] Extract the semantic elements of the target standard knowledge semantic factors, and sort out the relationships between the semantic elements of each knowledge semantic factor to form a standard factor element relationship system;

[0035] The matching degree is obtained by comparing the correlation system of the factor elements to be understood with the correlation system of the standard factor elements at each level; the missing knowledge elements are obtained based on the matching degree.

[0036] Preferably, the first semantic inference coefficient is obtained based on the knowledge-deficient elements of the knowledge-enhanced abnormal semantic factors, specifically including the following steps:

[0037] Obtain the missing element type and total number of missing elements for the knowledge missing elements;

[0038] Each missing knowledge element is assigned a missing element weight based on its missing element type, and the missing element weights are summed based on the total number of missing elements to obtain the factor element missing coefficient.

[0039] Set factor missing weights for knowledge-enhancing anomaly semantic factors based on the type of information to be understood;

[0040] The first semantic inference coefficient is obtained based on the factor missing weight and the factor element missing coefficient.

[0041] Preferably, the second semantic inference coefficient is obtained based on the semantic factors of the associated standard knowledge, specifically including the following steps:

[0042] Obtain the total number of associated factors and the total number of associated factor types for the semantic factors of the associated standard knowledge;

[0043] The first correlation inference coefficient is obtained based on the total number of correlation factors;

[0044] Assign knowledge type weights to the knowledge semantic factors based on the knowledge semantic factor types of the associated standard knowledge semantic factors, and accumulate the knowledge type weights based on the total number of associated factor types to obtain the second association inference coefficient;

[0045] The second semantic inference coefficient is obtained based on the first and second association inference coefficients.

[0046] Preferably, the multimodal digital human semantic reasoning result is output based on the first semantic reasoning coefficient and the second semantic reasoning coefficient, including the following steps:

[0047] The initial semantic determination result is obtained by constructing a coefficient fusion relationship based on the first semantic inference coefficient and the second semantic inference coefficient.

[0048] Perform semantic integrity verification on the initial semantic determination results;

[0049] The valid semantic information after verification is semantically adapted and sorted with the semantic factors of the associated standard knowledge to obtain the sorting result, and the core semantic orientation is determined based on the sorting result.

[0050] The localization result is obtained by localizing the multimodal digital human interaction intent based on the core semantic orientation;

[0051] Based on the positioning results, output multimodal digital human semantic reasoning results adapted to multimodal interaction scenarios.

[0052] Compared with the prior art, the present invention has the following beneficial effects:

[0053] This invention constructs a multimodal knowledge-enhanced semantic association architecture that incorporates standardized knowledge semantic factors. It decomposes dispersed and heterogeneous multimodal knowledge into standardized knowledge semantic factors, achieving structured knowledge accumulation and associative management. This fundamentally solves the problems of semantic confusion and lack of standards in multimodal interactive data, providing stable and reliable knowledge support for the entire semantic reasoning process and improving the consistency and standardization of semantic understanding. Through a progressive processing flow of generation and matching, it achieves the transformation and matching of interactive data to be understood into standardized semantic factors. This identifies missing knowledge elements in the semantics to be understood, effectively compensating for the deficiencies in knowledge completion in traditional semantic reasoning and improving the completeness and accuracy of semantic reasoning. By using a two-dimensional quantitative calculation of the first and second semantic reasoning coefficients, the degree of knowledge deficiency and the supporting ability of related knowledge are transformed into quantifiable reasoning indicators, enabling refined and data-driven evaluation of semantic reasoning. This avoids the subjective bias of traditional qualitative reasoning. By calling and integrating related standard knowledge semantic factors, knowledge is used to deeply enhance semantic reasoning. It can fill in missing semantics based on related knowledge, improve the understanding ability of multimodal digital humans of incomplete interactive data, and adapt to the semantic reasoning needs of multimodal interaction in all scenarios. Attached Figure Description

[0054] Figure 1 This invention provides a schematic diagram of a multimodal digital human semantic reasoning system based on knowledge enhancement, as shown in the embodiment of the invention.

[0055] Figure 2 This invention provides a schematic diagram illustrating the steps involved in generating semantic factors of knowledge to be understood in a knowledge-enhanced multimodal digital human semantic reasoning system. Detailed Implementation

[0056] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0057] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0058] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places throughout this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment selectively excluded from other embodiments.

[0059] Reference Figures 1-2 As shown.

[0060] The embodiments further illustrate the multimodal digital human semantic reasoning system based on knowledge enhancement proposed in this invention.

[0061] A knowledge-enhanced multimodal digital human semantic reasoning system includes:

[0062] Construction module: Constructs a multimodal knowledge-enhanced semantic association architecture containing standard knowledge semantic factors, wherein the standard knowledge semantic factors are associated with factor types, knowledge semantic factor elements, and factor inference tags determined based on the knowledge fusion order;

[0063] Generation module: Obtain the type of multimodal digital human interaction data to be understood, and generate corresponding semantic factors of knowledge to be understood based on the type of data to be understood;

[0064] Matching module: Based on the type of the semantic factors to be understood, the target standard knowledge semantic factors are matched from the multimodal knowledge-enhanced semantic association architecture.

[0065] Comparison module: The uncomprehended semantic factors of the knowledge semantic factors to be understood are compared with the knowledge semantic factors of the target standard knowledge semantic factors to obtain the knowledge missing elements. The uncomprehended semantic factors with knowledge missing elements are marked as knowledge enhancement anomalous semantic factors.

[0066] Processing module: The first semantic inference coefficient is obtained based on the knowledge-deficient elements of the knowledge-enhanced anomalous semantic factors; the associated standard knowledge semantic factors are obtained based on the factor inference tags of the knowledge-enhanced anomalous semantic factors; and the second semantic inference coefficient is obtained based on the associated standard knowledge semantic factors.

[0067] Output module: Outputs the semantic reasoning results of the multimodal digital human based on the first semantic reasoning coefficient and the second semantic reasoning coefficient.

[0068] Constructing a multimodal knowledge-enhanced semantic association architecture that includes standard knowledge semantic factors specifically includes the following steps:

[0069] Set up a standard template for digital human semantic understanding, which includes understanding types and digital human semantic understanding elements corresponding to the understanding types;

[0070] Set up a knowledge semantic factor template library containing knowledge semantic factor templates, where each knowledge semantic factor template corresponds to a template understanding type;

[0071] Construct a multimodal knowledge-enhanced semantic association architecture based on the matching results between the understanding type and the template understanding type.

[0072] The standard template for digital human semantic understanding is a standardized semantic understanding specification for multimodal digital human interaction scenarios. Understanding types define different semantic understanding scenarios for multimodal digital humans, encompassing various semantic understanding needs across all multimodal interaction scenarios. These include different sub-scenarios such as user command semantic understanding, emotion and intent recognition, multi-turn dialogue context understanding, and multimodal information fusion understanding. Digital human semantic understanding elements are the constituent elements supporting each understanding type, a set of elements matched to different understanding types. For example, semantic understanding elements for user command understanding include the command subject, command action, command constraints, and command target; semantic understanding elements for emotion and intent recognition include emotion tendency, emotion intensity, intent direction, and triggering scenario. The standard template establishes a unified and standardized framework for multimodal digital human semantic understanding, clarifying the core elements constituting different semantic understanding scenarios.

[0073] The knowledge semantic factor template library is a structured collection used to store and manage standardized knowledge semantic factor templates. Each knowledge semantic factor template corresponds to a unique template understanding type. The knowledge semantic factor template breaks down and encapsulates various types of knowledge required for multimodal digital human interaction, including domain-specific knowledge, general common sense knowledge, interaction scenario knowledge, and historical dialogue knowledge, into standardized semantic factor units. Each template encapsulates the semantic structure, element composition, and relational logic of the corresponding type of knowledge, achieving the structured and standardized accumulation of scattered knowledge. The template understanding type is the semantic understanding scenario type that the knowledge semantic factor template is adapted to. Its classification dimensions are consistent with the understanding types of the standard digital human semantic understanding templates, ensuring the adaptation of knowledge templates to semantic understanding scenarios and enabling the classified management and rapid retrieval of standardized knowledge.

[0074] This process involves matching each understanding type defined in the digital human semantic understanding standard template with the corresponding template understanding type in the knowledge semantic factor template library. This establishes a correspondence or many-to-many relationship between understanding types and their corresponding knowledge semantic factor templates, clarifying the standardized knowledge semantic factor set suitable for different semantic understanding scenarios. Based on this matching relationship, the standardized semantic understanding framework of the digital human semantic understanding standard template is deeply integrated with the standardized knowledge semantic factors in the knowledge semantic factor template library, ultimately constructing a multimodal knowledge-enhanced semantic association architecture. This architecture is a complete association system integrating the standard semantic understanding framework and standardized knowledge semantic factors, clarifying the core element requirements for different semantic understanding scenarios of the multimodal digital human, and matching corresponding standardized knowledge semantic factors for each understanding scenario.

[0075] Generate corresponding semantic factors of knowledge to be understood based on the type of knowledge to be understood, specifically including the following steps:

[0076] The basic semantic units are obtained by processing multimodal digital human interaction data to be understood, specifically including the following steps:

[0077] Semantic feature decomposition is performed on the multimodal digital human interaction data to be understood to obtain the interaction semantic dimensions of the types to be understood;

[0078] Feature aggregation is performed on each interactive semantic dimension to form basic semantic units;

[0079] Based on the type to be understood, each basic semantic unit is assigned a corresponding semantic association attribute. The basic semantic units with semantic association attributes are arranged in an orderly manner according to the interaction logic order to form an initial semantic combination.

[0080] After performing semantic integrity verification on the initial semantic combination, retain the semantically clear and logically coherent unit combination;

[0081] The primary and secondary relationships of unit combinations are distinguished according to their semantic importance, and the unit combinations are then regularized and integrated based on these relationships to generate semantic factors of the knowledge to be understood.

[0082] Multimodal digital human interaction data to be understood refers to multi-source heterogeneous data generated during the interaction between users and digital humans, including various modal information such as text, voice, images, actions, and expressions. Through semantic feature decomposition, the original multimodal interaction data is decomposed into multiple independent and clearly defined interaction semantic dimensions based on the preset type to be understood. Among them, the type to be understood is a predefined semantic understanding scenario classification.

[0083] The decomposed, scattered features are transformed into standardized semantic units. Each interactive semantic dimension still contains a large amount of scattered multimodal feature data. Through feature aggregation operations, all relevant multimodal features under the same interactive semantic dimension are integrated, refined, and standardized. Redundant and invalid feature information is removed, and semantic information under that dimension is extracted, ultimately forming the basic semantic unit of the corresponding dimension. Heterogeneous multimodal data is transformed into basic semantic units in a unified format, providing standardized basic units for subsequent semantic combination and factor generation.

[0084] First, each basic semantic unit is assigned a unique semantic association attribute based on the current type to be understood. This attribute clarifies the role, location, and association logic of the basic semantic unit within the semantic system of the corresponding type to be understood. The basic semantic units with semantic association attributes are then arranged in an orderly manner according to the natural logical sequence of multimodal interaction, connecting the scattered independent units into a complete semantic whole, forming an initial semantic combination.

[0085] The initial semantic combination is a preliminary integration of basic semantic units, which may contain issues such as semantic gaps, logical breaks, information redundancy, and semantic conflicts, making it unsuitable for subsequent processing. Therefore, a comprehensive verification of the initial semantic combination is performed based on pre-defined semantic integrity verification rules: the verification includes the semantic clarity of each basic semantic unit, i.e., whether the semantic information carried by the unit is clear and unambiguous; the logical coherence between units, i.e., whether the unit arrangement conforms to the interaction logic and whether there are logical breaks or conflicts; and the semantic integrity of the combination, i.e., whether it covers all the core semantic elements required for the type to be understood and whether there is any missing key information.

[0086] First, based on the semantic requirements of the type to be understood, the semantic importance of each basic semantic unit in the combination is assessed, distinguishing between core semantic units and auxiliary semantic units, and clarifying the primary and secondary relationships of the unit combination. Among them, the core semantic unit is the unit that carries the core semantic information of the type to be understood, and the auxiliary semantic units are used to supplement and improve the core semantics. The unit combination is regularized and integrated according to the primary and secondary relationships, and finally standardized semantic factors of the knowledge to be understood are generated. These semantic factors are standardized semantic units that integrate multimodal interaction information, meet the semantic requirements of the type to be understood, have a clear structure, and have a clear distinction between primary and secondary elements, and fully carry the semantic information of the multimodal interaction data to be understood.

[0087] Based on the type of semantic factors to be understood, the target standard knowledge semantic factors are matched from the multimodal knowledge-enhanced semantic association architecture. This process includes the following steps:

[0088] The standard set of semantic factors of knowledge is obtained by processing the types of factors to be understood, which includes the following steps:

[0089] Based on the type of semantic factors to be understood, hierarchical semantic processing is performed within the multimodal knowledge-enhanced semantic association architecture to obtain the processing results;

[0090] Based on the processing results, a set of standard knowledge semantic factors that match the types of factors to be understood is obtained;

[0091] The semantic relevance of each standard knowledge semantic factor in the standard knowledge semantic factor set is ranked to obtain the ranking result. Based on the ranking result, the factor with the highest semantic fit is selected as the initial standard knowledge semantic factor.

[0092] Verify the semantic compatibility between the initial selected standard knowledge semantic factors and the types of factors to be understood;

[0093] Based on semantic adaptability, the corresponding initial standard knowledge semantic factors are determined as target standard knowledge semantic factors.

[0094] The semantic factors of knowledge to be understood are standardized semantic units generated from multimodal digital human interaction data. The type of the factor to be understood is the semantic understanding scenario classification to which the semantic factor belongs. First, hierarchical semantic processing is carried out within a pre-constructed multimodal knowledge-enhanced semantic association architecture based on the type of the factor to be understood. Hierarchical semantic processing searches layer by layer from the top-level semantic category to the sub-categories according to the preset semantic hierarchy within the architecture. This ensures that the search process includes all semantic branches related to the type of the factor to be understood within the architecture, avoiding the omission of potential matching objects, and finally obtaining a complete hierarchical semantic processing result. Based on this processing result, all standard knowledge semantic factors that match the semantics of the type of the factor to be understood are selected and collected into a standard knowledge semantic factor set. This set contains all candidate standard knowledge semantic factors that meet the type matching requirements, ensuring the comprehensiveness and completeness of the matching process from the root.

[0095] The semantic relevance of each standard knowledge semantic factor within the standard knowledge semantic factor set is determined. The semantic relevance is used to characterize the degree of semantic fit between the standard knowledge semantic factor and the semantic factor to be understood. All standard knowledge semantic factors in the set are sorted in descending order of semantic relevance to obtain a complete sorting result. Based on the sorting result, the standard knowledge semantic factor with the highest semantic fit is selected as the initial standard knowledge semantic factor. The relevance sorting enables rapid screening of the candidate set, locking in the most promising matching object from a large number of candidates, greatly improving matching efficiency, while ensuring the semantic fit of the initial selection object.

[0096] The initial selection criteria for semantic factors are based on the optimal candidate objects obtained through semantic relevance ranking. However, they still need to undergo semantic adaptability verification to confirm their actual matching effectiveness with the type of factor to be understood, avoiding mismatches caused by deviations in relevance calculation. Semantic adaptability verification revolves around the semantic requirements of the type of factor to be understood, comprehensively verifying whether the type attributes, semantic element composition, and core content of the semantic logic framework of the initial selection criteria for semantic factors are fully adapted to the semantic understanding scenario of the type of factor to be understood. This ensures that the initial selection objects not only perform optimally in terms of relevance but also fully meet the requirements in terms of actual semantic adaptability, fundamentally avoiding the risk of matching errors.

[0097] If the initial standard knowledge semantic factor passes the verification, that is, it is confirmed that its semantic adaptability to the type of factor to be understood is fully in line with the requirements, then the initial standard knowledge semantic factor is officially determined as the target standard knowledge semantic factor; if the initial object fails the verification, the next candidate object is selected in turn according to the semantic relevance ranking result, and the adaptability verification process is repeated until the target standard knowledge semantic factor that meets the requirements is selected.

[0098] The missing knowledge elements are obtained by comparing the elements of the semantic factors of the knowledge to be understood with the elements of the semantic factors of the target standard knowledge factors. This process includes the following steps:

[0099] By analyzing the relationships between the various factor elements to be understood, a relationship system of the factor elements to be understood can be obtained;

[0100] Extract the semantic elements of the target standard knowledge semantic factors, and sort out the relationships between the semantic elements of each knowledge semantic factor to form a standard factor element relationship system;

[0101] The matching degree is obtained by comparing the correlation system of the factor elements to be understood with the correlation system of the standard factor elements at each level; the missing knowledge elements are obtained based on the matching degree.

[0102] The elements to be understood are the core semantic units contained within the semantic factors of the knowledge to be understood. Each element is defined within the overall semantic system, determining its position, role, and the subordinate, parallel, and progressive relationships between elements. Following the inherent laws of semantic logic, these elements are linked and integrated to form a complete system of relationships between the elements. This system fully restores the original semantic structure of the semantic factors of the knowledge to be understood and clearly presents the semantic connections between each element.

[0103] The target standard knowledge semantic factor is a pre-constructed standardized semantic reference benchmark. Its internal knowledge semantic factor elements are a complete set of semantic elements corresponding to a specific semantic understanding scenario. First, all knowledge semantic factor elements contained in the target standard knowledge semantic factor are extracted from the multimodal knowledge-enhanced semantic association architecture. The relationships between these elements are then analyzed to clarify the hierarchical relationships, semantic association logic, and overall semantic architecture of each element within the standard semantic system. Following the inherent rules of the standard semantic structure, these elements are integrated to form a standard factor element association system. This system fully presents the standard semantic structure of the corresponding semantic understanding scenario, containing complete core semantic elements and standardized semantic association logic.

[0104] The matching degree is obtained by comparing the association system of the factor elements to be understood with the association system of the standard factor elements layer by layer. Based on the matching degree, the missing knowledge elements are identified. The comparison process unfolds layer by layer according to the semantic hierarchy of the two association systems, ensuring that the comparison covers all semantic levels and avoiding omission of semantic element differences at any level. During the layer-by-layer comparison, the number of matching elements and the matching ratio at each level are statistically analyzed to obtain the overall matching degree and the matching degree at each level. The matching degree is used to quantitatively characterize the semantic fit between the association system of the factor elements to be understood and the association system of the standard factor elements. Based on the calculation results of the matching degree, the parts of the association system of the factor elements to be understood that do not match the standard factor elements are identified. These unmatched elements are the missing knowledge elements, which fully represent the semantic elements missing from the semantic factors of the knowledge to be understood compared to the standard semantic system. This hierarchical comparison achieves the identification of missing knowledge elements, ensuring the completeness and accuracy of semantic reasoning in multimodal digital humans.

[0105] The first semantic inference coefficient is obtained based on the knowledge-missing elements of the knowledge-enhanced anomaly semantic factor, specifically including the following steps:

[0106] Obtain the missing element type and total number of missing elements for the knowledge missing elements;

[0107] Each missing knowledge element is assigned a missing element weight based on its missing element type, and the missing element weights are summed based on the total number of missing elements to obtain the factor element missing coefficient.

[0108] Set factor missing weights for knowledge-enhancing anomaly semantic factors based on the type of information to be understood;

[0109] The first semantic inference coefficient is obtained based on the factor missing weight and the factor element missing coefficient.

[0110] The missing knowledge elements are identified by hierarchically comparing the semantic factors of the knowledge to be understood with the semantic factors of the target standard knowledge. All missing knowledge elements are classified and statistically analyzed to clarify the missing element type corresponding to each missing knowledge element, that is, the semantic category to which the missing element belongs. At the same time, the total number of all missing knowledge elements is counted to obtain the total number of missing elements.

[0111] Each missing knowledge element is assigned a missing element weight based on its type. These missing element weights are then summed based on the total number of missing elements to obtain a factor element missing coefficient. Different types of missing knowledge elements have varying degrees of impact on semantic reasoning. Therefore, according to a pre-defined weighting rule, a corresponding missing element weight is assigned to each missing knowledge element based on its type. The weight is positively correlated with the degree of influence of that type of element on semantic reasoning; the higher the influence, the larger the weight. The missing element weights of all missing knowledge elements are summed based on the total number of missing elements, integrating the scattered individual element weights into a unified factor element missing coefficient. This factor element missing coefficient quantifies the overall degree of missing elements in the semantic factors of the knowledge to be understood, and its magnitude is positively correlated with the severity of the missing elements.

[0112] Factor missing weights are assigned to knowledge-enhanced anomalous semantic factors based on their uncomprehended type. These factors are uncomprehended semantic factors containing missing knowledge elements, and their importance and influence weight in the corresponding semantic understanding scenario are determined by their uncomprehended type. The factor missing weights are assigned to knowledge-enhanced anomalous semantic factors based on the semantic importance of the uncomprehended type. The magnitude of the weight is positively correlated with the core importance of the uncomprehended type in the multimodal digital human interaction scenario; the higher the core importance, the larger the assigned factor missing weight.

[0113] The first semantic inference coefficient is obtained by combining the factor missing weights and factor element missing coefficients. This involves fusing the factor missing weights and factor element missing coefficients to calculate a comprehensive analysis of both the scenario weights at the factor level and the degree of missing elements at the element level. The first semantic inference coefficient is a comprehensive quantitative representation of the semantic missing status of anomalous semantic factors in knowledge enhancement. It reflects both the degree of element missingness of the semantic factor to be understood and the importance of that semantic factor in the corresponding semantic understanding scenario, thus reflecting the impact of knowledge missingness on the semantic reasoning of multimodal digital humans.

[0114] The second semantic inference coefficient is obtained based on the semantic factors of the association standard knowledge, specifically including the following steps:

[0115] Obtain the total number of associated factors and the total number of associated factor types for the semantic factors of the associated standard knowledge;

[0116] The first correlation inference coefficient is obtained based on the total number of correlation factors;

[0117] Assign knowledge type weights to the knowledge semantic factors based on the knowledge semantic factor types of the associated standard knowledge semantic factors, and accumulate the knowledge type weights based on the total number of associated factor types to obtain the second association inference coefficient;

[0118] The second semantic inference coefficient is obtained based on the first and second association inference coefficients.

[0119] The process involves obtaining the total number of associated factors and the total number of associated factor types for the associated standard knowledge semantic factors. These factors are factor inference tags based on knowledge-enhanced anomalous semantic factors. The process retrieves a set of standardized knowledge semantic factors that are knowledge-associated with anomalous semantic factors from the multimodal knowledge-enhanced semantic association architecture. The total number of associated standard knowledge semantic factors contained in the associated standard knowledge semantic factor set is then determined, thus obtaining the total number of associated factors. Simultaneously, the associated standard knowledge semantic factors within the associated standard knowledge semantic factor set are categorized by type, and the total number of different knowledge semantic factor types is counted, resulting in the total number of associated factor types.

[0120] The total number of association factors directly reflects the richness of associated knowledge resources that can be used to supplement knowledge gaps and support semantic reasoning. The more association factors there are, the more abundant the available associated knowledge and the stronger the support for semantic reasoning. A first association reasoning coefficient is generated based on the total number of association factors according to a preset quantitative dimension calculation rule. The coefficient is positively correlated with the total number of association factors; the more association factors there are, the higher the first association reasoning coefficient. This quantitatively characterizes the strength of associated knowledge's support for semantic reasoning in terms of quantity.

[0121] Based on the semantic factor types of the associated standard knowledge factors, knowledge type weights are assigned. These weights are then summed according to the total number of associated factor types to obtain a second association inference coefficient. Different types of associated standard knowledge semantic factors have varying supporting value and influence on semantic reasoning. According to a preset weight allocation rule, and considering the semantic factor type of each associated standard knowledge semantic factor, a corresponding knowledge type weight is assigned. The weight is positively correlated with the supporting value of that type of knowledge for semantic reasoning; the higher the supporting value, the larger the assigned weight. After assigning weights to individual factors, all knowledge type weights are summed according to the total number of associated factor types. The dispersed individual type weights are then integrated into the second association inference coefficient. This coefficient is used to quantitatively characterize the comprehensive supporting strength of associated knowledge for semantic reasoning in the type dimension, achieving a quantification of the supporting capability of associated knowledge at the type level.

[0122] The steps for deriving the second semantic reasoning coefficient based on the first and second association inference coefficients involve fusing the first and second association inference coefficients to calculate a comprehensive second semantic reasoning coefficient. This coefficient represents a quantifiable measure of the comprehensive support capability of semantic factors in relational standard knowledge, reflecting both the quantitative scale and the typological value of relational knowledge, and thus demonstrating the degree to which relational knowledge supports semantic reasoning in multimodal digital humans.

[0123] The multimodal digital human semantic reasoning result is output based on the first semantic reasoning coefficient and the second semantic reasoning coefficient, including the following steps:

[0124] The initial semantic determination result is obtained by constructing a coefficient fusion relationship based on the first semantic inference coefficient and the second semantic inference coefficient.

[0125] Perform semantic integrity verification on the initial semantic determination results;

[0126] The valid semantic information after verification is semantically adapted and sorted with the semantic factors of the associated standard knowledge to obtain the sorting result, and the core semantic orientation is determined based on the sorting result.

[0127] The localization result is obtained by localizing the multimodal digital human interaction intent based on the core semantic orientation;

[0128] Based on the positioning results, output multimodal digital human semantic reasoning results adapted to multimodal interaction scenarios.

[0129] An initial semantic judgment result is obtained by constructing a coefficient fusion relationship based on the first and second semantic inference coefficients. The first semantic inference coefficient represents the impact of the degree of knowledge deficiency of the semantic factors to be understood on semantic inference, while the second semantic inference coefficient represents the support strength of the associated standard knowledge semantic factors for semantic inference. By constructing a coefficient fusion relationship, the quantitative inference coefficients of the two dimensions are comprehensively calculated, integrating the reasoning logic of different dimensions, eliminating the limitations of a single coefficient, and integrating to obtain an initial semantic judgment result that can comprehensively reflect the impact of knowledge deficiency and the support of associated knowledge.

[0130] The initial semantic judgment result is a preliminary conclusion after coefficient fusion, and may contain incomplete semantic information, missing elements, or logical inconsistencies. Based on preset semantic integrity verification rules, the initial semantic judgment result is comprehensively reviewed, including the coverage of semantic information, the completeness of core elements, and the rigor of semantic logic. After verification, invalid or incomplete semantic information is removed, retaining only valid and complete semantic content.

[0131] The validated valid semantic information is semantically adapted and ranked against associated standard knowledge semantic factors to obtain a ranking result. Based on the ranking result, the core semantic orientation is determined, and the degree of fit, semantic relevance, and matching relationship between the valid semantic information and each standard knowledge semantic factor is assessed. The valid semantic information and standard knowledge semantic factors are ranked according to their semantic fit, clarifying the closeness of association between each semantic content. Based on the ranking result, the semantic content with the highest fit and strongest relevance is identified as the core semantic orientation. This orientation defines the core semantic connotation and knowledge logic of the multimodal digital human interaction data, providing a clear semantic guide for intent localization.

[0132] The localization result is obtained by localizing the interaction intent of multimodal digital humans based on core semantic orientation. The core semantic orientation clearly defines the core semantic connotation of the interaction data. Combined with the scene characteristics and semantic logic rules of multimodal digital human interaction, the user's interaction intent is accurately identified and located. The localization process revolves around the core semantic orientation, deconstructing the behavioral demands, operational needs, or interaction purposes behind the semantics, and transforming semantic information into clear interaction intent results. This achieves a deep mapping from the semantic level to the intent level, ensuring that the final inference result matches the user's actual interaction needs.

[0133] Based on the localization results, the system outputs multimodal digital human semantic reasoning results adapted to various multimodal interaction scenarios. Based on the interaction intent localization results, and combined with the diverse needs and scene characteristics of multimodal interaction scenarios, the reasoning content undergoes scenario-specific adaptation processing. The form, content, and logic of semantic expression are optimized to adapt to different multimodal interaction scenarios, such as text interaction, voice interaction, and image interaction. The final output semantic reasoning results possess both accurate semantic connotations and conform to the interaction specifications of multimodal scenarios, achieving a complete closed loop from semantic reasoning to scenario-specific implementation. This effectively improves the practicality and scenario adaptability of multimodal digital human semantic reasoning.

[0134] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple architectural units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0135] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or architectural device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A knowledge-enhanced multimodal digital human semantic reasoning system, characterized in that, include: Construction module: Constructs a multimodal knowledge-enhanced semantic association architecture containing standard knowledge semantic factors, wherein the standard knowledge semantic factors are associated with factor types, knowledge semantic factor elements, and factor inference tags determined based on the knowledge fusion order; Generation module: Obtain the type of multimodal digital human interaction data to be understood, and generate corresponding semantic factors of knowledge to be understood based on the type of data to be understood; Matching module: Based on the type of the semantic factors to be understood, the target standard knowledge semantic factors are matched from the multimodal knowledge-enhanced semantic association architecture. Comparison module: The uncomprehended semantic factors of the knowledge semantic factors to be understood are compared with the knowledge semantic factors of the target standard knowledge semantic factors to obtain the knowledge missing elements. The uncomprehended semantic factors with knowledge missing elements are marked as knowledge enhancement anomalous semantic factors. Processing module: The first semantic inference coefficient is obtained based on the knowledge-deficient elements of the knowledge-enhanced anomalous semantic factors; the associated standard knowledge semantic factors are obtained based on the factor inference tags of the knowledge-enhanced anomalous semantic factors; and the second semantic inference coefficient is obtained based on the associated standard knowledge semantic factors. Output module: Outputs the semantic reasoning results of the multimodal digital human based on the first semantic reasoning coefficient and the second semantic reasoning coefficient.

2. The knowledge-enhanced multimodal digital human semantic reasoning system according to claim 1, characterized in that, Constructing a multimodal knowledge-enhanced semantic association architecture that includes standard knowledge semantic factors specifically includes the following steps: A standard template for digital human semantic understanding is set, which includes understanding types and digital human semantic understanding elements corresponding to the understanding types; Set up a knowledge semantic factor template library containing knowledge semantic factor templates, where each knowledge semantic factor template corresponds to a template understanding type; Construct a multimodal knowledge-enhanced semantic association architecture based on the matching results between the understanding type and the template understanding type.

3. The knowledge-enhanced multimodal digital human semantic reasoning system according to claim 1, characterized in that, Generate corresponding semantic factors of knowledge to be understood based on the type of knowledge to be understood, specifically including the following steps: Basic semantic units are obtained by processing multimodal digital human interaction data to understand them; Based on the type to be understood, each basic semantic unit is assigned a corresponding semantic association attribute. The basic semantic units with semantic association attributes are arranged in an orderly manner according to the interaction logic order to form an initial semantic combination. After performing semantic integrity verification on the initial semantic combination, retain the semantically clear and logically coherent unit combination; The primary and secondary relationships of unit combinations are distinguished according to their semantic importance, and the unit combinations are then regularized and integrated based on these relationships to generate semantic factors of the knowledge to be understood.

4. A knowledge-enhanced multimodal digital human semantic reasoning system according to claim 3, characterized in that, The basic semantic units are obtained by processing multimodal digital human interaction data to be understood, specifically including the following steps: Semantic feature decomposition is performed on the multimodal digital human interaction data to be understood to obtain the interaction semantic dimensions of the types to be understood; Features from each interactive semantic dimension are aggregated to form basic semantic units.

5. A knowledge-enhanced multimodal digital human semantic reasoning system according to claim 1, characterized in that, Based on the type of semantic factors to be understood, the target standard knowledge semantic factors are matched from the multimodal knowledge-enhanced semantic association architecture. This process includes the following steps: The standard set of semantic factors of knowledge is obtained by processing the types of factors to be understood in the semantic factors of knowledge to be understood. The semantic relevance of each standard knowledge semantic factor in the standard knowledge semantic factor set is ranked to obtain the ranking result. Based on the ranking result, the factor with the highest semantic fit is selected as the initial standard knowledge semantic factor. Verify the semantic compatibility between the initial selected standard knowledge semantic factors and the types of factors to be understood; Based on semantic adaptability, the corresponding initial standard knowledge semantic factors are determined as target standard knowledge semantic factors.

6. A knowledge-enhanced multimodal digital human semantic reasoning system according to claim 5, characterized in that, The standard set of semantic factors of knowledge is obtained by processing the types of factors to be understood, which includes the following steps: Based on the type of semantic factors to be understood, hierarchical semantic processing is performed within the multimodal knowledge-enhanced semantic association architecture to obtain the processing results; Based on the processing results, a set of standard knowledge semantic factors that match the type of factors to be understood is obtained.

7. A knowledge-enhanced multimodal digital human semantic reasoning system according to claim 6, characterized in that, The missing knowledge elements are obtained by comparing the elements of the semantic factors of the knowledge to be understood with the elements of the semantic factors of the target standard knowledge factors. This process includes the following steps: By analyzing the relationships between the various factor elements to be understood, a relationship system of the factor elements to be understood can be obtained; Extract the semantic elements of the target standard knowledge semantic factors, and sort out the relationships between the semantic elements of each knowledge semantic factor to form a standard factor element relationship system; The matching degree is obtained by comparing the correlation system of the factor elements to be understood with the correlation system of the standard factor elements at each level; the missing knowledge elements are obtained based on the matching degree.

8. A knowledge-enhanced multimodal digital human semantic reasoning system according to claim 7, characterized in that, The first semantic inference coefficient is obtained based on the knowledge-missing elements of the knowledge-enhanced anomaly semantic factor, specifically including the following steps: Obtain the missing element type and total number of missing elements for the knowledge missing elements; Each missing knowledge element is assigned a missing element weight based on its missing element type, and the missing element weights are summed based on the total number of missing elements to obtain the factor element missing coefficient. Set factor missing weights for knowledge-enhancing anomaly semantic factors based on the type of information to be understood; The first semantic inference coefficient is obtained based on the factor missing weight and the factor element missing coefficient.

9. A knowledge-enhanced multimodal digital human semantic reasoning system according to claim 8, characterized in that, The second semantic inference coefficient is obtained based on the semantic factors of the association standard knowledge, specifically including the following steps: Obtain the total number of associated factors and the total number of associated factor types for the semantic factors of the associated standard knowledge; The first correlation inference coefficient is obtained based on the total number of correlation factors; Assign knowledge type weights to the knowledge semantic factors based on the knowledge semantic factor types of the associated standard knowledge semantic factors, and accumulate the knowledge type weights based on the total number of associated factor types to obtain the second association inference coefficient; The second semantic inference coefficient is obtained based on the first and second association inference coefficients.

10. A knowledge-enhanced multimodal digital human semantic reasoning system according to claim 1, characterized in that, The multimodal digital human semantic reasoning result is output based on the first semantic reasoning coefficient and the second semantic reasoning coefficient, including the following steps: The initial semantic determination result is obtained by constructing a coefficient fusion relationship based on the first semantic inference coefficient and the second semantic inference coefficient. Perform semantic integrity verification on the initial semantic determination results; The valid semantic information after verification is semantically adapted and sorted with the semantic factors of the associated standard knowledge to obtain the sorting result, and the core semantic orientation is determined based on the sorting result. The localization result is obtained by localizing the multimodal digital human interaction intent based on the core semantic orientation; Based on the positioning results, output multimodal digital human semantic reasoning results adapted to multimodal interaction scenarios.