Enterprise-level digital human question and answer method based on unified cognitive architecture

By using a multi-cognitive primitive network with a unified cognitive architecture and a comprehensive energy function, the fragmentation problem of enterprise-level question-answering systems is solved, enabling deep causal reasoning, dynamic emotional interaction, and continuous knowledge evolution, thereby improving the reliability and anthropomorphism of the question-answering system.

CN122489697APending Publication Date: 2026-07-31BEIJING RUNZEYUAN EDUCATION TECH (GRP) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING RUNZEYUAN EDUCATION TECH (GRP) CO LTD
Filing Date
2026-04-24
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

The modular architecture of existing enterprise-level intelligent question-answering systems leads to a disconnect between knowledge representation, logical reasoning, and interactive expression, making it difficult to achieve deep causal reasoning, dynamic emotional adaptation, and continuous knowledge evolution in complex question-answering scenarios for senior executives.

Method used

It adopts a unified cognitive architecture to construct a multi-cognitive primitive network. By driving the network's autonomous evolution through a comprehensive energy function, it achieves unified processing of the entire process from knowledge representation and problem reasoning to multimodal interaction. It introduces causal learning and sparsity constraints to perceive the user's psychological state in real time and supports large-scale parallel computing.

Benefits of technology

It achieves an inherent integration of knowledge, reasoning, and interaction, enhancing the depth, reliability, and naturalness of responses, and possesses continuous learning and adaptability to meet the needs of enterprise business development and personalized interaction by senior executives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122489697A_ABST
    Figure CN122489697A_ABST
Patent Text Reader

Abstract

This application provides an enterprise-level digital human question-answering method based on a unified cognitive architecture, belonging to the field of artificial intelligence and digital human interaction technology. It addresses the problem in traditional modular systems where knowledge management, reasoning generation, and digital human-driven processes are disconnected, resulting in insufficient depth, low reliability, and unnatural interaction in question-answering for corporate executives. This method constructs a unified cognitive primitive network to represent enterprise knowledge, encodes user questions as intent-driven processes within the network, and utilizes a comprehensive energy function that integrates multiple cognitive principles such as goal achievement, causal constraints, and empathic alignment to guide the network state to autonomously and continuously evolve to a stable state. Finally, it decodes and generates semantically consistent text answers and multimodal digital human driving signals in parallel from the same state. This method achieves the intrinsic integration of knowledge, reasoning, and expression, significantly improving the depth, reliability, human-like interactive experience, and continuous adaptive capability of enterprise-level question answering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence and digital human interaction technology, and in particular to an enterprise-level digital human question-answering method based on a unified cognitive architecture. Background Technology

[0002] Currently, intelligent question-answering systems designed for high-level corporate decision-making scenarios typically employ a modular architecture, combining multiple independent technology stacks such as knowledge base management, information retrieval, answer generation, and digital human presentation.

[0003] In existing technologies, retrieval-enhanced generation (RAG) technology combined with digital human-driven approaches is commonly used to construct question-answering systems. The construction of the enterprise knowledge base, the retrieval of user questions and generation of answers, and the voice and animation driving of the digital human are typically designed as independent processing modules, either serially or partially in parallel. Each module operates according to its own technological paradigm; for example, the knowledge base may contain vector databases and relational graphs, the RAG process relies on step-by-step retrieval and language model generation, and the digital human driving relies on independent speech synthesis and animation models.

[0004] However, such modular systems have inherent technical flaws. Information transfer between modules is fragmented and inefficient; knowledge representation, logical reasoning, and interactive expression fail to achieve internal unity. This results in insufficient overall synergy when dealing with complex corporate executive question-and-answer scenarios requiring deep causal reasoning, dynamic emotional adaptation, and continuous knowledge evolution. Consequently, the depth, reliability, and naturalness of responses are limited. Therefore, a technical solution that can fundamentally unify and integrate cognitive processes is urgently needed. Summary of the Invention

[0005] This application provides an enterprise-level digital human question-answering method based on a unified cognitive architecture, which can achieve unified processing of the entire process from knowledge representation and question reasoning to multimodal interaction in an endogenous collaborative and continuously evolving manner, significantly improving the depth, reliability and anthropomorphism of enterprise-level question answering.

[0006] Firstly, this application provides an enterprise-level digital human question-answering method based on a unified cognitive architecture. The method includes the following steps: constructing a unified cognitive primitive network composed of multiple cognitive primitives, wherein each cognitive primitive has a content state, a context state, and an intent state; receiving user-inputted question information and encoding the question information into the intent state of at least one target cognitive primitive in the cognitive primitive network; constraining and driving the overall state of the cognitive primitive network to autonomously evolve through a comprehensive energy function, wherein the comprehensive energy function includes at least one target achievement energy sub-item constructed based on the intent state of the target cognitive primitive; when the state evolution of the cognitive primitive network satisfies preset conditions, decoding the content state of at least one responding cognitive primitive in the network, generating and outputting the corresponding text answer and digital human multimodal driving signal.

[0007] By adopting the above technical solution, this application abandons the traditional module stacking paradigm and transforms the question-and-answer process into an autonomous evolutionary process of the network under comprehensive energy constraints by constructing a unified, stateful cognitive primitive network. Questions are encoded as the network's intention-driven force, driving the entire network state to continuously evolve towards a stable state containing the answer, and ultimately decoding the textual answer and digital human-driven signal from this unified state. This architecture ensures the inherent consistency and synergy of knowledge, reasoning, and expression, laying the foundation for achieving deep and reliable anthropomorphic question-and-answer.

[0008] Furthermore, the integrated energy function also includes multiple mutually cooperating energy sub-items, which include a first energy sub-item for maintaining the consistency of the network's internal state, a second energy sub-item for constraining network connection relationships, and a third energy sub-item for adapting to the user interaction state.

[0009] By adopting the above technical solution and introducing multiple synergistic energy sub-items, the evolution of the network is simultaneously constrained by multiple cognitive principles. The first energy sub-item ensures the logical consistency of knowledge representation and reasoning within the network; the second energy sub-item guides the network to form and follow a sparse causal structure, enhancing the reliability and interpretability of reasoning; and the third energy sub-item enables the network evolution to respond to user states in real time, providing a mechanism to ensure empathetic interaction.

[0010] Furthermore, the second energy sub-item is configured to enable the cognitive primitive network to autonomously learn and maintain the causal relationship between any two connected cognitive primitive state changes during its evolution, and to make the causal relationship sparsity.

[0011] By adopting the above technical solution, this configuration endogenizes causal reasoning capabilities within the network's dynamics. The network autonomously discovers and strengthens key causal relationships during learning and evolution, while sparsity constraints prevent overfitting and logical inconsistencies. This enables the system to automatically extract causal knowledge from enterprise data and make robust causal inferences based on this, meeting the requirements of executive decision-making for depth and logical consistency in answers.

[0012] Furthermore, the construction of the third energy sub-item is based on multimodal perception information generated during real-time user interaction that reflects their psychological or cognitive state, and the multimodal perception information is encoded into specific primitive states in the cognitive primitive network.

[0013] By adopting the above technical solution, a deep integration of real-time perception of user psychological state and internal network representation is achieved. Multimodal perceptual information (such as tone of voice and micro-expressions) is directly transformed into energy constraints that affect network evolution, enabling the final generated answer content and the digital human's expression to proactively adapt to the user's real-time psychological and cognitive state, thereby achieving a deeper level of empathy and natural interaction.

[0014] Furthermore, the steps for constructing the unified cognitive primitive network include: acquiring the enterprise's private operational data; and encoding the associations and patterns contained in the operational data into the initial connection structure and primitive state distribution of the cognitive primitive network through unsupervised learning.

[0015] By adopting the above technical solutions, enterprise-owned, multi-source operational data is naturally precipitated into "innate knowledge" within a unified cognitive architecture through unsupervised learning. This avoids complex artificial knowledge engineering, enabling the system to autonomously construct rich, interconnected, and pattern-rich intrinsic representations from raw data, providing a solid knowledge foundation for subsequent intelligent question answering.

[0016] Furthermore, the overall state of the cognitive primitive network undergoes autonomous evolution, guided by the integrated energy function, so that the network state continuously changes in the direction of decreasing energy.

[0017] By adopting the above technical solution, the core driving force and direction of network evolution were clarified. The entire question-answering process was formalized as a continuous dynamic process of searching for low-energy states (stable solutions) in the energy landscape. This controlled autonomous evolution mechanism enables the system to flexibly and robustly handle complex, unstructured problems.

[0018] Furthermore, the step of decoding the content state of the response cognitive primitive includes: starting from the same content state of the response cognitive primitive, generating the text answer and the digital human multimodal driving signal respectively through parallel text decoding path and multimodal signal decoding path.

[0019] By adopting the above technical solution, it is ensured that the generated text content and the digital human's multimedia performance, such as voice, facial expressions, and movements, originate from the same cognitive state. This homologous decoding mechanism fundamentally guarantees a high degree of consistency between the semantics of the answer and the digital human's expressive form, avoiding the problem of potential disconnect between content and performance in traditional technologies.

[0020] Furthermore, the method also includes the following steps: based on the network state trajectory and result feedback generated by the question-and-answer interaction, incrementally adjust the connection weights and state distributions of the associated cognitive primitives in the cognitive primitive network.

[0021] By adopting the above technical solutions, the system acquires the ability to continuously learn and evolve. Each interaction becomes an opportunity for the system to optimize its internal knowledge network, absorbing new experiences and adapting to new situations in an incremental and fine-tuned manner. This allows the system to grow alongside the company's business development, demonstrating human-like learning adaptability.

[0022] Furthermore, the steps for driving the state evolution of the cognitive primitive network are performed on processing units that support massively parallel and sparse computing.

[0023] By adopting the above technical solution, the adaptability of this method to high-performance computing hardware is demonstrated. The massively parallel computing and dynamic sparse connection characteristics in the unified cognitive architecture are highly compatible with novel processing units that support such computing (such as neuromorphic chips and high-dimensional tensor processors), which is conducive to achieving excellent computing performance and energy efficiency in practical deployments.

[0024] In summary, this application has at least the following beneficial effects: 1. It provides an enterprise-level digital human question-answering method based on a unified cognitive architecture, which realizes the intrinsic integration of knowledge, reasoning and interaction, and significantly improves the depth, reliability and naturalness of the answers; 2. Through endogenous causal learning and sparse constraint energy terms, the system is equipped with the ability to autonomously discover and apply causal logic for deep reasoning from data; 3. By using an empathy alignment energy term based on multimodal perception, digital human interaction achieves real-time and proactive adaptation to the user's psychological state; 4. Through an incremental adjustment mechanism, the system is able to continuously evolve its internal knowledge network based on interactive feedback, thus possessing the ability to continuously learn and adapt.

[0025] It should be understood that the description in the Summary Section is not intended to limit the key or essential features of the embodiments of this application, nor is it intended to restrict the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0026] The above and other features, advantages, and aspects of the embodiments of this application will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 A schematic diagram of an exemplary operating environment in which embodiments of this application can be implemented is shown; Figure 2 A flowchart of an enterprise-level digital human question-answering method based on a unified cognitive architecture is shown in an embodiment of this application. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0028] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0029] This application provides an enterprise-level digital human question-answering method based on a unified cognitive architecture. Through a unified knowledge-reasoning-expression architecture and an energy-minimization driving mechanism, it fundamentally solves the problem of collaborative fragmentation in traditional modular systems, realizes the inherent unity of deep reasoning of answers, dynamic emotional interaction and continuous knowledge evolution, and significantly improves the reliability, anthropomorphism and decision support capabilities of enterprise-level intelligent question answering.

[0030] Figure 1 A schematic diagram of an exemplary operating environment in which embodiments of this application can be implemented is shown. This environment provides the necessary physical computing, sensory interaction, and security support for the methods described.

[0031] Reference Figure 1 The operating environment includes a core computing cluster providing basic computing power and knowledge storage for the method, an interactive terminal system responsible for collecting user signals and presenting the digital human image, and a network and security system to ensure efficient and secure collaboration among the components. These components are interconnected through a high-speed internal network, forming a dedicated hardware support system that works collaboratively.

[0032] The core computing cluster consists of heterogeneous computing and storage units. The neuromorphic computing chip array is responsible for performing state evolution calculations of the cognitive primitive network in an event-driven, highly parallel, and sparse manner. The high-dimensional tensor processing unit specifically handles the large-scale, intensive tensor operations involved in network evolution. Computational non-volatile memory units are tightly coupled with the aforementioned computing units, enabling integrated storage and high-speed access to cognitive primitive states and network connection weights.

[0033] The interactive terminal system comprises two types of devices: input sensing devices and output presentation devices. Input sensing devices mainly include microphone arrays and multimodal sensors deployed on the user side, used to collect the user's voice and non-verbal interaction signals. Output presentation devices mainly refer to high-resolution display devices and audio playback devices, used to present the system-generated digital human video stream and synchronized voice to the user.

[0034] The network and security architecture includes a high-speed internal network switch connecting the core computing cluster and interactive terminal systems, as well as a security gateway and firewall deployed at the network boundary. The high-speed internal network ensures low-latency transmission of sensor data, control commands, and rendering stream data. The security gateway and firewall isolate, filter, and audit data exchange between the internal and external networks, ensuring the security boundaries of enterprise private data and core algorithm models.

[0035] The various components work collaboratively through clearly defined connections. The user interaction terminal is connected to a secure access point of the network and security system via the enterprise intranet or dedicated lines. The core computing cluster is directly interconnected at high speed with the core switches of the network system. Perceived data is sent to the core cluster for processing via the security system, and the resulting digital human signal stream is sent to the presentation terminal via a reverse path. The entire operating environment forms a closed loop from signal input, intelligent computing to multimodal output, providing the necessary physical support for the question-answering method based on the unified cognitive architecture.

[0036] Based on the above operating environment, this application discloses an enterprise-level digital human question-answering method based on a unified cognitive architecture. Figure 2 A flowchart of an enterprise-level digital human question-answering method based on a unified cognitive architecture is shown in an embodiment of this application.

[0037] Reference Figure 2 The method specifically includes the following steps S1 to S5: S1: Construct a unified cognitive primitive network consisting of multiple cognitive primitives, wherein each cognitive primitive has a content state, a context state, and an intent state.

[0038] The specific steps of this method include: acquiring the enterprise's private operational data; and encoding the relationships and patterns contained in the operational data into the initial connection structure and primitive state distribution of the cognitive primitive network through unsupervised learning. The private operational data includes, but is not limited to, internal enterprise documents, structured databases, meeting minutes, and historical communication records. This data is preprocessed and transformed into normalized text sequences and numerical feature vectors that can be processed by the model. Each cognitive primitive is represented by a triplet. Characterization, in which This is a content state vector used to encode the semantic concept represented by the primitive; This is a context state vector, used to represent the associated information of the concept in the current environment; This is the intention state vector, used to drive the network to evolve toward a specific goal. The preset embedding space dimension is a hyperparameter determined based on model capacity and computational resources, for example... or .

[0039] The core of network construction lies in determining the initial dynamic connection weights between primitives and the prior states of each primitive through unsupervised learning. Specifically, suppose the network has a total of Each cognitive primitive has an initial content state matrix denoted as... . The value is pre-set based on the enterprise's knowledge scale, for example, from several thousand to tens of thousands. By applying self-supervised methods such as contrastive learning or masked language modeling to the operational data, a mapping function can be learned. This makes each semantic unit extracted from the data All can be mapped to a content state vector These vectors constitute The initial estimate. The connectivity between primitives is determined by a dynamic sparse connectivity weight matrix. It indicates that its elements Defined from primitives To the basics The intensity of the impact. The initial values ​​are obtained by calculating the mutual information between primitive content states or by applying statistical correlation based on attention weights, and then... Regularization is used to maintain its sparsity, i.e., only those significantly above the threshold are retained. The connection, It is a preset small positive value (e.g.) This is used to control the initial connection density of the network, forming the initial sparse topology. Context state and intention state Initially, it can be set to a zero vector or a small random perturbation related to the content state, waiting for subsequent interaction processes to drive it. Ultimately, the initial state of the cognitive primitive network is determined by... The complete definition provides a unified, evolvable, distributed neural representation foundation for enterprise knowledge.

[0040] S2: Receive the question information input by the user and encode the question information into the intention state of at least one target cognitive primitive in the cognitive primitive network.

[0041] Specifically, the question information entered by the user The signal is captured via microphone or text interface; if it is an audio signal, it is converted into a text sequence via an Automatic Speech Recognition (ASR) module. To inject the question information into the unified cognitive primitive network, it needs to be mapped to the network's semantic space. First, the same embedding function as in step S1 is used. Text sequence Encode into a query vector Specifically, this can be achieved by averaging the word vectors in the sequence or through a dedicated query encoder. This query vector Used to locate relevant primitives in a network.

[0042] Next, calculate the query vector. The current content state of all cognitive primitives in the network The similarity is calculated using a dot product attention mechanism to generate an attention distribution. : , in For the embedding dimension consistent with step S1, This is a scaling factor used to prevent the gradient from vanishing due to an excessively large dot product. It also helps in selecting attention weights. The highest front Each element ( , For example, a preset positive integer. ) as a set of "goal cognitive primitives" For each target primitive Its intentional state vector It will be updated to encode the goal-driven information implied in the user's question. The update operation is performed through a learnable intent injection function. This function integrates the current intent state of the primitive with the query vector: , in and These are learnable parameters, randomly set during system initialization and optimized through subsequent learning. This represents vector concatenation. This update operation explicitly sets the semantics of the user query as the dynamic intent of these target primitives, thereby establishing a directional problem-solving "target potential field" at the network level. Simultaneously, to reflect the immediate context of the current dialogue, the contextual state of these target primitives... It can also be updated synchronously, for example, set as a query vector. Or its variants. At this point, the user's abstract question has been transformed into a concrete change in the state of a specific primitive within the network, providing a clear starting condition for subsequent goal-driven autonomous evolution.

[0043] S3: The overall state of the cognitive primitive network is constrained and driven to evolve autonomously through a comprehensive energy function, wherein the comprehensive energy function includes at least one target achievement energy sub-item constructed based on the intention state of the target cognitive primitive.

[0044] In this step, the integrated energy function further includes multiple mutually cooperating energy sub-items. These energy sub-items include a first energy sub-item for maintaining the consistency of the network's internal state, a second energy sub-item for constraining network connectivity, and a third energy sub-item for adapting to user interaction states. The second energy sub-item is configured to enable the cognitive primitive network to autonomously learn and maintain the causal relationship between any two connected cognitive primitive state changes during its evolution, ensuring that the causal relationship is sparse. The third energy sub-item is constructed based on multimodal perception information generated during real-time user interaction that reflects their psychological or cognitive state, and this multimodal perception information is encoded into specific primitive states within the cognitive primitive network. The autonomous evolution of the overall state of the cognitive primitive network is guided by the integrated energy function, causing the network state to continuously change in the direction of decreasing energy. The step of driving the evolution of the cognitive primitive network state is executed on a processing unit that supports massively parallel and sparse computation.

[0045] The comprehensive energy function Defined as the weighted sum of multiple energy components, its mathematical expression is: , in, , , , These correspond to the energy sub-items of goal achievement, prediction consistency, causal constraint, and empathy alignment, respectively. These are pre-defined non-negative weighting coefficients used to balance the relative importance of each sub-item in the evolution. These coefficients are determined by grid search or optimization algorithms on the validation set before system deployment.

[0046] The target achieves the energy sub-item It originates directly from the goal set in step S2. Let... For the set of target cognitive primitives, this sub-item is defined as the current content state of all target primitives. Its intended state The sum of squares of the differences between them: , in, This represents the L2 norm of the vector. Its function is to generate a potential gradient, driving the network's content state to evolve in a direction that matches the user's question intent; it is the core driving force of the question-answering process.

[0047] The first energy sub-item, namely the prediction consistency energy sub-item This is used to maintain the consistency of the dynamics within the network. It is based on the assumption that a primitive, given its own context state... Given the states of its connected neighbors, it should be able to predict changes in its own content state. For each primitive... The prediction error is defined as: , in, It is a parameterized prediction function (such as a small neural network). It is the connection weight threshold, which is a preset small positive number (e.g. ), Indicates the basic element Significant connections (weights) Greater than the threshold The set of current content states of all neighboring primitives. This energy component is the sum of the prediction errors of all primitives: Its function is to constrain changes in the network state to conform to the internal patterns it has learned, thus preventing the evolutionary process from falling into chaotic or meaningless states.

[0048] The second energy sub-item This is key to achieving endogenous causal reasoning. For any pair of connected primitives in the network... and (Right now or ), defined in a tiny time interval Within, the changes in its content state are respectively and , It is calculated from the state evolution history. Encouraging the existence of a sparse, explainable causal mechanism. This describes the relationship of change while imposing sparsity constraints on the mechanism itself. Its specific form is: , in, It is the set of all potential causal primitive pairs to be examined, typically containing all weights. Connection pairs with a value greater than zero. It is a hyperparameter that controls the sparsity intensity. ), preset before training. The L1 norm of a matrix (i.e., the sum of the absolute values ​​of its elements) is used to promote... A large number of elements tend to zero, thus yielding a sparse causal explanation. This sub-item allows the network to spontaneously organize its connections and state changes during evolution, causing important influencing relationships to be obscured by a few significant ones. Matrix element capture allows for the emergence of interpretable causal structures.

[0049] The third energy sub-item It is responsible for integrating real-time interactive emotions. First, it collects the user's raw multimodal signals (such as image frames and audio waveforms) through sensors such as cameras and microphones. Then, through a pre-trained feature extraction network (such as a CNN for facial expression action unit recognition and an acoustic model for speech emotion analysis), these signals are transformed into a fusion vector representing the user's current psychological and cognitive state. , This refers to the fixed dimension of the feature extraction network output. Within the network, a set of "empathy primitives" is pre-defined. their quantity and Dimensions There is a match or a mapping relationship. Encoding the content state of these primitives, for example (in (Preset fixed or finely adjustable projection parameters). .Then, Defined as the set of "expression primitives" related to the network's current generated response. The state and the set of "empathy primitives" A measure of the difference between states: , in, It can be Euclidean distance or cosine distance. This sub-item guides the network in generating the answer (influences...). When the primitive state is in a certain state, its internal expression state is coordinated with the perceived psychological state of the user, thus providing the underlying drive for the final digital human to express empathic tone and expression.

[0050] The autonomous evolution process of the network as a whole, that is, the state of all its primitives. Over time The change is due to the total energy mentioned above. The gradient flow dominates and is subject to random perturbations. Its continuous-time dynamics can be described by the Langevin equation of the following form: , in, Representative element joint state vector . It is the time constant, a preset positive parameter used to control the speed of state evolution. It is the equivalent "temperature" parameter ( ), control random noise The intensity of (usually modeled as Gaussian white noise with zero mean and variance as an identity matrix) is used to help the network escape local energy minima. This is the gradient of the total energy over all state variables of the primitive. The equation indicates that the network state drifts along the direction of energy decrease (the negative gradient direction) and is also affected by noise perturbations. By iteratively solving this equation at discrete time steps using numerical methods (such as the Euler-Maruyama method), the autonomous and continuous evolution of the network state can be achieved. The entire evolution process is executed on processing units that support massively parallel and sparse computation (such as neuromorphic computing chips), because the computation mainly involves large-scale sparse matrix-vector multiplication (calculating the gradient). ) and element-level operations are well-suited to the efficient parallel processing paradigm of this type of hardware.

[0051] S4: When the state evolution of the cognitive primitive network meets the preset conditions, the content state of at least one response cognitive primitive in the network is decoded to generate and output the corresponding text answer and digital human multimodal driving signal.

[0052] In this step, decoding the content state of the response cognitive primitive includes: starting from the same content state of the response cognitive primitive, generating the text answer and the digital human multimodal driving signal respectively through parallel text decoding path and multimodal signal decoding path.

[0053] The preset condition refers to the state evolution of the cognitive primitive network reaching a dynamic equilibrium or convergence state. Specifically, when the total energy of the network... The rate of change is less than a preset threshold ( ,For example ),Right now or network state vector The norm change is less than the threshold over multiple consecutive iteration steps. ( When the condition is met, the evolution is considered to satisfy the criteria. At this point, the network's state is considered to contain a stable internal "cognitive" result in response to the user's question.

[0054] "Response cognitive primitives" refer to the content state of a primitive at the end of its evolution. Related to the "Goal Achievement Energy Sub-item" The most closely related primitive subset Specifically, this can be achieved by checking the status of each primitive's content. The target intention state set in step S2 ( The cosine similarity is used to determine the order, and the top results with the highest similarity are selected. Each element ( For example, a preset positive integer. )constitute The content state of these primitives It aggregates comprehensive semantic information that has evolved to answer questions.

[0055] From the same content status The parallel decoding process is crucial for ensuring the inherent consistency between the answer content and its presentation in this scheme. The text decoding pathway is responsible for generating the natural language answer. First, the aggregated response content state... (For example, The weighted average of the inputs is used as the initial context, and a pre-trained autoregressive language model (such as the Transformer decoder) is input. The model predicts the probability distribution of the next word step by step based on this initial state and the generated preceding word sequence. Let the vocabulary size be... In generating the first When there are 100 words, the model's calculation is as follows: , , in, It is a word The embedding vectors are derived from the embedding layer of a pre-trained language model. It is the hidden state of the decoder at the current position. It is the hidden layer dimension of the decoder (a model architecture hyperparameter). and These are the parameters of the output layer, derived from a pre-trained model. They are sampled from this distribution using strategies such as beam search, ultimately generating the word sequence. That is, the text answer.

[0056] The multimodal signal decoding path is responsible for generating synchronization signals that drive the digital human's speech, lip movements, facial expressions, and body gestures. This path also aggregates content states. As a common input source. First, a speech synthesis decoder. Will Mapped to a sequence of acoustic features, such as a Mel spectrogram. ,in It is the number of time frames (determined by the length of the generated content). This is the Mel-band number (a fixed acoustic feature dimension, such as 80). This process can be represented as: , in These are the trainable parameters for the speech synthesis model. Simultaneously, a neural rendering-driven decoder... From the same The visual driving parameter sequence is decoded. These parameters include: facial action unit intensity vector sequence. ( (Number of action units, e.g., 46), head rotation Euler angle sequence and the rotation sequence of key joints in the upper body ( (Number of joints). For the number of frames in the visual sequence, and the audio duration Alignment is achieved through a preset frame rate. The decoding process is as follows: , in These are the trainable parameters for the vision-driven model. It is a temporal generative model (such as a conditional variational autoencoder or a diffusion model) that ensures that the generated visual parameters are consistent with the speech features. Align with in time, and with The semantic content (such as emotion and emphasis) is consistent. Ultimately, Converted to waveform via vocoder ,and The input is then fed into the digital human rendering engine, which drives it to generate corresponding lip movements, facial expressions, and actions to synthesize the video stream. Text answer Audio and video The output is synchronized to form a complete digital human response with a high degree of unity in content and presentation.

[0057] S5: Based on the network state trajectory and result feedback generated by the question-and-answer interaction, the connection weights and state distributions of the associated cognitive primitives in the cognitive primitive network are incrementally adjusted.

[0058] Specifically, each complete question-and-answer interaction (from step S2 to S4) generates a learning opportunity. The network state trajectory refers to the historical record of the changes in the network's key state variables over time throughout the entire evolution process (step S3), which can be abstractly represented as a sequence. ,in Includes time The state of all primitives and connection weight matrix A snapshot of the moment. This is the time when the evolution terminates. The results are then fed back. It is a scalar reward signal used to evaluate the overall effectiveness of this interaction. It can be obtained in several ways: explicit ratings from users (such as five-star ratings, normalized to...) The results can be obtained through intervals, implicit behavior analysis (such as the duration a user listens to an answer, whether they immediately ask follow-up questions or interrupt), or by an automatic evaluation module based on indicators such as the consistency between the answer and the facts in the knowledge base and the completeness of information.

[0059] The goal of incremental adjustments is to leverage the experience gained from a single interaction. The network parameters are optimized in a minimal, online manner, enabling the network to evolve to a better solution or provide a better interactive experience when faced with similar problems in the future. The adjustments primarily target two types of parameters: one is the dynamic connection weight matrix between primitives. It encodes knowledge associations and causal structures; secondly, it represents the prior distribution of key primitive states, which can be reflected in the state of primitive content. bias terms Or mean vector.

[0060] The adjustment process follows a gradient-based online learning paradigm. First, an immediate loss function related to the result of this interaction is defined. An efficient construction method is to combine it with the comprehensive energy function in step S3. Final state value and feedback signal Connect them. For example: , in, It is the total energy value at the end of the evolution. It is a positive reward (the higher the better). It is a positive coefficient used to balance the two objectives of feedback reward and energy minimization, and its value is a preset hyperparameter.

[0061] Next, the loss function is calculated. Relative to the parameter to be adjusted (Include Effective connection weights and associated state bias parameters in gradient of ) .because Since the gradient calculation depends on the entire state trajectory, it requires backpropagation along the trajectory, which can be achieved using a time-truncated backpropagation algorithm or a more advanced stochastic computation graph method. Using the calculated gradient, the parameters are then slightly updated using stochastic gradient descent (SGD) or its variants (such as the Adam optimizer with momentum). , in, It is a very small learning rate (e.g.) arrive (on a scale of magnitude), which ensures the "incremental" nature of updates, meaning that each interaction causes only a small perturbation to the parameters, avoiding catastrophic forgetting.

[0062] Specifically, for connection weights After the update, sparsity constraints are usually applied again, for example, by thresholding to remove values ​​whose absolute values ​​are less than a certain small amount. ( ,For example The weights of the network can be reset to zero, or a soft threshold operator can be applied, to maintain the causal interpretability and computational efficiency of the network structure. Adjustments to the state distribution, such as those of a specific primitive... It remains activated upon successful response, and its state prior (bias) It can be appropriately enhanced to make the primitive easier to activate in the future.

[0063] This incremental adjustment mechanism enables the unified cognitive primitive network to fine-tune itself based on each "experience," much like a biological system, achieving continuous, lifelong learning and adaptation. The network's knowledge and reasoning patterns are no longer static but evolve and optimize with ongoing interaction with executive users, thus becoming increasingly aligned with the company's specific business context and the executive's personalized interaction preferences.

[0064] To enable those skilled in the art to more fully understand this application, a simplified additional embodiment is provided below to demonstrate how the unified cognitive architecture is applied in a specific business scenario.

[0065] Suppose a tech company executive asks, "Why did our market share in cloud computing decline last quarter?"

[0066] In step S1, the network has constructed cognitive primitives and their initial associations that include concepts such as "R&D investment", "marketing expenses", "product performance", "market share" and "customer satisfaction" by learning data such as the company's annual financial reports, market analysis reports and competitor dynamics.

[0067] In step S2, the problem is encoded, relevant primitives such as "market share", "cloud computing", and "last quarter" in the network are activated as target primitives, and their intent state is set to "seeking the cause of the decline".

[0068] In step S3, the network evolves under the influence of integrated energy. The driving network is used to find the reasons associated with the "market share decline"; This prompts the network to follow learned causal chains (such as "marketing expense reduction") Market voice decline The reduction in new customer acquisition will be traced. Based on the frowning expressions of executives captured by cameras, the network was guided to adopt a more cautious and analytical internal expression when interpreting information.

[0069] In step S4, after the network evolution stabilizes, the response primitive (whose content state integrates information such as "marketing expense data," "competitor dynamics," and "product iteration cycle") is decoded. The text decoder generates a structured report: "Mainly due to a 15% year-on-year decrease in Q2 marketing budget, while competitor A released a major upgrade..." Simultaneously, the multimodal decoder, based on the same content state, drives the digital human to simultaneously broadcast the answer with a serious, concerned expression, accompanied by chart pointing gestures.

[0070] In step S5, the causal path from "marketing expenses" to "market share" was strongly activated during this interaction. Based on the positive feedback from executives (such as nods of approval) received in this response, the system fine-tuned the relevant causal connections along this interaction trajectory. The weight of the "marketing expenses" primitive was increased, and the state prior of the primitive was slightly enhanced, making the response to similar problems in the future more sensitive.

[0071] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to the embodiments of this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.

[0072] By constructing a unified cognitive primitive network and encoding enterprise privatized operational data into its initial connection structure and state distribution, this method first establishes a distributed neural foundation capable of deeply representing the intrinsic connections and patterns of enterprise knowledge. This step transforms the originally scattered and heterogeneous enterprise knowledge into a unified, computable, and dynamically evolving semantic network, providing a structured knowledge foundation for subsequent intelligent question answering.

[0073] When a user's question is received and encoded into the intent state of a specific target primitive in the network, the system accurately maps an abstract natural language question into a quantifiable driving goal within the network. This establishes a clear starting point for goal-driven network evolution, ensuring that subsequent complex computational processes always revolve around the user's intent, and guaranteeing the relevance of question answering from a mechanistic perspective.

[0074] Subsequently, the network's autonomous evolution process, guided by the comprehensive energy function, becomes the core of achieving deep reasoning and intelligent decision-making. The goal-achievement energy sub-item continuously drives the network state to converge in a direction consistent with the user's intentions; the prediction consistency energy sub-item constrains the network to follow its learned internal logical patterns, avoiding the generation of illogical answers; the causal constraint energy sub-item prompts the network to autonomously discover and strengthen sparse causal relationships during evolution, making the reasoning process interpretable; and the empathy alignment energy sub-item incorporates the real-time perceived user psychological state into the evolutionary constraints, aligning the network's internal state with the user's emotions. The synergistic effect of these energy sub-items enables the network to spontaneously organize information, trace causality, and weigh contradictions through a continuous dynamic process, ultimately stabilizing in an optimized state that satisfies the problem objective, conforms to knowledge logic, and adapts to the user's emotions, all while balancing multiple cognitive principles.

[0075] When the network evolves to a stable state, it decodes the content states of the same set of response primitives in parallel to generate text answers and multimodal driving signals for the digital human, fundamentally ensuring the inherent unity of the output content in terms of semantics and presentation. The text decoding pathway generates accurate and structured language descriptions, while the multimodal decoding pathway generates matching speech, facial expressions, and gestures based on the same semantic source. This allows the digital human to present not only broadcasts but also emotional and focused performances, greatly improving the naturalness of the interaction and the efficiency of information transmission.

[0076] Finally, an incremental adjustment mechanism based on each interaction trajectory and feedback transforms the entire system from a static knowledge base and model into an organism capable of continuous learning from experience. Subtle optimizations to network connection weights and state distribution allow it to gradually adapt to the evolving business needs of the enterprise and the personal preferences of senior executives, thereby achieving continuous evolution in question-answering capabilities and interactive experience.

[0077] In summary, from unified knowledge representation, precise intent mapping, multi-constraint co-evolution, to homogeneous multimodal generation and continuous learning optimization, this series of technical approaches are interconnected and progressive. Together, they enable the system to generate deep, reliable, and causally traceable answers to complex enterprise problems, and to express itself naturally through a highly anthropomorphic and emotionally resonant digital human. Ultimately, this achieves a systematic improvement in the question-answering system's knowledge depth, reasoning reliability, naturalness of interaction, and adaptability in the context of enterprise executives.

[0078] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the foregoing disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A unified cognitive architecture based enterprise level digital human question answering method, characterized in that, Includes the following steps: Construct a unified cognitive primitive network consisting of multiple cognitive primitives, wherein each cognitive primitive has a content state, a context state, and an intent state; Receive user-inputted question information and encode the question information as the intention state of at least one target cognitive primitive in the cognitive primitive network; The overall state of the cognitive primitive network is constrained and driven to evolve autonomously through a comprehensive energy function, wherein the comprehensive energy function includes at least one target achievement energy sub-item constructed based on the intention state of the target cognitive primitive; When the state evolution of the cognitive primitive network meets the preset conditions, the content state of at least one response cognitive primitive in the network is decoded to generate and output the corresponding text answer and digital human multimodal driving signal.

2. The method of claim 1, wherein, The comprehensive energy function also includes several mutually synergistic energy sub-terms. The multiple energy sub-items include a first energy sub-item for maintaining the consistency of the network's internal state, a second energy sub-item for constraining network connectivity, and a third energy sub-item for adapting to user interaction states.

3. The method of claim 2, wherein, The second energy sub-item is configured as follows: This enables the cognitive primitive network to autonomously learn and maintain the causal relationship between any two connected cognitive primitive state changes during its evolution, and makes the causal relationship sparse.

4. The method of claim 2, wherein, The construction of the third energy sub-item, Based on multimodal perception information generated during real-time user interaction that reflects their psychological or cognitive state, the multimodal perception information is encoded into specific primitive states in the cognitive primitive network.

5. The method of claim 1, wherein, The steps for constructing the unified cognitive primitive network include: Obtain the company's private operational data; Through unsupervised learning, the associations and patterns contained in the operational data are encoded into the initial connection structure and primitive state distribution of the cognitive primitive network.

6. The method of claim 1, wherein, The cognitive primitive network undergoes an autonomous evolution of its overall state. Guided by the aforementioned integrated energy function, the network state continuously changes in the direction of decreasing energy.

7. The method of claim 1, wherein, The steps for decoding the content state of the response cognitive primitive include: Starting from the same content state of the response cognitive primitive, the text answer and the digital human multimodal driving signal are generated respectively through parallel text decoding path and multimodal signal decoding path.

8. The method of claim 1, wherein, The method further includes the following steps: Based on the network state trajectory and result feedback generated by question-and-answer interaction, the connection weights and state distributions of the associated cognitive primitives in the cognitive primitive network are incrementally adjusted.

9. The method of claim 1, wherein, The steps that drive the state evolution of the cognitive primitive network. Executes on processing units that support massively parallel and sparse computing.