Digital human tour guide voice generation method, system and device and storage medium

By building a user mental model and real-time calibration mechanism, the problem of lagging explanation strategies in the digital human tour guide system was solved, and the stability of user experience and the continuous optimization of personalized effects were achieved.

CN120803263APending Publication Date: 2025-10-17NANJING NICEBRIDGE INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510903843.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

The existing digital human tour guide system lacks the forward-looking prediction capability and self-calibration mechanism for the user's future mental state, resulting in a lag in the adjustment of explanation strategies and difficulty in continuously optimizing personalized effects.

Method used

By establishing a closed-loop mechanism of prediction-planning-verification-calibration, we build a user mental model, predict the future impact of different narrative paths, generate multimodal explanation content, and calibrate the model through real-time feedback to achieve dynamic calibration of the user's mental state.

Benefits of technology

It achieves forward-looking planning of explanation strategies, improves the stability and personalization of user experience, and the system has the ability to self-correct and learn, and can continuously optimize user cognition and interest preference judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803263A_ABST
    Figure CN120803263A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent speech synthesis and emotion calculation, and discloses a digital human tour guide speech generation method, system and device and a storage medium, the digital human tour guide speech generation method comprises the following steps: S1, constructing a user mental model comprising an initial knowledge state of a user for a knowledge graph; s2, on the basis of the model, predicting and evaluating candidate narrative paths, planning an optimal path and determining an expected mental state; s3, multi-modal explanation content is generated and broadcasted according to the optimal path; s4, collecting real-time feedback of the user to obtain a real mental state; and S5, comparing the real state with the expected state, calculating a prediction deviation, and dynamically calibrating the user mental model for subsequent planning according to the prediction deviation. According to the method, prospective path planning is carried out by constructing the mental model, closed-loop calibration and robustness evaluation are combined, and personalized explanation which is accurate, stable and free of lag adjustment is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent speech synthesis and affective computing, and particularly relates to a digital guide voice generation method, system, device and storage medium. BACKGROUND

[0002] With the advancement of digital construction of cultural sites such as museums and tourist attractions, digital guide as a new carrier that can provide personalized and interactive interpretation services has attracted widespread attention. How to make the interpretation content of digital guide not only accurate in information, but also dynamically adapt to the cognitive level and interest changes of different users, so as to provide continuous and high-quality interactive experience, is a key problem that the current technical field is committed to solving.

[0003] In the prior art, there are some technical solutions for realizing personalized interpretation of digital guide. A common solution is to recommend fixed interpretation routes and content based on user-pre-set interest tags (such as "history", "architecture", "person"). Another solution introduces affective recognition technology in the interaction process, analyzes the user's facial expressions or voice tone to judge his current emotional state (such as happy, confused), and selects the corresponding interpretation script from the pre-set content library for broadcast accordingly, for example, when detecting that the user shows confusion, the current content is simplified or supplemented.

[0004] Although the above prior art realizes the personalization of the explanation to some extent, there are still some deficiencies: the interactive adjustment mechanism of the prior art has inherent hysteresis. This is because the running logic of these technical solutions is reactive, that is, the system must first observe that the user's negative state (such as declining interest or difficulty in understanding) has occurred before triggering the adjustment strategy. The fundamental reason for this "happening first and then making up" mode is that the system lacks the ability to predict the trend of the user's mental state, and cannot foresee the comprehensive impact of a specific sequence of explanation content on the user's future cognitive load and interest, so its adjustment cannot be anticipatory to avoid it after the user experience has been damaged. In addition, the self-adaptive ability of the prior art is limited, and it is difficult to deepen the understanding of the user during use. The reason for this is that the system lacks a closed-loop verification mechanism for self-evaluation and correction, which cannot quantitatively compare the expected results of the system on the user's state with the actual feedback of the user, and use this deviation to calibrate its internal user model. This results in the initial judgment of the system on the user once deviating, which will continue to exist in the whole interaction process, and cannot realize dynamic and accurate self-adaptation. Finally, the decision basis of the prior art when selecting content is relatively single, which makes the stability of the explanation process insufficient. The decision model is usually based on maximizing a single expected indicator (such as interest degree), without considering the uncertainty or risk of the decision itself, so the system may choose an explanation path that has high expected return but also has high potential cognitive load, resulting in a sharp fluctuation in user experience. SUMMARY

[0005] The purpose of the present application is to provide a digital human tour guide voice generation method, system, device and storage medium, which solves the problem that the existing digital human tour guide voice generation method lacks the ability to predict the future mental state of the user and the self-calibration mechanism, resulting in the hysteresis of the explanation strategy adjustment, and the difficulty in continuously optimizing the personalization effect.

[0006] To achieve the above purpose, the present application is realized by the following technical solutions: the first aspect of the present application provides a digital human tour guide voice generation method, which realizes the anticipatory prediction of the user's mental state and the active planning of the narrative strategy by establishing a prediction-planning-verification-calibration closed-loop mechanism, and has self-calibration ability.

[0007] The method comprises the following steps: firstly, constructing an initial user mental model as a basis for subsequent planning and calibration, the initial user mental model comprising an initial knowledge state of a user on a knowledge node in a preset knowledge graph; then, using the initial user mental model, predicting and evaluating the future influence of different candidate narrative paths on the user to plan an optimal narrative path and synchronously determine an expected mental state associated with the optimal narrative path; then, generating multi-modal explanation content according to the planned optimal narrative path and playing the multi-modal explanation content to the user; thereafter, collecting real-time feedback of the user to obtain a real mental state of the user in response to the playing; finally, comparing the real mental state with the expected mental state, calculating a prediction deviation therebetween, and dynamically calibrating the user mental model based on the prediction deviation, the calibrated user mental model being used for subsequent narrative path planning.

[0008] In an optional embodiment, the step of using the initial user mental model to predict and evaluate the future influence of different candidate narrative paths on the user to plan an optimal narrative path specifically comprises: generating a group of candidate narrative paths based on the user mental model; for each candidate narrative path, constructing an emotional potential field thereof to predict a probability distribution of the user mental state under the candidate narrative path; performing robustness evaluation on the emotional potential field of each candidate narrative path, and selecting a candidate narrative path with an optimal robustness evaluation result as the optimal narrative path.

[0009] In an optional embodiment, the robustness evaluation on the emotional potential field of each candidate narrative path is performed by an evaluation function calculated by the following formula: In the formula, R(P i ) is the evaluation function; P i is the candidate narrative path; is an expected value; Var[·] is a variance; IL is an interest arousing degree; CL is a cognitive load; S u is the user mental state; w IL , w CL and λ are preset weights or coefficients. The evaluation function selects a narrative path with the lowest risk and the highest comprehensive benefit by comprehensively considering the expected interest benefit, the cognitive load cost and the uncertainty of the prediction result of the explanation content.

[0010] In an optional implementation, the step of calculating the prediction deviation between the real mental state and the expected mental state, and dynamically calibrating the user mental model based on the prediction deviation, specifically comprises: representing the real mental state as a real state vector, and representing the expected mental state as an expected state vector containing the same dimension indicators; obtaining the prediction deviation vector by performing vector subtraction operation on the real state vector and the expected state vector; performing attribution analysis based on the prediction deviation vector to determine the cause of the deviation; and updating the user mental model according to the determined cause.

[0011] In an optional implementation, the calculation of the prediction deviation vector is expressed by the following formula: wherein, is the prediction deviation vector; S actual is the real mental state collected at time point τ; is the expected mental state at time point τ. The prediction deviation vector quantifies the difference between the prediction result and the actual result, providing a quantitative basis for subsequent attribution analysis and model calibration.

[0012] The second aspect of the present application provides a digital human tour guide voice generation system for executing the method of any one of the preceding aspects, the system comprising: a mental model construction unit configured to construct an initial user mental model; a narrative path planning unit connected to the mental model construction unit and configured to plan an optimal narrative path and determine an expected mental state corresponding to the optimal narrative path based on the user mental model by predicting the influence of different candidate narrative paths on the user's future mental state; a multi-modal content generation unit connected to the narrative path planning unit and configured to generate multi-modal explanation content according to the optimal narrative path and play it back; a user state acquisition unit configured to acquire the real mental state of the user after playback; a model calibration unit connected to the narrative path planning unit and the user state acquisition unit and configured to perform narrative closed-loop verification and dynamically calibrate the user mental model by comparing the real mental state with the expected mental state.

[0013] The third aspect of the present application provides a computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the method of the first aspect of the present application.

[0014] The fourth aspect of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method of the first aspect of the present application.

[0015] In summary, the present application includes at least one of the following beneficial technical effects: 1. The present application realizes the forward-looking planning of explanation strategy by constructing the user's mental model and predicting the impact of different paths on the user's future mental state before narrative planning. This active prediction and planning mechanism can avoid the explanation paths that may lead to excessive cognitive load or decreased interest of the user in advance, rather than passively responding after the negative state occurs, thereby maintaining the user's better cognitive and emotional state during the explanation process and avoiding the damage to the user experience caused by adjustment lag.

[0016] 2. The present application introduces the mechanism of narrative closed-loop verification and mental model self-calibration. After each explanation, the prediction deviation is calculated by comparing the user's real mental state with the system's expected mental state, and the user's mental model is dynamically calibrated accordingly. This closed-loop feedback process enables the system to have the ability of self-correction and learning, and can continuously improve the accuracy of judging the user's cognitive and interest preferences during use, making the subsequent personalized explanation more and more accurate.

[0017] 3. The present application evaluates the emotional potential field of the candidate narrative path by robustness, and its decision not only considers the expected return, but also quantifies the uncertainty and risk of the prediction result. This robust-based decision-making method makes the system tend to choose the scheme with more stable comprehensive performance and lower potential risk when planning the explanation path, improves the smoothness and reliability of the explanation process, and avoids the user's discomfort caused by drastic or inappropriate style switching. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 Fig. 1 is a schematic diagram of the system structure of the present application; Figure 2 Fig. 2 is a schematic diagram of the process of constructing the initial user mental model of the present application; Figure 3 Fig. 3 is a schematic diagram of the process of narrative path planning and expected state determination of the present application; Figure 4 Fig. 4 is a schematic diagram of the process of multi-modal explanation content generation and broadcasting of the present application; Figure 5 Fig. 5 is a schematic diagram of the process of real mental state acquisition of the present application; Figure 6 Fig. 6 is a schematic diagram of the process of narrative closed-loop verification and mental model dynamic calibration of the present application. DETAILED DESCRIPTION

[0019] The following will be described in combination with the drawingsFigure 1 -Appendix Figure 6 The application will be described in further detail.

[0020] Please refer to the Appendix Figure 1 , Figure 1 is a structural schematic diagram of a system according to an embodiment of the application. The application provides a digital human tour guide voice generation system, which is deployed in a computing environment including a client device and a server. The client device can be a smartphone, a tablet computer, AR glasses or an interactive display screen. The server can be a single server or a cluster composed of multiple servers. The system can include: A mental model construction unit 10 configured to perform step S1 of the method of the application. The unit receives user registration information from the client device, or acquires a static portrait of the user through initial interaction, which includes data such as cultural background and knowledge level. At the same time, the unit retrieves and invokes a preset knowledge graph from the database of the server according to the specific scenic spot identifier selected by the user. The knowledge nodes in the knowledge graph are annotated with cognitive dependency relationships. The mental model construction unit 10 assesses the initial mastery of each knowledge node by the user based on the static portrait and the knowledge graph, and determines the interest weight of the user for different topics in combination with the historical interaction data of the user, and finally outputs a structured initial user mental model.

[0021] A narrative path planning unit 20 configured to perform step S2 of the method of the application. The unit receives the user mental model output by the mental model construction unit 10. Based on the model, the unit generates a set of candidate narrative paths, each path being a sequence of explanation nodes in the future period of time. For each candidate narrative path, the unit predicts the probability distribution of the user's mental state under the path by constructing an emotional potential field.

[0022] To select the optimal one from the candidate narrative paths, the narrative path planning unit 20 performs robust evaluation on the emotional potential field of each candidate narrative path. In a specific embodiment, the evaluation is realized by an evaluation function R(P i ), which is calculated as follows: In the formula, R(P i ) is the evaluation function; P i is the candidate narrative path; is the expected value; Var[·] is the variance; IL is the interest arousal degree; CL is the cognitive load; S u is the user's mental state; w IL , w CL and λ are preset weights or coefficients. The narrative path planning unit 20 selects R(P i) the path with the highest calculation result as the optimal narrative path, and determine the mental state probability distribution corresponding to the optimal narrative path as the expected mental state, and then output the optimal narrative path to the multi-modal content generation unit 30 and output the expected mental state to the model calibration unit 50.

[0023] The multi-modal content generation unit 30 is configured to perform step S3 of the method of the present application. The unit receives the optimal narrative path output by the narrative path planning unit 20. It first expands and converts the sequence of knowledge nodes in the path into a natural language form of a script. Subsequently, the unit performs cross-cultural adaptation processing on the script, and performs speech synthesis through a multi-voice scene-based text-to-speech (TTS) engine according to the emotional tone of the script. At the same time, the generated speech signal is used to drive the lip movements and facial expressions of a digital human visual model, and finally the synthesized speech and synchronized visual animation are combined into multi-modal explanation content, which is played to the user through the client device.

[0024] The user state acquisition unit 40 is configured to perform step S4 of the method of the present application. The unit acquires real-time feedback signals containing user speech or facial expressions through the sensors (such as microphones or cameras) of the client device. The unit extracts features from the acquired signals to obtain emotional features or cognitive state features, and quantifies the cognitive load and interest arousal degree of the user based on these features, and finally combines these indicators into a real mental state vector consistent with the expected mental state dimension, and outputs it to the model calibration unit 50.

[0025] The model calibration unit 50 is configured to perform step S5 of the method of the present application. The unit receives the expected mental state output by the narrative path planning unit 20 and the real mental state output by the user state acquisition unit 40. The unit calculates the prediction deviation by comparing the two, and dynamically calibrates the user mental model based on the prediction deviation. In a specific embodiment, the prediction deviation is quantified by a prediction deviation vector , which is calculated as follows: In the formula, is the prediction deviation vector; S actual is the real mental state acquired at time point τ; is the expected mental state at time point τ.

[0026] The model calibration unit 50 performs attribution analysis based on the prediction deviation vector to determine the cause of the deviation, and updates the internal parameters of the user mental model maintained by the mental model construction unit 10 according to the determined cause.

[0027] The calibrated user mental model will be used by the narrative path planning unit 20 in the next round of narrative path planning, thus forming a continuous optimization closed loop.

[0028] Please refer to the attached Figure 2 , Figure 2 is a flowchart of the initial user mental model construction according to an embodiment of the present application. First, step S1 is performed, i.e. constructing an initial user mental model as the basis for subsequent planning and calibration, which is performed by the mental model construction unit 10 in the aforementioned system implementation. The specific process is as follows: First, when the user uses the service provided by the present application through the client device, for example, when registering an account or first configuring, the system obtains the static portrait of the user. The static portrait is a set of data used to describe the stable characteristics of the user. In a specific embodiment, it can include nationality information (used to determine the cultural background) and education level information (used to preliminarily determine the knowledge level) input by the user, as well as the selected interface language, etc. These information constitutes the basis for the preliminary understanding of the user.

[0029] Next, the system obtains the unique identifier of a specific scenic spot (e.g. the Palace Museum) selected by the user on the client interface. The mental model construction unit 10 uses the identifier to retrieve and call the knowledge graph matching the scenic spot from the database deployed on the server side. The knowledge graph is a structured knowledge base, which includes: multiple knowledge nodes, each node representing a specific explanation entity, such as a scenic spot, a historical figure or a historical event; and cognitive dependency relationships annotated between knowledge nodes, which clearly indicate the prerequisite knowledge required to understand a node.

[0030] Then, the mental model construction unit 10 generates the initial knowledge state of the user based on the obtained user static portrait and the called knowledge graph. This process is a quantitative evaluation of the user's current knowledge mastery level of the scenic spot. In an embodiment, the system assigns an initial mastery probability value to each knowledge node in the knowledge graph according to the knowledge level in the user's static portrait. For example, for basic or common sense nodes, a higher initial mastery probability is assigned; for professional or detailed nodes, a lower initial mastery probability is assigned, thus forming a probability vector representing the initial knowledge state of the user.

[0031] Meanwhile, the mental model building unit 10 also generates an initial interest profile of the user. This profile is used to describe the user’s preference degree for different explanation topics. The generation of this profile can combine the static profile and the user’s historical interaction data. The historical interaction data can include the user’s listening time length, like behavior, or collection behavior for different categories of content (such as architectural art, historical stories, and biographies of figures) when using similar applications in the past. The system assigns corresponding interest weights to different topic fields by analyzing these data, forming an initial interest profile vector.

[0032] Finally, the mental model building unit 10 structurally combines the static profile, the initial knowledge state, and the initial interest profile generated in the foregoing steps to form a complete initial user mental model.

[0033] Please refer to the accompanying Figure 3 , Figure 3 is a narrative path planning and expected state determination flowchart according to an embodiment of the present application. After the initial user mental model is built, the system performs step S2, that is, using the user mental model, the future impact of different candidate narrative paths on the user is predicted and evaluated to plan the optimal narrative path and simultaneously determine the expected mental state associated with the optimal narrative path. This step is performed by the narrative path planning unit 20 in the foregoing system implementation, and the specific process is as follows: First, the narrative path planning unit 20 generates a set of candidate narrative paths based on the user mental model output by the mental model building unit 10 and the position of the current explanation in the knowledge graph. Each candidate narrative path is an ordered sequence composed of multiple knowledge nodes. In an embodiment, the generation of the path follows the cognitive dependency relationship marked in the knowledge graph, for example, a path that deeply explains the details of the current node, a path that explains interesting historical anecdotes related to the current node, and a path that switches to the next logically associated scenic spot can be generated, thereby forming a set of candidate paths containing multiple paths with logical differences.

[0034] Next, for each candidate narrative path, the narrative path planning unit 20 constructs an emotional potential field for it. This process is a forward-looking simulation aimed at predicting the possible changes in the user’s mental state in the future period of time if the system chooses this path for explanation.

[0035] The construction of the emotional potential field is based on the current user mental model and combines the attributes (such as complexity and interest) of each knowledge node in the path. The output result is a probability distribution describing the user’s future mental state, which can be represented by a multi-dimensional vector containing at least two indicators of cognitive load and interest arousal degree.

[0036] Then, the narrative path planning unit 20 performs a robustness evaluation on the affective potential field of each candidate narrative path to select the optimal one. The evaluation process is not simply to choose the path with the highest expected return, but to consider the return, cost and risk of the path comprehensively. In a specific embodiment, the robustness evaluation is done by executing an evaluation function R(P i ) which is calculated as follows: where R(P i ) is the evaluation function; P i is the candidate narrative path; is the expected value; Var[·] is the variance; IL is the interest level; CL is the cognitive load; S u is the user’s mental state; w IL , w CL and l are preset weights or coefficients. The narrative path planning unit 20 calculates the R(P i ) score of all candidate paths and selects the one with the highest score as the optimal narrative path.

[0037] Finally, after the optimal narrative path is determined, the narrative path planning unit 20 stores the predicted mental state probability distribution in the affective potential field of the optimal narrative path as the expected mental state. The expected mental state is a time series which records the system’s expectation of the user’s mental state indicators (such as the expected value and variance of the cognitive load and interest level) at future time points. This data will be output to the model calibration unit 50 for subsequent comparison and verification. Meanwhile, the determined optimal narrative path is output to the multi-modal content generation unit 30 for generating specific explanation content.

[0038] Please refer to the accompanying Figure 4 , Figure 4 is a flowchart of the multi-modal explanation content generation and delivery according to an embodiment of the present application. After the optimal narrative path is planned, the system performs step S3, i.e. generates multi-modal explanation content according to the planned optimal narrative path and delivers it to the user. This step is performed by the multi-modal content generation unit 30 in the aforementioned system implementation, and the specific process is as follows: First, the multi-modal content generation unit 30 receives the optimal narrative path output by the narrative path planning unit 20. The path is an ordered sequence containing multiple knowledge nodes. The unit expands each knowledge node in the sequence into a complete natural language sentence or paragraph according to the detailed text description stored in the knowledge graph, and organizes these sentences or paragraphs in sequence to form an initial explanation script which is structurally complete and logically coherent.

[0039] Next, the multi-modal content generation unit 30 performs a cross-cultural adaptation process on the generated initial script. This process aims to eliminate potential understanding barriers caused by cultural differences. In one specific embodiment, the system retrieves the user's static profile, particularly the cultural background information therein, from the user's mental model. The system uses this information to scan the script for specific terms, allusions, or metaphors that are incompatible with the cultural background. If such content is detected, the system queries a pre-defined equivalent expression library and replaces it with a more understandable vocabulary or sentence structure in the user's cultural context.

[0040] Then, the system performs a multi-modal content synthesis based on the adapted script. This process includes two parallel parts: audio synthesis and visual driving. In terms of audio synthesis, the system inputs the script into a multi-voice prosody-based TTS engine. The engine selects matching voice timbre, tone, and speed parameters according to the emotional tone (e.g., flat statement, interesting explanation, serious introduction) pre-labeled for different paragraphs in the script, to generate a high-performance speech stream.

[0041] In terms of visual driving, the system uses the generated speech stream and its timestamp information to synchronously drive a three-dimensional digital human visual model. Specifically, the phoneme sequence in the speech signal is used to calculate and drive the digital human model's lip animation in real time, achieving precise synchronization of sound and lip movement. At the same time, the prosodic features (e.g., pitch, energy) in the speech signal and the emotional tone labeled in the script are used to drive the digital human model's facial expressions and head posture, generating visual performance consistent with the emotional tone of the speech.

[0042] Finally, the multi-modal content generation unit 30 encapsulates the synthesized speech stream and the synchronized visual animation stream to form the final multi-modal explanation content. This content is transmitted to the user's client device and played on its display screen and audio output device, providing the user with a visually and aurally consistent tour service.

[0043] Please refer to the accompanying drawings Figure 5 , Figure 5 is a flowchart of real mental state acquisition according to one embodiment of the present application. After the multi-modal explanation content is played to the user, the system performs step S4, i.e., in response to the playing, acquires the user's real-time feedback to obtain his real mental state. This step is performed by the user state acquisition unit 40 in the aforementioned system implementation, and the specific process is as follows: First, the user state acquisition unit 40 collects real-time feedback signals containing user reactions through sensors integrated on the client device. In one specific embodiment, the unit invokes the camera of the client device to capture a video stream containing facial expressions of the user, and invokes the microphone to capture an audio stream containing speech (e.g., spontaneous comments, questions or sighs) of the user. These two data streams constitute the raw real-time feedback signals.

[0044] Next, the user state acquisition unit 40 processes the collected real-time feedback signals to extract quantified affective features or cognitive state features. For the video stream, the system applies a facial expression analysis model that first detects key feature points of the face, and then identifies the activation status and intensity of facial action units (AUs) associated with specific emotions or cognitive states (e.g., confusion, surprise, focus). For the audio stream, the system applies a speech signal processing module that analyzes prosodic features in the speech, such as pitch variation, energy fluctuation and speech rate, and can further obtain the textual content of the user’s utterance through speech recognition.

[0045] Then, the user state acquisition unit 40 quantifies high-level semantic indicators such as the user’s cognitive load and interest arousal degree based on the extracted features. This process maps low-level features to high-level state indicators.

[0046] For example, if the facial expression analysis model detects that the user’s eyebrows are tightly furrowed (specific AUs are activated), while the speech signal processing module detects that the user utters a tone or an utterance indicating a question, the system combines these features and maps them to a higher “cognitive load” indicator value. Conversely, if the user’s mouth is detected to be upturned and the user’s gaze is detected to be steadily fixed on the screen, it is mapped to a higher “interest arousal” indicator value.

[0047] Finally, the user state acquisition unit 40 combines these quantified indicators into a structured real mental state vector. The dimensionality of this vector is consistent with that of the expected mental state vector generated in step S2, to ensure the subsequent comparability of the two. The real mental state vector is then output to the model calibration unit 50 for performing the next step of the method of the present application.

[0048] Please refer to the accompanying drawings Figure 6 , Figure 6 is a flowchart of the narrative closed-loop verification and dynamic calibration of the mental model according to an embodiment of the present application. After collecting the real mental state of the user, the system performs step S5, i.e., by comparing the real mental state with the expected mental state, calculates the prediction deviation between the two, and based on the prediction deviation, dynamically calibrates the user’s mental model. This step is performed by the model calibration unit 50 in the aforementioned system implementation, and the specific process is as follows: First, the model calibration unit 50 receives the real mental state output by the user state acquisition unit 40, and retrieves from the storage the expected mental state corresponding to the announced path, which is determined by the narrative path planning unit 20 in step S2. Both states are represented as vectors containing the same dimension indicators, for example, both contain cognitive load and interest arousal degree as two components.

[0049] Next, the model calibration unit 50 obtains a quantitative prediction deviation vector by performing vector subtraction operation on the real state vector and the expected state vector. In a specific embodiment, the prediction deviation vector is calculated as follows: wherein, is the prediction deviation vector; S actual (τ) is the real mental state acquired at time point τ; is the expected mental state at time point τ.

[0050] The value of each component of the prediction deviation vector directly reflects the degree and direction of the difference between the system prediction and the actual user response.

[0051] Then, the model calibration unit 50 performs attribution analysis based on the obtained prediction deviation vector to determine the most likely cause of the deviation.

[0052] The purpose of attribution analysis is to trace the observed deviation to a specific internal module or parameter of the system; for example, if the prediction deviation vector shows that the actual cognitive load of the user is significantly higher than the expected value, the system will analyze the reason: one possibility is that the mastery level of a certain prerequisite knowledge node in the user's mental model is overestimated; another possibility is that the complexity of the current explanation content exceeds the preset value of the model. This analysis process can be performed based on a preset rule base or a probability model.

[0053] Finally, the model calibration unit 50 performs targeted dynamic calibration on the user's mental model according to the cause determined by the attribution analysis.

[0054] This calibration is specific and targeted; for example, if the attribution analysis determines that the deviation is due to overestimation of the user's mastery of a certain background knowledge, the model calibration unit 50 will modify the user's mental model by specifically reducing the user's mastery probability value for that specific knowledge node. If the cause is determined to be the absence of a cognitive dependency relationship between two nodes in the knowledge graph, the system can add a to-be-corrected label to this relationship.

[0055] After dynamic calibration, the updated user mental model will replace the original initial user mental model and be used in the next round of narrative path planning. By repeatedly performing steps S2 to S5, the method of the present application forms a complete prediction-planning-verification-calibration closed loop, so that the system can continuously improve its accuracy of user state prediction and adaptability of personalized explanation in continuous interaction with the user.

[0056] The present application also provides a computer device, comprising a processor and a memory, the memory storing a computer program executable by the processor, the computer program being executed by the processor to perform the method as above.

[0057] The present application also provides a computer device, comprising a memory, a processor and a computer program stored on the memory and executable on the processor, the processor executing the computer program to implement the method of the present application.

[0058] The present application also provides a computer readable storage medium, having a computer program stored thereon, the computer program being executed by a processor to implement the method of the present application.

[0059] Although embodiments of the present application have been shown and described, it is to be understood that various modifications, substitutions, replacements and changes can be made to these embodiments without departing from the principles and spirit of the present application, the scope of the present application being defined by the appended claims and their equivalents.

Claims

1. A method for generating digital human tour guide voice, characterized in that: The following steps are involved: S1. Construct an initial user mental model as a basis for subsequent planning and calibration, wherein the initial user mental model includes the user's initial knowledge state of knowledge nodes in a preset knowledge graph; S2. Utilizing the initial user mental model, predicting and evaluating the future impact of different candidate narrative paths on the user, thereby planning an optimal narrative path, and simultaneously determining an expected mental state associated with the optimal narrative path; S3. Generate multimodal explanation content based on the planned optimal narrative path and broadcast it to the user; S4. In response to the broadcast, collecting real-time feedback from the user to obtain the user's true mental state; S5. Comparing the actual mental state with the expected mental state, calculating the prediction deviation between the two, and dynamically calibrating the user mental model based on the prediction deviation. The calibrated user mental model is used for subsequent narrative path planning.

2. A method for generating digital human tour guide voice according to claim 1, characterized in that: In step S1, the step of constructing an initial user mental model as a basis for subsequent planning and calibration includes: Obtaining a static profile of the user through the information entered by the user during client registration, the static profile including cultural background and knowledge level; Retrieving and retrieving a knowledge graph matching the identifier of the specific scenic spot from a database of the server; the knowledge graph includes: a plurality of knowledge nodes representing scenic spots, people, or events, and cognitive dependency relationships annotated between the knowledge nodes for indicating an order of understanding; Based on the static portrait and the preset knowledge graph, the user's initial mastery of each knowledge node in the knowledge graph is evaluated to generate the user's initial knowledge state; Combining the static profile with the user's historical interaction data, determining the user's interest weights in different subject areas to generate an initial interest profile of the user; The static portrait, the initial knowledge state and the initial interest portrait are combined to form the initial user mental model.

3. The method for generating digital human tour guide voice according to claim 2, wherein: In step S2, the steps of using the initial user mental model to predict and evaluate the future impact of different candidate narrative paths on the user to plan an optimal narrative path and simultaneously determining the expected mental state associated with the optimal narrative path include: generating a set of candidate narrative paths based on the user mental model; For each candidate narrative path, construct its emotional potential field to predict the probability distribution of the user's mental state under the candidate narrative path; Performing a robustness evaluation on the emotional potential field of each candidate narrative path, and selecting the candidate narrative path with the best robustness evaluation result as the optimal narrative path; The mental state probability distribution predicted in the emotional potential field corresponding to the optimal narrative path is determined and stored as the expected mental state.

4. The method for generating digital human tour guide voice according to claim 3, characterized in that: The robustness evaluation of the emotional potential field of each candidate narrative path is performed, and the calculation of the evaluation function is expressed by the following formula: Where, R(P i ) is the evaluation function; P i for candidate narrative paths; is the expected value; Var[·] is the variance; IL is the interest stimulation; CL is the cognitive load; S u is the user's mental state; w IL 、w CL and λ are preset weights or coefficients.

5. The method for generating digital human tour guide voice according to claim 4, characterized in that: In step S3, the step of generating multimodal explanation content based on the planned optimal narrative path and broadcasting it to the user includes: Expanding the knowledge node sequence in the optimal narrative path into an explanation script; Performing cross-cultural adaptation on the explanation script according to the static portrait in the user mental model; According to the emotional tonality of the explanation script, multi-timbre situational speech synthesis is performed, and the lip shape and facial expression of the digital human model are driven synchronously to complete the generation and broadcasting of the multimodal explanation content.

6. The method for generating digital human tour guide voice according to claim 5, characterized in that: In step S4, the step of collecting the user's real-time feedback in response to the broadcast to obtain the user's true mental state includes: Collect real-time feedback signals including user voice or facial expressions through microphones or cameras; Processing the real-time feedback signal to extract emotional features or cognitive state features; Based on the emotional characteristics or cognitive state characteristics, indicators of the user's cognitive load and interest arousal are quantitatively generated, and the indicators are combined into the real mental state.

7. The method for generating digital human tour guide voice according to claim 6, characterized in that: In step S5, the steps of comparing the actual mental state with the expected mental state, calculating a prediction deviation between the two, and dynamically calibrating the user mental model based on the prediction deviation include: Representing the actual mental state as an actual state vector and representing the desired mental state as an desired state vector containing the same dimensional indices; The prediction deviation vector is obtained by performing a vector subtraction operation on the real state vector and the expected state vector; an attribution analysis is performed based on the prediction deviation vector to determine the cause of the deviation, and the prediction deviation between the real mental state and the expected mental state is calculated by calculating the prediction deviation vector The formula for calculating the prediction deviation vector is: Where, is the prediction deviation vector; S actual (τ) is the actual mental state collected at time τ; is the desired mental state at time τ; The user mental model is updated according to the determined reasons to complete the dynamic calibration of the user mental model.

8. A digital human tour guide voice generation system, characterized in that: A method for generating a digital human tour guide voice according to any one of claims 1 to 7, comprising: A mental model building unit, used to build an initial user mental model; a narrative path planning unit, configured to plan an optimal narrative path based on the user mental model by predicting the impact of different candidate narrative paths on the user's future mental state, and determine an expected mental state corresponding to the optimal narrative path; A multimodal content generation unit, configured to generate multimodal explanation content according to the optimal narrative path and broadcast the content; A user state collection unit is used to collect the user's real mental state after the broadcast; The model calibration unit is used to perform narrative closed-loop verification and dynamically calibrate the user's mental model by comparing the actual mental state with the expected mental state.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Automatic scenic spot explanation content pushing method integrating multi-source perception

    CN121434397A