Intelligent dialogue system and method based on AI multimodal large model

Through multimodal data fusion and reinforcement learning analysis, a user digital avatar model is constructed, which solves the shortcomings of existing intelligent dialogue systems in personalized expression and dynamic interaction, realizes the accurate capture and multi-dimensional evaluation of user emotions and physiological reactions, and improves the adaptability and realism of interaction.

CN120491834BActive Publication Date: 2025-09-12NANCHANG YIJING INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510977222.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-09-12
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

Existing intelligent dialogue systems have shortcomings in user personalized expression cloning and dynamic interaction adjustment. It is difficult to fully capture the user's emotional state and physiological reactions. The evaluation dimension is single and lacks collaborative analysis of multimodal indicators.

Method used

By collecting users' biometric data in real time, using a multimodal fusion encoder to extract speech rhythm, language structure and facial dynamic features, a user digital avatar model is constructed, and personalized expression patterns are cloned using a generative adversarial network. The interaction process is analyzed through a reinforcement learning model to generate a multidimensional evaluation report.

Benefits of technology

It achieves accurate capture of users’ emotional states and subtle physiological reactions, generates highly realistic audio avatars, improves the adaptability and realism of interactions, and provides a scientific interaction capability evaluation and optimization path.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120491834B_ABST
    Figure CN120491834B_ABST
Patent Text Reader

Abstract

The present invention relates to an intelligent dialogue system and method based on an AI multimodal large model. The method comprises collecting a user's biometric data in real time; utilizing a generative adversarial network to clone the user's personalized expression pattern; driving an audio avatar to engage in multiple rounds of dialogue interaction with the user, capturing the user's real-time physiological signals through an expression recognition module, and generating dynamic response content suggestions based on the dialogue context; and analyzing the interaction process based on a reinforcement learning model. This system achieves a multimodal deep fusion of speech prosodic features, language structure features, and facial dynamic features, thus breaking through the limitations of single-modality or simple feature splicing in existing technologies. It can accurately and comprehensively capture a user's emotional state, expression intentions, and subtle physiological reactions, laying a solid foundation for building high-fidelity user representations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and specifically to an intelligent dialogue system and method based on an AI multimodal large model. Background Art

[0002] With the advancement of artificial intelligence (AI), multimodal interaction has become a key development direction for intelligent dialogue systems. Currently, dialogue systems based on large AI models are evolving from single-modality systems (such as text or voice) to systems that integrate multimodal data, including voice, image, and text. Users are increasingly demanding authentic, personalized, and context-sensitive interactions. The industry has gradually recognized that relying solely on single-modal features is insufficient to fully capture a user's emotional state and expressive intent. Consequently, multimodal fusion technologies, personalized expression cloning, and dynamic interaction analysis have become research hotspots, but practical applications still face technical bottlenecks.

[0003] Existing technical solutions typically employ single-modality or simple multimodal fusion architectures, such as dialogue systems based solely on speech recognition and text analysis, or using pre-trained models to implement basic multimodal feature splicing. While these solutions can accomplish basic conversational interactions, they often rely on fixed speech templates or simple prosodic adjustments to personalize user expressions, lacking in-depth modeling of user biometrics (such as facial micro-expressions and dialect accents). During conversational interactions, it is difficult to dynamically adjust response strategies based on real-time physiological signals, and the evaluation of the interaction process often remains limited to a single dimension (such as language fluency), lacking the collaborative analysis of multimodal indicators. Summary of the Invention

[0004] In response to the technical problems existing in the prior art, the present invention provides an intelligent dialogue system and method based on an AI multimodal large model.

[0005] The present invention solves the above technical problems with the following technical solution: an intelligent dialogue method based on an AI multimodal large model, the method comprising:

[0006] Collect user biometric data in real time, including voice stream, facial video stream and interactive text, extract voice prosodic features, language structure features and facial dynamic features through a multimodal fusion encoder, and output a joint feature vector;

[0007] A user digital avatar model is constructed based on the joint feature vector, and a generative adversarial network is used to clone the user's personalized expression pattern, generating an audio avatar that includes mappings of language rhythm, idiomatic vocabulary, and regional accents.

[0008] Loading a preset scenario knowledge base, driving the audio avatar to conduct multiple rounds of conversational interactions with the user, and using the expression recognition module to capture the user's real-time physiological signals, generating dynamic response content suggestions based on the conversation context;

[0009] The interaction process is analyzed based on the reinforcement learning model, and a multi-dimensional evaluation report is generated that includes language redundancy, logical coherence, and micro-expression management indicators. The optimal performance clips are simultaneously stored in the style sample library.

[0010] Another object of the present invention is to provide an intelligent dialogue system based on an AI multimodal large model, the system comprising:

[0011] The biometric data acquisition module is used to collect the user's biometric data in real time, including voice stream, facial video stream and interactive text, and extract the speech prosody features, language structure features and facial dynamic features through a multimodal fusion encoder to output a joint feature vector;

[0012] The avatar model construction module is used to build a user digital avatar model based on the joint feature vector, using a generative adversarial network to clone the user's personalized expression pattern and generate an audio avatar that includes mappings of language rhythm, idiomatic vocabulary, and regional accents;

[0013] The conversation interaction module is used to load the preset scenario knowledge base, drive the audio avatar to conduct multiple rounds of conversational interaction with the user, and capture the user's real-time physiological signals through the expression recognition module, and generate dynamic response content suggestions based on the conversation context;

[0014] The interaction process analysis module is used to analyze the interaction process based on the reinforcement learning model, generate a multi-dimensional evaluation report including language redundancy, logical coherence, and micro-expression management indicators, and simultaneously store the best performance clips in the style sample library.

[0015] The beneficial effects of the present invention are:

[0016] 1. The present invention uses a biometric data acquisition module, a directional beamforming microphone array to capture voice streams, deploys a 3D facial key point detection algorithm to quantify facial dynamics, and parses interactive text based on a dependency syntax analyzer. Finally, a cross-modal Transformer encoder is used to generate a dimensionally aligned joint feature vector, thereby achieving multimodal deep fusion of speech prosodic features, language structure features, and facial dynamic features, thus breaking through the limitations of single modality or simple feature splicing in the existing technology. It can accurately and comprehensively capture the user's emotional state, expression intentions, and subtle physiological reactions, laying a solid foundation for building high-fidelity user representation.

[0017] 2. The present invention uses an avatar model construction module to construct a user digital avatar model based on the joint feature vector, uses a conditional adversarial generative network (GAN) and its discriminator to optimize the similarity of regional accents through a spectral graph dynamic time warping algorithm, and combines the established personalized speech rule library to store user-specific language behavior patterns, thereby achieving high-fidelity cloning of personalized expression patterns such as user language rhythm, idiomatic vocabulary, and regional accent mapping relationships, generating highly realistic audio avatars, and effectively solving the problems of insufficient personalization and style distortion caused by traditional speech synthesis technology relying on fixed templates or only being able to perform simple rhythm adjustments.

[0018] 3. The present invention uses a dialogue interaction module to load a preset scenario knowledge base and drive the audio avatar to conduct multiple rounds of dialogue interactions with the user. In the interview scenario, it activates the industry knowledge graph to recursively generate a technical question chain with dynamically adjusted abstract levels. During speech training, the speech text is converted in real time and the virtual confusion trigger signal of the virtual audience is injected. At the same time, the expression recognition module captures the user's real-time physiological signals (such as confused micro-expressions caused by changes in facial microcurrents), and generates dynamic response content suggestions (such as simplified responses) based on the dialogue context. This significantly improves the system's adaptability and interactive realism in complex and changeable scenarios (such as in-depth technical interviews and high-pressure speech training), and can provide highly targeted interactive support based on the user's real-time status.

[0019] 4. The present invention uses an interaction process analysis module to analyze the entire interaction process based on a reinforcement learning model. It calculates language redundancy indicators (statistically analyzing the proportion of repetitive phrases and the frequency of occurrence of vague qualifiers), constructs logical coherence topology maps and indicators, generates micro-expression management indicators (quantifies the proportion of uncontrolled micro-expressions) and performs dynamic weighted evaluation. It generates a comprehensive evaluation report containing multi-dimensional indicators (language redundancy, logical coherence, and micro-expression management). At the same time, it synchronously marks and stores the best performance segments in the style sample library, forming a closed-loop feedback and capability improvement mechanism of real-time multimodal diagnosis-dynamic strategy optimization-high-quality experience accumulation. This overcomes the shortcomings of existing technologies in terms of single evaluation dimensions and lack of in-depth collaborative analysis of non-verbal signals and logical structures, and provides users with a scientific, comprehensive and accumulative interaction capability evaluation and optimization path. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A flowchart of an intelligent dialogue method based on an AI multimodal large model provided by an embodiment of the present invention;

[0021] Figure 2 A flowchart of collecting user's biometric data in real time provided by an embodiment of the present invention;

[0022] Figure 3A flowchart for generating an audio avatar that includes mappings of language rhythm, idiomatic vocabulary, and regional accents, provided by an embodiment of the present invention;

[0023] Figure 4 A structural block diagram of an intelligent dialogue system based on an AI multimodal large model provided by an embodiment of the present invention;

[0024] Figure 5 A structural block diagram of a biometric data acquisition module provided in an embodiment of the present invention;

[0025] Figure 6 A structural block diagram of an avatar model construction module provided in an embodiment of the present invention;

[0026] Figure 7 A structural block diagram of a dialogue interaction module provided in an embodiment of the present invention;

[0027] Figure 8 This is a structural block diagram of the interaction process analysis module provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0028] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0029] Figure 1 The flowchart of the intelligent dialogue method based on AI multimodal large model provided by the embodiment of the present invention is as follows: Figure 1 As shown, the method includes:

[0030] S100 collects the user's biometric data in real time, including voice stream, facial video stream, and interactive text, extracts speech prosody features, language structure features, and facial dynamic features through a multimodal fusion encoder, and outputs a joint feature vector;

[0031] A directional beamforming microphone array is used to capture voice streams in real time. The array focuses on the user's voice source through an acoustic phase difference algorithm, effectively suppressing ambient noise. Combined with linear predictive coding technology, it can accurately extract the fundamental frequency trajectory (reflecting the fluctuation of intonation) and resonance peak features (characterizing timbre characteristics), and then construct a phoneme model that includes dialect-specific acoustic markers. For example, in dialect scenarios, it can capture regional pronunciation characteristics such as the inability to distinguish between "n" and "l" or the lack of retroflex sounds.

[0032] At the same time, a 3D facial key point detection algorithm is deployed. Based on the stereoscopic vision principle in computer vision, the three-dimensional coordinates of 68 key nodes in the facial video stream (such as the corners of the eyes, nose wings, corners of the mouth, etc.) are tracked, and the eye blinking frequency (reflecting the state of attention), the nasolabial groove displacement vector (reflecting the intensity of emotions) and the dynamic sequence of zygomatic muscle contraction intensity (associated with pleasure) are quantified. For example, when the user is thinking, the algorithm can capture the subtle movement combination of raising the eyebrows and slightly closing the eyelids.

[0033] In interactive text analysis, the dependency parser not only parses the subject-verb-object structure of sentences, but also annotates the distribution density of interrogative sentences (reflecting communication initiative) and the intensity of sentiment-related connectives (such as the transition strength of "although...but..."). For example, in a debate scenario, it can identify users' frequent use of interrogative sentences such as "Isn't it?" to enhance persuasiveness. Finally, through the cross-modal Transformer encoder, a self-attention mechanism is used to align the dimensions of three types of features: speech, face, and text. This encoder uses a multi-head attention mechanism to process data from different modalities in parallel. For example, it maps the acoustic features of speech (such as the Mel-spectrogram), facial action unit features (AU features), and word vector features of text into the same high-dimensional space, generating a joint feature vector that contains temporal dynamic correlations, achieving a multi-dimensional feature fusion of "speech rhythm, facial expression, and semantic logic."

[0034] This step builds a three-dimensional perception system combining physiological signals, verbal behavior, and emotional state. The synergy between directional microphones and 3D facial detection enables the system to accurately capture non-verbal signals even in noisy environments. For example, in a conference room, beamforming technology can focus on the target user's voice, even amidst background chatter, while facial landmark detection eliminates distracting facial expressions from other participants.

[0035] The multimodal fusion encoder breaks the semantic limitations of single-modal features. For example, when a user says "I'm fine", if frequent eye blinking and tense displacement of the nasolabial groove are detected at the same time, the system can combine the falling tone characteristics in the speech rhythm to determine that the user may have hidden confusion rather than a literally expressed affirmative attitude.

[0036] This cross-modal feature fusion provides a high-fidelity user representation foundation for the subsequent construction of a digital avatar model. For example, in a speech training scenario, the system continuously collects the speaker's speech pause patterns (such as the filler word "hmm" that appears once every 15 seconds), the frequency of zygomatic muscle contraction (8-10 smiles per minute), and the use of logical conjunctions such as "first... secondly..." in the text. This can generate highly personalized expression pattern clones, enabling the audio avatar to reproduce the user's characteristic of faster speech speed when nervous (the number of syllables per second increases by 20%).

[0037] In addition, the real-time dynamic feature extraction mechanism supports millisecond-level response. For example, in an interview scenario, when the interviewer asks technical questions, the system can complete the multimodal feature capture of the user's voice fundamental frequency mutation (increase of 50Hz), pupil diameter reduction (change range of 1.2mm), and the high-frequency appearance of vague qualifiers such as "maybe" and "probably" in the text within 0.5 seconds, providing instant data support for the subsequent dialogue interaction module to generate targeted response suggestions. This multi-dimensional perception capability significantly enhances the intelligent dialogue system's depth of understanding of the user's true intentions and avoids semantic misjudgments that may be caused by single-modal analysis.

[0038] like Figure 2 As shown, the real-time collection of the user's biometric data specifically includes:

[0039] S110 uses a directional beamforming microphone array to capture speech streams, extracts fundamental frequency trajectory and formant features through linear predictive coding, and constructs a phoneme model that includes dialect-specific acoustic markers;

[0040] S120 deploys a 3D facial landmark detection algorithm to quantify the eye blink frequency, nasolabial fold displacement vector, and zygomatic muscle contraction intensity dynamic sequence from facial video streams;

[0041] S130: Parse the interactive text based on the dependency parser and annotate the distribution density of rhetorical questions and sentiment-oriented connectives;

[0042] S140: Input the above-mentioned voice stream, facial video stream and interactive text into the cross-modal Transformer encoder to generate a dimensionally aligned joint feature vector.

[0043] S200 builds a user digital avatar model based on the joint feature vector, using a generative adversarial network to clone the user's personalized expression pattern and generate an audio avatar that includes mappings of language rhythm, idiomatic vocabulary, and regional accents.

[0044] The conditional generative adversarial network (cGAN) is the core framework. Its generator receives the joint feature vector output by the multimodal fusion encoder (covering speech prosody, facial dynamics and text semantic features) and scene labels (such as "interview", "speech", etc.), and generates dialect text sequences with prosody tags through a multi-layer transposed convolutional network.

[0045] The discriminator uses the spectrogram dynamic time warping (DTW) algorithm to align and compare the Mel spectrogram of the generated samples with that of the real speech samples in terms of time series, and focuses on optimizing the similarity of regional accent feature points (such as the conversion probability of flat and retroflex sounds, the proportion of pronunciation duration of nasal finals, etc.). For example, when processing the audio clone of Cantonese users, the discriminator will iteratively optimize the pronunciation features of the initial consonant "ng" to gradually approximate the accent similarity of the generated speech to the real samples. The personalized speech rule library adopts a dynamic update mechanism to store the user's unique language behavior patterns in real time, such as the association model between the occurrence frequency of the filler word "um" and the thinking duration, the change curve of the rising amplitude of the interrogative sentence ending with the emotional fluctuation, and the non-linear mapping relationship of the speech rate fluctuation under the tense state.

[0046] In specific implementation, when the user is making a technical report, the system will synchronously record the usage frequency of "so" when elaborating on key points, the duration distribution of pauses between sentences (such as a 1.5-second pause every 10 seconds), and the change in the contraction intensity of the zygomaticus major muscle. After clustering analysis, these data form personalized rules to guide the prosody generation of the audio clone.

[0047] This step constructs a three-layer mapping system of "physiological characteristics - language patterns - scene adaptability". The adversarial training mechanism of cGAN enables the audio clone to dynamically fit the subtle differences in the user's expression. For example, in the cross-dialect communication scenario, the system can automatically generate speech with the prosody features of the target dialect according to the acoustic markers in the joint feature vector, while retaining the user's original language rhythm habits, achieving the effect of "changing the accent while maintaining the expression style".

[0048] The introduction of the personalized speech rule library endows the clone model with the ability to remember long-tail language features. For example, it can reproduce the non-linear change that the speech rate increases while the filler words decrease when the user is excited. This dynamic mapping ability is difficult to achieve by traditional voice cloning technologies. In practical applications, this technology can significantly enhance the interaction realism of virtual assistants. For example, the audio clone can reproduce the unique soothing intonation (such as the rising sentence ending and the frequent use of "please rest assured") of customer service staff when handling complaints based on historical interaction data, and dynamically adjust the speech rate according to the current user's emotional state; in the field of education and training, the clone model can simulate the teaching style of excellent lecturers, including the speech pause pattern during writing on the blackboard, the stress pattern of key content, and even the slight speech errors when nervous, providing an immersive learning experience for students.

[0049] This high-fidelity personalized expression cloning breaks through the bottleneck of traditional speech synthesis technology of "similar timbre but distorted style", enabling the audio clone to maintain consistent personality characteristics in cross-scene interactions, enhancing the user's emotional resonance and trust.

[0050] such as Figure 3As shown, generating an audio clone containing a mapping relationship between language rhythm, idiomatic vocabulary, and regional accents specifically includes:

[0051] S210, constructing a conditional adversarial generation network;

[0052] The generator receives the joint feature vector and scene label and outputs a dialect text sequence with prosody tags.

[0053] The discriminator compares real user speech samples with generated samples and optimizes regional accent similarity using the spectrogram dynamic time warping algorithm;

[0054] S220: Establish a personalized speech rule library to store the user's unique filler word usage frequency, question ending tone rising pattern, and speech speed fluctuation curve under stress.

[0055] S300 loads a preset scenario knowledge base, drives the audio avatar to engage in multiple rounds of conversational interaction with the user, and uses the expression recognition module to capture the user's real-time physiological signals and generate dynamic response content suggestions based on the conversation context;

[0056] The pre-set scenario knowledge base uses a hybrid architecture of knowledge graphs and dynamic rules. For example, in an interview scenario, the industry knowledge graph includes hierarchical relationships in the technology stack (e.g., the knowledge chain of "deep learning-convolutional neural network-attention mechanism"). Based on technical keywords mentioned in the user's response (e.g., "Transformer architecture"), the system generates a chain of follow-up questions through a recursive traversal algorithm. This dynamically adjusts the level of abstraction from fundamental principles to engineering applications. Furthermore, based on the user's proficiency in "model compression" from past interactions, the system automatically avoids knowledge points that the user has already mastered, forming a personalized follow-up question path.

[0057] In the speech training scenario, the real-time speech text conversion unit converts the written speech into a mixed spoken style text based on the user's accustomed spoken vocabulary library, and at the same time injects the confusion trigger signal of the virtual audience - this signal is generated by simulating the cognitive load threshold of the real audience. For example, when three professional terms are detected in succession in the speech, the virtual confusion signal is automatically triggered, driving the virtual confusion signal monitoring unit to capture the changes in microcurrent on the user's face (collected by the skin electrode activity sensor). When the confused micro-expression combination of continuous contraction of the corrugator muscle (accompanied by drooping eyelids) is identified and the duration exceeds the preset threshold, the system generates a simplified response suggestion based on the semantic similarity algorithm, such as rewriting "back propagation of convolutional neural network" as "core adjustment process of neural network learning."

[0058] The facial expression recognition module uses multi-channel physiological signal fusion technology. It uses infrared cameras to capture changes in thermal radiation generated by micro-twitching facial muscles. This information, combined with facial skin deformation data acquired by fiber optic sensors, constructs dynamic feature vectors for micro-expressions. For example, in a medical consultation scenario, when a user describes their symptoms, a complex expression of nasal flaring (increased breathing rate) and orbicularis oris muscle tension (slightly pursed lips) may appear. Based on the contextual description of "persistent pain," the system identifies possible anxiety and adjusts the audio avatar's response tone (slowing down the speech and adding soothing phrases). It also generates an interactive suggestion: "Do you need more detailed treatment options?" The dynamic response generation process relies on a cross-modal attention mechanism, which weightedly fuses physiological signal features captured in real time (such as heart rate variability), semantic representations in the conversation history (extracted through the BERT model), and preset response templates in the scenario knowledge base. For example, in legal debate training, when a user's Adam's apple rolls frequently when citing a law (a signal of tension), the system will prioritize retrieving the simplified explanation template of "Legal Application Cases" in the knowledge base, and at the same time add transition words such as "as previously stated" to the response text to enhance logical coherence.

[0059] The dynamic activation mechanism of the preset scenario knowledge base enables the audio avatar to adapt to the professional conversation needs in different fields. For example, in technical interviews, the knowledge graph-driven question chain can accurately locate the candidate's technical blind spots. In sales negotiation simulations, the customer objection handling templates in the knowledge base can push response techniques in real time based on the hesitation characteristics in the user's voice (such as a sudden increase in the frequency of filler words).

[0060] The joint modeling of physiological signals and conversation context enables the system to have the ability to "perceive implicit needs". For example, in education and training scenarios, when a student answers a question and shows a combination of expressions such as dilated pupils (focused attention) and slight contraction of the zygomatic muscles (pleasant feedback), the system will automatically increase the difficulty of the question. Traditional dialogue systems can only respond based on text content and cannot capture such non-verbal signals.

[0061] Virtual confusion signal injection technology creates a realistic interactive pressure environment, enabling trainees to adapt to real audience feedback within a simulated scenario. For example, under the stimulation of the system's continuous injection of confusion signals, a speaker can gradually improve their ability to express professional content in a colloquial manner. This immersive training effect is difficult to achieve with static text-based dialogue systems. In practical applications, this technology can significantly improve the efficiency and realism of skills training.

[0062] For example, in emergency response training for aviation security officers, the audio clone can dynamically trigger different types of conflict scenarios (such as simulating drunk passengers causing trouble) based on the trainees' responses. At the same time, it uses eye tracking technology to capture the trainees' gaze shift patterns (to determine whether they are paying attention to key handling steps) and generate responses containing action instruction suggestions. In language learning scenarios, the system adjusts the difficulty of the conversation in real time based on the learners' pronunciation characteristics (obtained through voice stream analysis) and facial expressions (to determine their understanding of grammatical difficulties). For example, when a confused expression is detected, it automatically switches to an explanation mode with pictures and text.

[0063] This dynamic response mechanism, which integrates multimodal physiological signals and scene knowledge, breaks through the limitations of the traditional intelligent dialogue system's "preset script-matching response" model, enabling the interaction process to possess situational adaptability and emotional sensitivity similar to human dialogue, and providing more effective technical support for areas such as professional skills training and psychological intervention.

[0064] In this embodiment, the discriminator compares the real user speech sample with the generated sample and optimizes the regional accent similarity using the spectrogram dynamic time warping algorithm, specifically including:

[0065] S2101 performs multi-scale voiceprint resonance analysis on real user speech samples, extracting a set of dialect-specific acoustic invariants to generate a dialect voiceprint quantum map. The dialect voiceprint quantum map includes the quantized energy level distribution of tone glide trajectories, the chaotic attractor characteristics of plosive oral air pressure pulses, and the resonant cavity topology fingerprint of nasalized vowels.

[0066] S2102: Obtain a synthesized speech sample, map the synthesized speech sample to a dialect acoustic manifold space, perform variational encoding based on the dialect voiceprint quantum map, identify spectral feature offset regions, and quantify the degree of acoustic distortion to form an acoustic anomaly tensor field containing distortion hotspots;

[0067] S2103: Constructing a spatiotemporal regular grid for adaptive dialect linking within an acoustic anomaly tensor field containing distortion hotspots. This generates a dialect acoustic conservation manifold through synchronous calibration of tone energy levels using Riemannian manifolds, Lie group isomorphisms of chaotic attractors, and conformal mapping compensation of resonant cavity fingerprints.

[0068] S2104, based on the dialect acoustic conservation manifold, drives the quantum state reorganization of sound waves, implants the quantum entanglement characteristics of real samples, reconstructs the chaotic synchronization parameters of blast pulses, injects the topological constraints of nasal resonance, and outputs the optimized sound field propagation operator;

[0069] S2105: Import the optimized sound field propagation operator into the acoustic holographic modulation engine to calculate the sound pressure gradient field distribution of dialect feature points, generate destructive interference wavefronts to counter mechanical resonance, synthesize the quantum coherent state sequence of regional intonation, and form a dialect acoustic holographic projection;

[0070] S2106, recursively comparing the dialect acoustic holographic projection with the acoustic invariant set of the real user speech sample through a multi-channel feedback loop, capturing the residual spectral phase deviation and quantifying the formant shift error, and generating an acoustic error correction vector;

[0071] S2107, based on the acoustic error correction vector, reconstructs the glottal pulse wave conduction path, optimizes the vocal cord relaxation oscillation parameters through the laryngeal turbulence field inversion algorithm, and embeds regional oropharyngeal closure inertia characteristics to obtain the dialect bioacoustic optimization operator;

[0072] S2108 uses a dialect bioacoustic optimization operator to perform multiple rounds of iterative convergence calculations, eliminates synthetic harmonic distortion through quantum coherent state interference, reconstructs the quantized energy level distribution of the tone glide trajectory, and generates high-fidelity dialect speech samples;

[0073] S2109, based on high-fidelity dialect speech samples, ultimately optimizes regional accent similarity.

[0074] In this embodiment, the present invention generates a dialect voiceprint quantum map carrying bioacoustic markers by performing multi-scale voiceprint resonance analysis on real user voice samples, thereby solving the problem of incomplete acoustic feature modeling in traditional dialect cloning; by mapping the synthesized speech to the dialect acoustic manifold space to form an acoustic anomaly tensor field, and constructing a spatiotemporal regular grid for adaptive dialect linking, the acoustic feature offset error caused by linking is avoided; by implanting tone quantum entanglement features and reconstructing chaotic synchronization parameters through sound wave quantum state recombination, combining with an acoustic holographic modulation engine to generate destructive interference wavefronts, dynamically compensating for harmonic distortion caused by mechanical resonance; by recursive comparison of a multi-channel feedback loop to generate an acoustic error correction vector, and reconstructing the glottal pulse wave conduction path to implant regional oropharyngeal closure features, reflecting the real physical constraints of dialect bioacoustics; using the dialect bioacoustic optimization operator to perform multiple rounds of iterative convergence calculations to eliminate synthetic harmonic distortion, and finally generating high-fidelity dialect speech samples through quantized energy level distribution recombination, solving the problem of acoustic distortion accumulation in regional accent similarity optimization.

[0075] Driving the audio avatar to conduct multiple rounds of conversational interaction with the user specifically includes:

[0076] S310 activates the industry knowledge graph in interview scenarios, recursively generates a technical question chain based on the depth of the user's answers, and dynamically adjusts the question abstraction level;

[0077] S320: Real-time conversion of speech scripts during speech training, retaining the user's accustomed spoken vocabulary while injecting virtual confusion trigger signals from the virtual audience;

[0078] S330, when a virtual confusion trigger signal is detected, captures the microcurrent changes on the user's face, identifies the confused micro-expression, and generates a simplified response suggestion when the confused micro-expression persists.

[0079] S400 analyzes the interaction process based on a reinforcement learning model, generates a multi-dimensional evaluation report including language redundancy, logical coherence, and micro-expression management indicators, and simultaneously stores the best performance clips in the style sample library.

[0080] The reinforcement learning model uses interaction rounds as its state space and response strategies as its action space. By setting multi-dimensional reward functions such as language quality, emotional expression, and logical rationality, it dynamically optimizes the weight distribution of evaluation indicators. The speech-to-text data processing unit employs a sliding window algorithm to analyze the speech transcription and interaction text sentence by sentence. It uses semantic fingerprinting technology to identify repetitive phrases (such as the frequent occurrence of "that is to say") and calculates a redundancy index using a dictionary of fuzzy qualifiers (including words like "maybe" and "probably"). For example, in a business negotiation simulation, the system can identify patterns in negotiators frequently using ambiguous expressions when elaborating on core terms.

[0081] The logical coherence construction unit uses a graph neural network to construct a dialogue topology graph, parsing the content of each round of responses into argument nodes (such as "market demand analysis") and evidence edges (such as "user survey data support"). Through a traversal algorithm, it detects nodes with missing evidence (such as putting forward an opinion but not providing data support) and broken causal reasoning edges (such as the semantic incoherence of the previous and subsequent sentences connected by "therefore". For example, in an academic defense scenario, it can locate the logical fault in the respondent's discussion of "research innovation points" that lacks support from comparative experimental data.

[0082] The micro-expression management indicator generation unit uses a dynamic time warping algorithm to align the real-time captured facial action unit (AU) sequence with the benchmark expressions in the sample library, quantifying the proportion of micro-expression out-of-control duration (such as the duration of lip biting when nervous). For example, in a stress interview simulation, the system can identify the complex micro-expression out-of-control state of pupil constriction (AU4) and frontalis muscle contraction (AU1) that occurs when a candidate answers sensitive questions.

[0083] The dynamic weighted evaluation unit is based on the policy gradient algorithm of reinforcement learning and automatically adjusts the weight of each indicator according to the user's historical performance. For example, in the novice training stage, the penalty weight for language redundancy is increased, and in the advanced stage, the reward weight for logical coherence is increased. When the weighted total score exceeds a certain proportion of the historical best record, the optimal segment marking mechanism is triggered, and the attention mechanism is used to locate the interaction segments with outstanding performance (such as the period in a round of conversation when the logical topology diagram is complete and the micro-expressions are stable).

[0084] The multi-dimensional assessment report generation unit visually associates the three-dimensional radar chart with improvement suggestions. For example, when the language redundancy dimension in the radar chart is low, it automatically pushes a targeted training plan of "using specific data instead of vague expressions". At the same time, the speech rhythm features (such as the ups and downs of rhythm), text structure (such as the three-part structure of "viewpoint-evidence-conclusion") and facial expression patterns (such as the degree of eyebrow raising when expressing confidence) of the optimal segment are stored in the style sample library to form a reusable personalized expression template.

[0085] The dynamic weighting mechanism driven by reinforcement learning enables the evaluation criteria to adapt to the ability stages of different users. For example, in the training of new employees, the system will focus on language fluency indicators, while for senior managers, the assessment will focus on logical depth and micro-expression control. This personalized evaluation model is more targeted than traditional fixed-weight assessment tools.

[0086] The collaborative analysis of multi-dimensional indicators breaks through the limitations of single-dimensional evaluation. For example, when a user's speech has high language redundancy but good logical coherence, the system can identify the characteristics of "detailed information but insufficient expression efficiency" rather than simply judging it as poor overall performance. This refined evaluation can avoid the misjudgment of "one loss for all".

[0087] The incremental storage of optimal performance fragments forms a personalized experience knowledge base. For example, the "technical difficulty response fragments" accumulated by users in multiple interview simulations can form a style sample library. When encountering similar scenarios again, the system can quickly call historical high-quality models based on transfer learning to assist in generating response suggestions. This experience inheritance mechanism enables the training effect to show a continuous cumulative effect.

[0088] In practical applications, this technology can significantly improve the accuracy and efficiency of skills training. For example, in government spokesperson training, the system analyzes spokespersons' language redundancy indicators (such as the frequency of vague expressions like "relevant departments"), logical coherence topology (such as the completeness of the causal chain when interpreting policies), and micro-expression management indicators (such as changes in blinking frequency when responding to sensitive questions) during simulated reporter questions. It then generates a multi-dimensional report with real-time feedback. It also stores a sample of a spokesperson's outstanding three-part response ("factual statement - emotional resonance - solution") from a crisis response in a sample library for reuse in subsequent training. In psychological counselor training scenarios, the system assesses the duration of a student's uncontrolled micro-expressions during simulated consultations (such as the duration of a frown when facing a client's negative emotions). Combined with the logical coherence of the conversation (such as the consistency of the counseling strategy), it generates personalized improvement plans. It also identifies the prosodic characteristics of outstanding students' empathetic expressions (such as a gentle tone and appropriate pauses) as optimal learning references.

[0089] This evaluation mechanism, which integrates reinforcement learning and multimodal analysis, breaks through the limitations of traditional subjective scoring and provides users with a scientific and personalized path to improve their abilities. It is particularly suitable for training in professional fields that require high accuracy of expression and emotional control.

[0090] In this embodiment, the industry knowledge graph is activated in the interview scenario, a technical question chain is recursively generated based on the depth of the user's answer, and the question abstraction level is dynamically adjusted. Specifically, the following steps are involved:

[0091] S3101: Extract technical entities and action predicates from user responses, annotate the functional roles of entities in the semantic network through dependency syntax analysis, and generate a technical entity-action mapping table with weighted labels;

[0092] S3102: Based on the weighted technical entity-action mapping table, detect the role transition of the same technical entity in adjacent responses, record the direction of change in the concept abstraction level, and construct the drift vector of the user's cognitive path;

[0093] S3103: After activating the industry knowledge graph, perform subgraph traversal along the drift vector direction of the user's cognitive path. If the vector points to the practice layer, prune the theoretical branch; if it points to the abstract layer, prune the engineering detail branch to generate a scenario-adapted condensed knowledge subgraph.

[0094] S3104, combining real-time speech pauses and facial micro-expressions, marking nodes that conflict with physiological signals in the scene-adapted condensed knowledge subgraph to generate an interference-resistant knowledge node set;

[0095] S3105: Assemble the anti-interference knowledge node set into a tree structure according to logical dependencies. The root node is the last technical point mentioned by the user. The child nodes are expanded according to the three-level abstraction of "principle-implementation-application". Each node generates an industry standard question template to obtain a tree-shaped question chain prototype.

[0096] S3106: Dynamically compress the levels of the tree-shaped question chain prototype based on historical answer quality. If the user completes the application layer for two consecutive rounds, delete the principle layer nodes. If semantic repetition occurs at the implementation layer, collapse redundant nodes at the same level to generate an optimized question sequence.

[0097] S3107 acquires real-time speech rate fluctuation and palm sweat data, inserting buffer questions into the optimized follow-up question sequence: When a high-pressure signal is detected, low-abstraction verification questions are inserted between consecutive technical questions to generate a stress-resistant follow-up question chain;

[0098] S3108 loads the stress-resistant follow-up chain into the dialogue engine, and deeply regulates subsequent questions based on the real-time response semantics. When upgrading, it jumps from "how to achieve it" to "why it is better than the traditional solution", and when downgrading, it inserts "please give an example to illustrate the basic process" to achieve closed-loop regulation and ultimately dynamically adjust the abstract level of the problem.

[0099] In this embodiment, the present invention generates a weighted technical entity-action mapping table by extracting technical entities and action predicates from user responses, thereby solving the problem of the lack of semantic association in traditional question chains. The invention also prunes knowledge graph branches along the direction of the user's cognitive path drift vector (pruning theoretical branches at the practice layer or pruning engineering details at the abstract layer) and generates an interference-resistant knowledge set by combining physiological signal-labeled conflicting nodes to avoid question loss caused by user cognitive jumps. The invention assembles a tree-shaped question chain prototype according to the three-level abstraction of "principle-implementation-application" and dynamically compresses the levels (deleting principle-level nodes or collapsing redundant nodes) based on historical response quality to correct the mismatch between the question abstraction level and user capabilities. The invention also inserts buffer questions into the question sequence using real-time speech rate and sweat data (high-pressure signals trigger low-abstraction verification questions) and combines the response semantics to deeply close the loop to regulate the question's order of advancement or reduction, thereby solving the problem of insufficient adaptability of the question chain in high-pressure scenarios.

[0100] The generation of a multi-dimensional evaluation report including language redundancy, logical coherence, and micro-expression management indicators specifically includes:

[0101] S410, obtaining text data converted from the voice stream, and calculating language redundancy indicators based on the interactive text, and counting the proportion of repetitive phrases and the frequency of occurrence of vague qualifiers;

[0102] S420: Record the dynamic responses of multiple rounds of dialogue interactions, construct a logical coherence topology and logical coherence indicators, and mark nodes with missing argument evidence and broken edges in causal reasoning.

[0103] S430, comparing facial video stream data with the content in the sample library, generating micro-expression management indicators, and quantifying the proportion of time when micro-expressions are out of control;

[0104] S440, dynamically weighting the redundancy index, the coherence index, and the micro-expression index. When the weighted total score exceeds the interval of the historical best record, marking the current interaction segment as the best performance segment;

[0105] S450 integrates redundancy indicators, coherence indicators, micro-expression indicators and optimal performance segments to form a multi-dimensional evaluation report that includes a three-dimensional radar chart and improvement suggestions.

[0106] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0107] In this embodiment, the text data converted from the voice stream is obtained, and the language redundancy index is calculated in combination with the interactive text, and the proportion of repetitive phrases and the frequency of occurrence of vague qualifiers are counted, specifically including:

[0108] S4101: Acquire speech-to-text data, convert the acoustic features detected in the speech-to-text data into visual text markers, insert a spacer when a pause timeout is detected, add a conclusion marker where a pitch drop exceeds a preset frequency, and mark an acceleration segment above the speaking speed range when the speaking speed increases, thereby generating enhanced text data;

[0109] S4102 performs double filtering based on the enhanced text data. First, it scans for repeated subject-predicate structures within consecutive sentences. If the repeated region contains a separator, it is retained as effective emphasis. Then, it detects the repetition of synonyms within the same paragraph. If rapid repetition occurs within an accelerated segment, it is marked as tense redundancy, and a list of repeated phrases is obtained.

[0110] S4103: Acquire spatiotemporal features, locate fuzzy qualifiers in the enhanced text data, and label severity levels based on the spatiotemporal features. Fuzzy words following the conclusion marker are marked as red (high risk), repeated at adjacent intervals are marked as yellow (habitual), and the rest are marked as blue (normal), thus forming a fuzzy word distribution map with three-dimensional spatiotemporal coordinates.

[0111] S4104, semantically purify the list of repeated phrases, exempt repetitions in figurative rhetorical structures, do not count repetitions of professional terms, retain only phrases that are purely redundant, and delete items that are justified by acoustic markers to obtain a purified core set of redundant phrases;

[0112] S4105: Compare the fuzzy word distribution map of the three-dimensional space-time coordinates with the enhanced text data. If a red high-risk fuzzy word appears before a separator, it is upgraded to a double contradiction point. If a yellow habitual fuzzy word increases with the frequency of pitch, it is marked as a pseudo-deterministic expression, and a list of high-risk contradiction points is generated.

[0113] S4106: Determine and obtain the scenario knowledge base, repeatedly exempt the technical terms in the purified core redundant phrase set in the interview scenario, and apply weight addition to the items in the high-risk contradiction list in the speech scenario to generate the final scenario-adapted data set;

[0114] S4107 compares the user's historical high-quality interaction clips to obtain the user's typical expression style. When repeated patterns in the purified core redundant phrase set match the user's typical expression style, they are added to the personalized whitelist. Furthermore, the user's frequently used cautious and ambiguous words in the high-risk conflict point list are marked as reasonable expressions to generate a personalized pardon whitelist.

[0115] S4108: Integrate the final scenario-adapted dataset with the personalized pardon whitelist. Count the percentage of repeated phrases, excluding whitelist entries. Calculate the frequency of fuzzy words, focusing on high-risk contradictions. This yields the percentage of repeated phrases and the frequency of fuzzy qualifiers that contain fraud warnings.

[0116] S4109, based on the proportion of repetitive phrases containing fraud warnings and the frequency of occurrence of vague qualifiers, the proportion of repetitive phrases and the frequency of occurrence of vague qualifiers are finally obtained by statistics.

[0117] In this embodiment, the present invention generates enhanced text data by converting speech acoustic features into visual text markers (inserting spacers for pauses and timeouts, or adding conclusion markers for sudden drops in pitch, or marking accelerated segments with faster speech speeds). This solves the problem of traditional text analysis ignoring acoustic semantics. Based on a double filtering mechanism, the present invention scans subject-verb repetition structures (including spacers, which are retained for emphasis) and synonym repetition (accelerated segments are marked as tense and redundant), and combines three-dimensional spatiotemporal coordinates to mark the severity level of fuzzy words (marking them red for high risk after the conclusion marker or yellow for habitual recurrence) to avoid misjudgment in redundancy statistics. It generates a scenario-adapted dataset (exempting technical terms or weighted high-risk points in interviews) by semantically purifying repeated phrases (exempting metaphors or professional terms) and identifying fuzzy word contradictions (upgrading high-risk words before spacers to double contradictions or marking them with false certainty when accompanied by pitch increases). It also establishes a personalized whitelist by comparing historical high-quality user clips (matching repetitive patterns of classic styles into the whitelist or using habitual cautious fuzzy words for reasonable expressions). When counting the proportion of repeated phrases, it deducts whitelist entries and focuses on high-risk contradictions, thus solving the problem of insufficient scenario and individual adaptability in language redundancy assessment.

[0118] In this embodiment, the comparison of facial video stream data with the content in the sample library to generate micro-expression management indicators and quantify the proportion of time duration of uncontrolled micro-expressions specifically includes:

[0119] S4301: Extracting standard values ​​of blink frequency, nasolabial fold displacement stability distance, and zygomatic muscle relaxation strength of the user in a calm state from a sample library to form a balance interval between blinking and nasolabial fold, and a coordination interval between zygomatic muscle and blinking as reference benchmarks, thereby obtaining a reference balance band.

[0120] S4302: Obtain current facial video stream data and compare the actual blink frequency, actual nasolabial fold displacement, and real-time zygomatic muscle strength of the current facial video stream data with the reference balance band. A red signal is recorded when the blink frequency exceeds the standard range, a yellow signal is recorded when the nasolabial fold displacement exceeds the stable distance, and a blue signal is recorded when the zygomatic muscle strength deviates from the coordination range. A three-color signal record is obtained based on the red, yellow, and blue signals.

[0121] S4303, when the red and yellow signals are present simultaneously but the blue signal is absent in the three-color signal recording, it is determined that the nasolabial groove is out of control and the zygomatic muscle is not responding; when the blue and red signals are present simultaneously but the yellow signal is absent, it is determined that the zygomatic muscle is out of control and blinking is not inhibited. The time period of the above states is continuously recorded to obtain the period of compensatory loss of control;

[0122] S4304: During the period of uncontrolled compensation, detect whether blink frequency remains constant when the nasolabial groove displacement increases continuously and rapidly, or whether blink frequency recovers slowly after a sudden decrease in zygomatic muscle strength. Mark the specific time points when such movements are lost to obtain a neural disconnection sequence.

[0123] S4305: Taking each node in the neural breakpoint sequence as the center, trace back within a preset time period and extend back within a preset time period. Count the number of nasolabial groove displacement direction reversals, the difference in zygomatic muscle strength fluctuations, and the amplitude of blink frequency fluctuations within the window during the sum of the traced and extended periods to obtain the three elements of oscillation intensity.

[0124] S4306: When the three elements of oscillation intensity show frequent reversals of the nasolabial groove direction and violent fluctuations of the zygomatic muscle, it is marked as severe loss of control; when the amplitude of blink fluctuations significantly exceeds the standard, it is marked as widespread loss of control; other cases are marked as mild loss of control, and a graded loss of control label is obtained;

[0125] S4307: Accumulate all graded loss of control label periods, multiply severe loss of control by 1.5 times the duration, extensive loss of control by 2 times the duration, and mild loss of control by the original duration. Simultaneously filter out pseudo-loss of control segments with stable zygomatic muscles and normal blinking to obtain the total effective loss of control duration.

[0126] S4308: Divide the total effective out-of-control duration by the duration of the entire interaction process to obtain the basic proportion and obtain the user's best historical performance. When the out-of-control proportion in the user's best historical performance is low, the proportion of micro-expression out-of-control duration is dynamically adjusted according to the out-of-control proportion interval to finally quantify the proportion of micro-expression out-of-control duration.

[0127] In this embodiment, the present invention compares the measured values ​​of blink frequency, nasolabial fold displacement, and zygomatic muscle strength with a reference balance belt to generate a three-color signal record (red for exceeding the standard range, yellow for breaking the stability distance, and blue for deviation from the coordination range). This solves the problem of inaccurate micro-expression assessment using a single indicator. The present invention also analyzes the characteristics of motor disconnection during periods of uncontrolled compensation (blinking stagnation during a surge in nasolabial fold displacement or delayed blink recovery after a sudden drop in the zygomatic muscle) to mark neural breakpoint sequences, thus avoiding misjudgment due to muscle compensation effects. The present invention also expands the time window around the neural breakpoint to calculate the three elements of oscillation intensity (number of nasolabial fold reversals, zygomatic muscle fluctuation difference, or blink amplitude). The degree of uncontrolled micro-expression is then graded (frequent reversals plus severe fluctuations, or extensive blinking beyond the standard), correcting the coarse-grained nature of traditional duration statistics. The present invention also accumulates the graded uncontrolled micro-expression duration (severe × 1.5 or extensive × 2.0), filters out pseudo-uncontrolled micro-expression segments (stable zygomatic muscle plus normal blinking), and dynamically adjusts the basic percentage based on the user's historical best performance to accurately quantify the true physiological significance of the duration of uncontrolled micro-expression.

[0128] Figure 4 The structural block diagram of the intelligent dialogue system based on the AI ​​multimodal large model provided by the embodiment of the present invention is as follows: Figure 4 As shown, the system includes:

[0129] The biometric data collection module 100 is used to collect the user's biometric data in real time, including voice stream, facial video stream and interactive text, extract speech prosodic features, language structure features and facial dynamic features through a multimodal fusion encoder, and output a joint feature vector;

[0130] Avatar model construction module 200 is used to construct a user digital avatar model based on the joint feature vector, clone the user's personalized expression pattern using a generative adversarial network, and generate an audio avatar that includes mappings of language rhythm, idiomatic vocabulary, and regional accents;

[0131] The conversation interaction module 300 is used to load a preset scenario knowledge base, drive the audio avatar to conduct multiple rounds of conversational interaction with the user, and capture the user's real-time physiological signals through the expression recognition module, and generate dynamic response content suggestions based on the conversation context;

[0132] The interaction process analysis module 400 is used to analyze the interaction process based on the reinforcement learning model, generate a multi-dimensional evaluation report including language redundancy, logical coherence, and micro-expression management indicators, and simultaneously store the best performance fragments in the style sample library.

[0133] like Figure 5 As shown, the biometric data collection module 100 specifically includes:

[0134] The microphone capture unit 110 is used to capture the speech stream using a directional beamforming microphone array, extract the fundamental frequency trajectory and formant features through linear predictive coding, and construct a phoneme model containing dialect-specific acoustic markers;

[0135] A facial key point detection unit 120 is used to deploy a 3D facial key point detection algorithm to quantify the eye blink frequency, nasolabial groove displacement vector, and zygomatic muscle contraction strength dynamic sequence from the facial video stream;

[0136] The interactive text analysis unit 130 is used to parse the interactive text based on the dependency syntax analyzer and mark the distribution density of rhetorical questions and sentiment-oriented connectives;

[0137] The cross-modal encoder unit 140 is configured to input the above-mentioned voice stream, facial video stream and interactive text into a cross-modal Transformer encoder to generate a dimensionally aligned joint feature vector.

[0138] like Figure 6 As shown, the avatar model construction module 200 specifically includes:

[0139] The conditional adversarial generative network construction unit 210 is used to construct a conditional adversarial generative network, wherein:

[0140] The generator receives the joint feature vector and scene label and outputs a dialect text sequence with prosody tags;

[0141] The discriminator compares real user speech samples with generated samples and optimizes regional accent similarity using the spectrogram dynamic time warping algorithm;

[0142] The personalized speech rule base establishing unit 220 is used to establish a personalized speech rule base to store the user's unique filler word usage frequency, question ending tone rising pattern, and speech speed fluctuation curve under tension.

[0143] like Figure 7 As shown, the dialogue interaction module 300 specifically includes:

[0144] The industry knowledge graph activation unit 310 is used to activate the industry knowledge graph in an interview scenario, recursively generate a technical question chain based on the depth of the user's answer, and dynamically adjust the question abstraction level;

[0145] The speech text real-time conversion unit 320 is used to convert the speech text in real time during speech training, retaining the user's accustomed spoken vocabulary while injecting virtual confusion trigger signals of the virtual audience;

[0146] The virtual confusion signal monitoring unit 330 is used to capture the changes in microcurrent on the user's face and identify confused micro-expressions when a virtual confusion trigger signal is detected, and generate a simplified response suggestion when the confused micro-expression lasts for more than 2 seconds.

[0147] like Figure 8 As shown, the interaction process analysis module 400 specifically includes:

[0148] The speech-to-text data processing unit 410 is used to obtain text data converted from the speech stream, and calculate the language redundancy index in combination with the interactive text, and count the proportion of repetitive phrases and the frequency of occurrence of vague qualifiers;

[0149] Logical coherence construction unit 420 is used to record the dynamic response content of multiple rounds of dialogue interaction, construct a logical coherence topology map and logical coherence indicators, and mark the nodes where argument evidence is missing and the edges where causal reasoning is broken;

[0150] A micro-expression management index generating unit 430 is used to compare facial video stream data with the content in the sample library to generate a micro-expression management index to quantify the proportion of time when micro-expressions are out of control;

[0151] Dynamic weighted evaluation unit 440 is used to dynamically weight the redundancy index, coherence index, and micro-expression index. When the total weighted score exceeds 90% of the historical best record, the current interaction segment is marked as the best performance segment;

[0152] The multi-dimensional evaluation report generating unit 450 is used to integrate the redundancy index, the coherence index, the micro-expression index and the best performance segment to form a multi-dimensional evaluation report including a three-dimensional radar chart and improvement suggestions.

[0153] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.

[0154] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0155] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. An intelligent dialogue method based on AI multimodal large model, characterized by: The method comprises: Collect user biometric data in real time, including voice stream, facial video stream and interactive text, extract voice prosodic features, language structure features and facial dynamic features through a multimodal fusion encoder, and output a joint feature vector; A user digital avatar model is constructed based on the joint feature vector, and a generative adversarial network is used to clone the user's personalized expression pattern, generating an audio avatar that includes mappings of language rhythm, idiomatic vocabulary, and regional accents. Loading a preset scenario knowledge base, driving the audio avatar to conduct multiple rounds of conversational interactions with the user, and using the expression recognition module to capture the user's real-time physiological signals, generating dynamic response content suggestions based on the conversation context; Analyze the interaction process based on the reinforcement learning model, generate a multi-dimensional evaluation report including language redundancy, logical coherence, and micro-expression management indicators, and simultaneously store the best performance clips in the style sample library; Generating an audio avatar containing mappings of language rhythm, idiomatic vocabulary, and regional accents specifically includes: Construct a conditional adversarial generation network, where: The generator receives the joint feature vector and scene label and outputs a dialect text sequence with prosody tags; The discriminator compares real user speech samples with generated samples and optimizes regional accent similarity using the spectrogram dynamic time warping algorithm; Build a personalized speech rule library to store user-specific filler word usage frequency, rising tone pattern at the end of question sentences, and speech speed fluctuation curve under stress; The discriminator compares the real user speech samples with the generated samples and optimizes the regional accent similarity through the spectrogram dynamic time warping algorithm, specifically including: Perform multi-scale voiceprint resonance analysis on real user speech samples, extract a set of dialect-specific acoustic invariants to generate a dialect voiceprint quantum map. The dialect voiceprint quantum map includes the quantized energy level distribution of tone glide trajectories, the chaotic attractor characteristics of plosive oral air pressure pulses, and the resonant cavity topology fingerprint of nasalized vowels. Acquiring a synthesized speech sample, mapping the synthesized speech sample to a dialect acoustic manifold space, performing variational encoding based on the dialect voiceprint quantum map, identifying spectral feature offset regions and quantifying the degree of acoustic distortion, thereby forming an acoustic anomaly tensor field containing distortion hotspots; A spatiotemporal regular grid for adaptive dialect linking is constructed within an acoustic anomaly tensor field containing distortion hotspots. Through Riemannian manifold synchronization calibration of tone energy levels, Lie group isomorphism transformation of chaotic attractors, and conformal mapping compensation of resonant cavity fingerprints, a dialect acoustic conservation manifold is generated. Based on the dialect acoustic conservation manifold, the acoustic wave quantum state is driven to reorganize, the tone quantum entanglement characteristics of real samples are implanted, the chaotic synchronization parameters of the blast sound pulse are reconstructed, the topological constraints of nasal resonance are injected, and the optimized sound field propagation operator is output; The optimized sound field propagation operator is imported into the acoustic holographic modulation engine to calculate the sound pressure gradient field distribution of dialect feature points, generate destructive interference wavefronts to counter mechanical resonance, synthesize the quantum coherent state sequence of regional intonation, and form a dialect acoustic holographic projection; The dialect acoustic holographic projection is recursively compared with the acoustic invariant set of the real user speech sample through a multi-channel feedback loop to capture the residual spectral phase deviation and quantify the formant shift error to generate an acoustic error correction vector. The glottal pulse wave conduction path is reconstructed based on the acoustic error correction vector. The vocal cord relaxation oscillation parameters are optimized through the laryngeal turbulence field inversion algorithm. The regional oropharyngeal closure inertia characteristics are implanted to obtain the dialect bioacoustic optimization operator. Using the dialect bioacoustic optimization operator to perform multiple rounds of iterative convergence calculations, quantum coherent state interference is used to eliminate synthetic harmonic distortion, reconstruct the quantized energy level distribution of the tone glide trajectory, and generate high-fidelity dialect speech samples; Based on high-fidelity dialect speech samples, the ultimate goal is to optimize the similarity of regional accents.

2. The intelligent dialogue method based on AI multimodal large model according to claim 1 is characterized in that: The real-time collection of the user's biometric data specifically includes: A directional beamforming microphone array is used to capture the speech stream, and linear predictive coding is used to extract the fundamental frequency trajectory and formant features to construct a phoneme model that includes dialect-specific acoustic markers. Deploy a 3D facial landmark detection algorithm to quantify the eye blink frequency, nasolabial fold displacement vector, and zygomatic muscle contraction intensity dynamic sequence from facial video streams; Parse interactive texts based on a dependency parser, annotating the distribution density of interrogative sentences and sentiment-based connectives; The above-mentioned voice stream, facial video stream and interaction text are input into the cross-modal Transformer encoder to generate a dimensionally aligned joint feature vector.

3. The intelligent dialogue method based on AI multimodal large model according to claim 1 is characterized in that: Driving the audio avatar to conduct multiple rounds of conversational interaction with the user specifically includes: Activate the industry knowledge graph in interview scenarios, recursively generate technical question chains based on the depth of user responses, and dynamically adjust the level of question abstraction; During speech training, the speech text is converted in real time, retaining the user's accustomed spoken vocabulary while injecting virtual confusion trigger signals from the virtual audience; When a virtual confusion trigger signal is detected, the system captures the changes in microcurrent on the user's face, identifies the confused micro-expression, and generates simplified response suggestions when the confused micro-expression persists.

4. The intelligent dialogue method based on AI multimodal large model according to claim 3 is characterized in that: The above mentioned method activates the industry knowledge graph in the interview scenario, recursively generates a technical question chain based on the depth of the user's answer, and dynamically adjusts the question abstraction level, specifically including: Extract technical entities and action predicates from user responses, annotate the functional roles of entities in the semantic network through dependency syntax analysis, and generate a technical entity-action mapping table with weighted labels; Based on a weighted technical entity-action mapping table, we detect the role transition of the same technical entity in adjacent responses, record the direction of change in the conceptual abstraction level, and construct the drift vector of the user's cognitive path. After activating the industry knowledge graph, the subgraph is traversed along the drift vector direction of the user's cognitive path. If the vector points to the practice layer, the theoretical branch is pruned; if it points to the abstract layer, the engineering detail branch is pruned to generate a condensed knowledge subgraph adapted to the scenario. By combining real-time speech pauses and facial micro-expressions, we mark nodes that conflict with physiological signals in a scene-adapted condensed knowledge subgraph, generating an interference-resistant knowledge node set. The anti-interference knowledge node set is assembled into a tree structure according to logical dependencies. The root node is the last technical point mentioned by the user. The child nodes are expanded according to the three-level abstraction of "principle-implementation-application". Each node generates an industry-standard question template to obtain a tree-shaped question chain prototype. Based on the quality of historical responses, the tree-shaped question chain prototype is dynamically compressed. If the user completes the application layer for two consecutive rounds, the principle layer node is deleted. If semantic repetition occurs at the implementation layer, redundant nodes at the same level are collapsed to generate an optimized question sequence. Real-time speech rate fluctuation and palm sweat data are obtained, and buffer questions are inserted into the optimized question sequence: when high-pressure signals are detected, low-abstraction verification questions are inserted between consecutive technical questions to generate a stress-resistant question chain; The stress-resistant follow-up question chain is loaded into the dialogue engine, and subsequent questions are regulated according to the real-time response semantic depth, ultimately obtaining a dynamically adjusted question abstraction level.

5. The intelligent dialogue method based on AI multimodal large model according to claim 4 is characterized in that: The generation of a multi-dimensional evaluation report including language redundancy, logical coherence, and micro-expression management indicators specifically includes: Obtain text data converted from voice streams, and combine it with interactive text to calculate language redundancy indicators, including the proportion of repetitive phrases and the frequency of vague qualifiers. Record the dynamic response content of multiple rounds of dialogue interactions, construct a logical coherence topology diagram and logical coherence indicators, and mark the nodes where argument evidence is missing and the edges where causal reasoning is broken; Compare facial video stream data with the content in the sample library to generate micro-expression management indicators and quantify the proportion of time spent in micro-expressions that are out of control; Dynamically weight the redundancy index, coherence index, and micro-expression index. When the weighted total score exceeds the interval of the historical best record, the current interaction segment is marked as the best performance segment. Integrate redundancy indicators, coherence indicators, micro-expression indicators and optimal performance segments to form a multi-dimensional evaluation report that includes a three-dimensional radar chart and improvement suggestions.

6. The intelligent dialogue method based on AI multimodal large model according to claim 5 is characterized in that: The method of obtaining text data converted from the voice stream and calculating the language redundancy index in combination with the interactive text, and counting the proportion of repetitive phrases and the frequency of occurrence of vague qualifiers, specifically includes: Acquire speech-to-text data, convert the acoustic features detected in the speech-to-text data into visual text markers, insert a spacer when a pause timeout is detected, add a conclusion marker when the pitch drops beyond a preset frequency, and mark the acceleration segment above the speaking speed range when the speaking speed increases, thereby generating enhanced text data; Double filtering is performed based on enhanced text data. First, repeated subject-predicate structures within consecutive sentences are scanned. If the repeated region contains separators, they are retained as effective emphasis. Then, the repetition of synonyms in the same paragraph is detected. When rapid repetition occurs in the accelerated segment, it is marked as tense redundancy, and a list of repeated phrases is obtained. Obtain spatiotemporal features, locate fuzzy qualifiers in enhanced text data, and label severity levels based on spatiotemporal features. Fuzzy words following the conclusion marker are marked as red, high-risk; repeated occurrences at adjacent intervals are marked as yellow, habitual, and the remainder are marked as blue, forming a fuzzy word distribution map with three-dimensional spatiotemporal coordinates. The list of repeated phrases was semantically purified, with repetitions in figurative rhetoric structures exempted and repetitions of professional terms not counted. Only phrases that were purely redundant were retained, while items that were justified by acoustic markers were deleted, resulting in a purified core set of redundant phrases. Comparing the fuzzy word distribution map of the three-dimensional space-time coordinates with the enhanced text data, red high-risk fuzzy words are upgraded to double contradiction points if they appear before the separator, and yellow habitual fuzzy words are marked as pseudo-deterministic expressions when accompanied by an increase in pitch frequency, thus generating a list of high-risk contradiction points; Determine and obtain the scenario knowledge base. In the interview scenario, the technical terms in the purified core redundant phrase set are repeatedly exempted. In the speech scenario, weights are added to the items in the high-risk contradiction list to generate the final scenario-adapted dataset. Compare the user's historical high-quality interaction clips to obtain the user's typical expression style. When the repeated patterns in the purified core redundant phrase set match the user's typical expression style, they are added to the personalized whitelist. The cautious and ambiguous words commonly used by the user in the high-risk conflict point list are marked as reasonable expressions to generate a personalized pardon whitelist. Integrate the final scenario-adapted dataset with the personalized pardon whitelist. Count the percentage of repeated phrases, deduct the whitelist entries. Calculate the frequency of fuzzy words, focusing on high-risk contradictions. This yields the percentage of repeated phrases and the frequency of fuzzy qualifiers that contain fraud warnings. Based on the proportion of repetitive phrases and the frequency of occurrence of vague qualifiers containing fraud warnings, the proportion of repetitive phrases and the frequency of occurrence of vague qualifiers are finally obtained through statistics.

7. The intelligent dialogue method based on AI multimodal large model according to claim 6 is characterized in that: The comparison of facial video stream data with the content in the sample library to generate micro-expression management indicators and quantify the proportion of micro-expression out-of-control duration specifically includes: Extract the standard values ​​of eye blink frequency, nasolabial fold displacement stability distance, and zygomatic muscle relaxation strength of the user in a calm state from the sample library to form the balance interval between blinking and nasolabial fold, and the coordination interval between zygomatic muscle and blinking as reference benchmarks to obtain the reference balance belt; Obtain the current facial video stream data, and compare the actual blink frequency, actual nasolabial fold displacement distance, and real-time zygomatic muscle strength value of the current facial video stream data with the reference balance belt; record a red signal when the blink frequency exceeds the standard range, a yellow signal when the nasolabial fold displacement exceeds the stable distance, and a blue signal when the zygomatic muscle strength deviates from the coordination range; based on the red, yellow, and blue signals, a three-color signal record is obtained; When the red and yellow signals are present simultaneously but the blue signal is absent in the three-color signal recording, it is determined that the nasolabial groove is out of control and the zygomatic muscle is not responding; when the blue and red signals are present simultaneously but the yellow signal is absent, it is determined that the zygomatic muscle is out of control and blinking is not inhibited. The time period of the above state is continuously recorded to obtain the period of compensatory loss of control; During periods of uncontrolled compensation, we examined whether blink frequency remained constant when nasolabial fold displacement increased rapidly and continuously, or whether blink frequency recovered after a sudden decrease in zygomatic muscle strength. We then marked the specific time points when such movements were lost to obtain a neural disconnection sequence. Taking each node in the neural breakpoint sequence as the center, trace forward within a preset time period and extend backward within a preset time period. Count the number of nasolabial groove displacement direction reversals, the difference in zygomatic muscle strength fluctuations, and the amplitude of blink frequency fluctuations within the window in the sum of the traceback period and the backward extension period to obtain the three elements of oscillation intensity. When the three elements of oscillation intensity show frequent reversals of the nasolabial groove direction and violent fluctuations of the zygomatic muscle, it is marked as severe loss of control; when the amplitude of blink fluctuations significantly exceeds the standard, it is marked as widespread loss of control; other cases are marked as mild loss of control, and a graded loss of control label is obtained; All the time periods of graded loss of control labels were accumulated, with severe loss of control multiplied by 1.5 times the duration, extensive loss of control multiplied by 2 times the duration, and mild loss of control kept at the original duration. Pseudo-loss of control segments with stable zygomatic muscles and normal blinking were simultaneously filtered out to obtain the total effective loss of control duration. Divide the total effective out-of-control duration by the duration of the entire interaction process to obtain the basic proportion and obtain the user's best historical performance. When the out-of-control proportion in the user's best historical performance is low, dynamically adjust the proportion of micro-expression out-of-control duration according to the out-of-control proportion range to ultimately quantify the proportion of micro-expression out-of-control duration.

8. An intelligent dialogue system based on AI multimodal large model, characterized by: The system applies the intelligent dialogue method based on the AI ​​multimodal large model as described in any one of claims 1 to 7 above, and the system includes: The biometric data acquisition module is used to collect the user's biometric data in real time, including voice stream, facial video stream and interactive text, and extract the speech prosody features, language structure features and facial dynamic features through a multimodal fusion encoder to output a joint feature vector; The avatar model construction module is used to build a user digital avatar model based on the joint feature vector, using a generative adversarial network to clone the user's personalized expression pattern and generate an audio avatar that includes mappings of language rhythm, idiomatic vocabulary, and regional accents; The conversation interaction module is used to load the preset scenario knowledge base, drive the audio avatar to conduct multiple rounds of conversational interaction with the user, and capture the user's real-time physiological signals through the expression recognition module, and generate dynamic response content suggestions based on the conversation context; The interaction process analysis module is used to analyze the interaction process based on the reinforcement learning model, generate a multi-dimensional evaluation report including language redundancy, logical coherence, and micro-expression management indicators, and simultaneously store the best performance clips in the style sample library.

Citation Information

Patent Citations

  • Intelligent multi-mode virtual digital human interaction system based on AI language large model, interaction method and application

    CN120259499A