Agent role switching method, system and related components based on multimodal perception

Through the multimodal perception module and intelligent role switching decision-making mechanism, the problems of cumbersome role switching and insufficient context perception in existing children's educational equipment are solved, the convenience of intelligent agent role switching and the processing of multi-role concurrent scenarios are realized, and the interactive experience is improved.

CN120197139BActive Publication Date: 2025-09-26SHENZHEN BOYUE DOMESTIC GOODS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510685686.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-26
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

Existing children's educational devices have cumbersome role switching operations, lack of context awareness, inability to handle multiple role concurrent scenarios, and single role behavior patterns, resulting in a poor interactive experience.

Method used

User information is collected through the multimodal perception module, candidate roles are determined using the intelligent role switching decision mechanism, and the target role is selected or synthesized through the multi-role conflict resolution mechanism, and role switching is achieved in combination with the preset role model library.

Benefits of technology

It improves the convenience of role switching and context awareness, can handle multi-role concurrent scenarios, and enhances the user experience of the intelligent agent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197139B_ABST
    Figure CN120197139B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, and related components for intelligent agent role switching based on multimodal perception. The method comprises: collecting a user's multimodal perception information through a multimodal perception module; determining a candidate role based on the multimodal perception information through an intelligent role switching decision mechanism; when there is only one candidate role, setting the candidate role as the target role; when there are multiple candidate roles, selecting one of the candidate roles as the target role or combining the candidate roles into the target role through a multi-role conflict resolution mechanism; and switching the current role of the intelligent agent to the target role based on a preset role model library. The present invention can improve the role switching effect of the intelligent agent and enhance the user experience of the intelligent agent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, system, and related components for intelligent agent role switching based on multimodal perception. Background Art

[0002] With the development of artificial intelligence technology, the application of intelligent educational devices in early childhood education is becoming increasingly widespread. Existing children's educational devices usually adopt a single-role or multi-role interaction mode. However, in actual application, existing children's educational devices still have certain defects, such as:

[0003] (1) The role switching operation is cumbersome. Existing devices usually require users to manually switch roles, which usually takes an average of 3-5 steps. This will seriously affect children's learning experience and interest maintenance. (2) Lack of context perception ability. Existing devices cannot automatically switch to the appropriate role according to the conversation context, resulting in a fragmented interactive experience. (3) Unable to handle multi-role concurrent scenarios. When children want to interact with multiple characters at the same time, existing devices cannot effectively handle role response conflicts, often resulting in a confusing interactive experience. (4) The role behavior pattern is single. The characters in existing devices usually adopt preset fixed behavior patterns, lack personalization and adaptability, and cannot adjust their performance according to the characteristics of different children.

[0004] Therefore, how to overcome the above-mentioned defects of the prior art and improve the role interaction effect of smart devices is a problem that those skilled in the art need to solve. Summary of the Invention

[0005] The embodiments of the present invention provide a method, system, intelligent terminal and storage medium for intelligent agent role switching based on multimodal perception, aiming to improve the role switching effect of the intelligent agent and enhance the user experience of the intelligent agent.

[0006] In a first aspect, an embodiment of the present invention provides an agent role switching method based on multimodal perception, comprising:

[0007] Collecting the user's multimodal perception information through the multimodal perception module;

[0008] Based on the multimodal perception information, determining candidate roles through an intelligent role switching decision mechanism;

[0009] When the number of the candidate roles is one, setting the candidate role as the target role;

[0010] When there are multiple candidate roles, one of the candidate roles is selected as the target role through a multi-role conflict resolution mechanism, or the candidate roles are combined into the target role;

[0011] Switch the agent's current role to the target role based on the preset role model library.

[0012] In a second aspect, an embodiment of the present invention provides an agent role switching system based on multimodal perception, comprising:

[0013] An information collection unit, configured to collect the user's multimodal perception information through a multimodal perception module;

[0014] a role determination unit, configured to determine a candidate role through an intelligent role switching decision mechanism based on the multimodal perception information;

[0015] a first setting unit, configured to set the candidate role as a target role when the number of the candidate role is one;

[0016] A second setting unit is configured to select one of the candidate roles as a target role through a multi-role conflict resolution mechanism when there are multiple candidate roles, or to combine the candidate roles into a target role;

[0017] The role switching unit is used to switch the current role of the agent to the target role based on the preset role model library.

[0018] In a third aspect, an embodiment of the present invention provides an intelligent terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for intelligent agent role switching based on multimodal perception as described in the first aspect is implemented.

[0019] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the agent role switching method based on multimodal perception as described in the first aspect.

[0020] The embodiment of the present invention provides a method, system, intelligent terminal and storage medium for intelligent agent role switching based on multimodal perception. The embodiment of the present invention collects the user's multimodal perception information through a multimodal perception module, and then uses an intelligent role switching decision mechanism to determine the candidate role. If there is only one candidate role, the candidate role is set as the target role; if there are multiple candidate roles, a multi-role conflict resolution mechanism is used to select one as the target role, or multiple candidate roles are combined into one target role. Finally, the role of the intelligent agent is switched to the target role according to a preset role model library. In this way, not only can the problems of cumbersome user operations and lack of contextual awareness in the role switching process be overcome, but also the multi-role concurrent scenario can be handled, thereby improving the role switching effect of the intelligent agent and enhancing the user experience of the intelligent agent. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0022] Figure 1 A schematic diagram of a process flow of an agent role switching method based on multimodal perception provided by an embodiment of the present invention;

[0023] Figure 2 A schematic block diagram of an agent role switching system based on multimodal perception provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0025] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0026] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0027] It should be further understood that the term "and / or" used in the present description and appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0028] See below Figure 1 An embodiment of the present invention provides an agent role switching method based on multimodal perception, which specifically includes: steps S101~S105.

[0029] Step S101: Collecting multimodal perception information of the user through a multimodal perception module;

[0030] Step S102: Determine candidate roles through an intelligent role switching decision mechanism based on the multimodal perception information;

[0031] Step S103: When the number of candidate roles is one, set the candidate role as the target role;

[0032] Step S104: When there are multiple candidate roles, one of the candidate roles is selected as the target role through a multi-role conflict resolution mechanism, or the candidate roles are combined into the target role;

[0033] Step S105: Switch the current role of the agent to the target role based on the preset role model library.

[0034] In this embodiment, a multimodal perception module first collects the user's multimodal perception information, then utilizes an intelligent role-switching decision-making mechanism to determine candidate roles. If there is only one candidate role, that candidate role is set as the target role. If there are multiple candidate roles, a multi-role conflict resolution mechanism is used to select one as the target role, or multiple candidate roles are combined into one target role. Finally, the agent's role is switched to the target role based on a pre-set role model library. This approach not only overcomes issues such as cumbersome user operations and a lack of contextual awareness during role switching, but also addresses multiple concurrent role scenarios, thereby improving the agent's role switching effectiveness and enhancing the user experience.

[0035] In particular, the multimodal sensing-based agent role switching method provided in this embodiment is particularly suitable for children's education. For example, the multimodal sensing-based agent role switching method can be applied to an agent device, which can interact with users such as children or parents to enable children to experience different agent roles in different time periods, thereby enhancing the effectiveness of children's education. Of course, the method provided in this embodiment can also be applied in many other scenarios, such as:

[0036] (1) Family education auxiliary equipment or applications. As intelligent auxiliary tools for family education, they can automatically switch roles such as teacher, friend, storyteller, etc. according to the child’s emotional state and learning progress, provide personalized learning content and emotional companionship, and effectively enhance the family learning experience;

[0037] (2) School education supplement. As an intelligent supplement to school education, it can identify children's confusion points in different subjects, automatically switch to the role of expert to provide targeted explanations, and at the same time, switch to the role of motivator when children's attention is distracted, thereby improving learning efficiency;

[0038] (3) Special education support. Provide customized support for children with special educational needs (such as autism, attention deficit, etc.), adjust the interaction mode according to the real-time response of the children, and switch to professional roles such as psychological counseling, speech therapy or behavioral guidance when needed;

[0039] (4) Children’s entertainment platform. As an intelligent partner of the children’s entertainment platform, it can seamlessly switch roles such as game partner, knowledge guide, and creative inspirer according to children’s interests and emotions during the game, balancing entertainment and educational value;

[0040] (5) Children’s medical environment. In a children’s medical environment, the device can automatically switch to being a narrator (explaining medical procedures), a comforter (relieving anxiety), or a distracting playmate, depending on the treatment stage and the child’s emotional state, significantly improving the child’s medical experience.

[0041] (6) Museums and science and technology museums. As an intelligent explanation system for museums and science and technology museums, it can intelligently switch between roles such as professional guides, storytellers, and adventure guides according to children's age, interests, and questions, providing an immersive knowledge exploration experience;

[0042] (7) Assistance for families with multiple children. In families with multiple children of different ages, the system can simultaneously address the needs of different children, provide each child with a personalized interactive experience through multi-role concurrency and conflict resolution technology, and reduce the educational burden on parents;

[0043] (8) Cross-cultural education environment. In a cross-cultural education environment, teachers can switch to the role of cultural guide or language teacher based on the children’s cultural background and language habits, helping them better adapt to the new environment and promoting multicultural understanding.

[0044] In one embodiment, the multimodal perception information includes voice information, visual information, and touch information; and collecting the user's multimodal perception information through the multimodal perception module includes:

[0045] Extracting features from the speech information using a GFCC algorithm to obtain speech features;

[0046] Capturing the user's facial expression information from the visual information based on a temporal attention mechanism to obtain visual features;

[0047] performing touch intention recognition on the touch information by using an adaptive interference filtering algorithm to obtain touch features;

[0048] Multimodal fusion is performed on the voice features, visual features, and touch features to obtain the multimodal perception information.

[0049] In this embodiment, multimodal perception is performed across three dimensions: the user's voice, facial expressions, and touch operations on the agent, to obtain multimodal perception information. Specifically, for voice information, features are extracted using the GFCC algorithm to obtain voice features. For visual information, facial expressions are captured using a temporal attention mechanism to obtain visual features. For touch information, features are extracted using an adaptive interference filtering algorithm to obtain touch features. The resulting voice, visual, and touch features are then fused to form multimodal perception information.

[0050] In actual application scenarios, the multimodal perception process can be optimized for child users. For example, when extracting features through the GFCC algorithm, the high-frequency cutoff point can be set to 7000Hz (not the 4000Hz commonly used for adult speech). This is more suitable for capturing the high-frequency characteristics of children's voices, thereby improving the accuracy of children's speech recognition. The use of a temporal attention mechanism can also capture the rapid changes in children's expressions, thereby improving the sensitivity of emotion recognition. Touch intention analysis takes into account the instability of children's touch behavior and improves the accuracy of touch intention recognition through adaptive threshold adjustment.

[0051] Specifically, in practical applications, children's speech recognition faces the following unique challenges:

[0052] (1) Differences in acoustic characteristics: Children have short vocal tracts (about 9-12 cm for preschoolers and 14-18 cm for adults), resulting in significantly higher fundamental frequencies (about 250-400 Hz for children and 85-250 Hz for adults) and resonance peak frequencies than adults, making the spectral characteristics of children's speech significantly different from those of adults; (2) Instability of pronunciation: non-standard pronunciation, incomplete grammar, frequent pauses, etc.; (3) Background noise interference: Children's educational scenes are usually accompanied by complex noises such as toy sounds and multiple people talking; (4) Rich emotional expression: emotions change quickly and the tone fluctuates greatly.

[0053] To address this, this embodiment designs a specialized speech enhancement preprocessing process to improve recognition accuracy. This process includes four main steps: noise suppression (e.g., through voice activity detection (VAD) algorithms), speech enhancement, speech rate normalization, and feature extraction. Here, the traditional VAD algorithm is optimized for the characteristics of children's speech. The specific implementation is as follows:

[0054] Procedure ChildVAD(x[n], fs), where the input x[n] represents the audio signal and fs represents the sampling rate (Hz); parameter settings include a frame length of 20ms, which is shorter than that of adults, a frame shift of 10ms, and a child speech energy threshold adjustment factor α child is 1.3, and the children's speech zero-crossing rate threshold adjustment factor β child←0.8. When extracting features, calculate the short-time energy Em and zero-crossing rate Zm of each frame, and further perform threshold calculation. When TE←α child When the energy threshold is adaptive, the energy threshold is lowered to capture weak syllables. When TZ←β child When adaptively setting the zero-crossing rate threshold, the zero-crossing rate threshold is adjusted to accommodate high-frequency characteristics. Speech segment detection is then performed, using a dual-threshold decision and state machine for segment labeling. A minimum duration constraint (50ms) is applied to prevent false detection, and a maximum silence interval constraint (300ms) is applied to prevent over-segmentation. Finally, a sequence of segment labels is output.

[0055] Children's expression recognition faces the following unique challenges:

[0056] (1) Differences in facial features: The proportions of children's facial features are different from those of adults (their eyes are relatively larger and their facial contours are rounder), which results in a 15-20% decrease in the accuracy of adult expression recognition models on children; (2) Rapid changes in facial expressions: Children's emotional switching frequency is 2-3 times that of adults, and their expressions last for a short time, making it difficult to capture them with traditional sampling rates; (3) Exaggerated expressions: Children's expressions are usually more exaggerated than those of adults, and their emotional expressions are more direct, requiring a wider recognition range; (4) Large individual differences: Children of different ages have significant differences in the way they express their expressions, requiring highly adaptable recognition algorithms.

[0057] In view of the characteristics of children's expressions, this embodiment designs an expression recognition preprocessing process to improve the recognition accuracy. The process includes four main steps: high frame rate acquisition, facial feature enhancement, micro-expression capture and age adaptive processing. Specifically: According to the characteristics of children's facial features, the traditional facial detection algorithm is optimized to obtain a child face detection and feature enhancement algorithm. Among them, the input includes I-image frame, age_range-child age range; parameter settings include age adaptive parameters αchild←GetAgeAdaptiveParams(age_range); minimum face size min_face_size←0.1·frame_height; child face aspect ratio range aspect_ratio_range←[0.8, 1.2]. When detecting faces, the MTCNN optimized for children is used, that is, faces←MTCNN(I, α child) ; Then feature enhancement is performed, including eye area enhancement, mouth area enhancement, and age adaptive processing to achieve facial enhancement.

[0058] To capture children's rapidly changing expressions, this embodiment designs a corresponding micro-expression capture algorithm. The input includes a continuous facial image sequence face_sequence; parameter settings include a high frame rate sampling of 60fps (sampling_rate←60), a micro-expression detection window (frames) (detection_window←5), and a change detection threshold threshold of threshold←0.15. Keypoint tracking and micro-expression detection are then performed as shown below:

[0059] (1) Key point tracking:

[0060] Track the facial key points of the face sequence through the TrackFacialLandmarks function to obtain the key point sequence;

[0061] (2) Micro-expression detection:

[0062] Initialize an empty micro-expression set;

[0063] Traverse the key point sequence:

[0064] Calculate the motion amplitude within the specified detection window. If the motion amplitude exceeds a preset threshold, extract the micro-expression from the corresponding face sequence and add the extracted micro-expression to the micro-expression collection;

[0065] (3) Return the set of all detected micro-expressions

[0066] The algorithm mainly detects micro-expressions by analyzing the movement changes of facial key points. When a movement exceeding the threshold is detected, it is considered that a micro-expression has occurred and is extracted.

[0067] To handle the rapidly changing nature of children’s expressions, this embodiment designs a temporal attention mechanism, specifically as follows:

[0068] (1) Input expression feature sequence [f1, f2, ..., fT];

[0069] (2) Temporal coding: Apply position coding to the expression feature sequence to generate coding features containing temporal information;

[0070] (3) Attention calculation. First, the encoded features are converted into query (Q), key (K), and value (V) matrices through the weight matrix; then the attention score is calculated: the query is multiplied by the transpose of the key, divided by the scaling factor √dk, and the softmax function is applied. Then, the attention feature is generated: the attention score is multiplied by the value matrix;

[0071] (4) Multi-scale fusion. Perform multi-scale aggregation processing on attention features;

[0072] (5) Emotion prediction. Use the emotion classifier to process multi-scale features and generate emotion probability distribution;

[0073] (6) Output: Returns the probability distribution of emotions.

[0074] Furthermore, to improve the model's ability to recognize children's expressions, this embodiment uses the following data augmentation techniques:

[0075] (1) Age-stratified sampling: training data was constructed by stratification according to age groups: 0-1 years old, 2-3 years old, 4-5 years old, and 6-8 years old; (2) Expression intensity transformation: simulating the range of variation in children's expression intensity and generating expression samples of different intensities; (3) Micro-expression synthesis: generating intermediate transition expressions through interpolation to enhance the recognition ability of rapidly changing expressions; (4) Multi-angle enhancement: considering the characteristics of children's high mobility, generating multi-angle expression samples within the range of ±30°.

[0076] This embodiment designs a deep learning model architecture specifically for the characteristics of children's expressions. Through the above optimization, this embodiment can accurately identify children's seven basic emotions (happiness, sadness, anger, fear, surprise, disgust, and neutrality) and four complex emotions (confusion, concentration, boredom, and excitement), providing a reliable emotional basis for role switching decisions.

[0077] In practical applications, data is first input into the input layer, followed by the child face CNN, and then the spatial attention module, bidirectional LSTM, temporal attention module, multi-feature fusion layer and emotion classification layer are used to obtain the child expression recognition results.

[0078] Analyzing children's touch behavior faces the following unique challenges:

[0079] (1) Operation instability: Children have weak control over fine motor movements, and touch operations are often unstable and imprecise; (2) Frequent accidental touches: Children's frequency of accidental touches when using devices is 3-5 times that of adults, requiring stronger interference filtering capabilities; (3) Multi-touch confusion: Children tend to use multiple fingers to operate simultaneously, making it difficult to identify touch intentions; (4) Special touch patterns: Children's unique non-standard touch patterns such as patting and swiping require special identification.

[0080] In view of the characteristics of children's touch behavior mentioned above, this embodiment designs a special touch analysis and processing process to improve the accuracy of touch intention recognition. The process includes four main steps: touch data preprocessing, accidental touch filtering, touch pattern recognition, and intention inference. Specifically:

[0081] To address children's frequent accidental touches, this embodiment designs an adaptive noise filtering algorithm. The implementation procedure includes: procedure AdaptiveNoiseFilter(touch_events, age_range). The input includes the touch event sequence touch_events and the child's age range age_range; parameter settings include the minimum valid touch duration min_duration←GetAgeAdaptiveThreshold(age_range, "duration"), the minimum valid movement distance min_distance←GetAgeAdaptiveThreshold(age_range, "distance"), and the maximum pressure change max_pressure_var←GetAgeAdaptiveThreshold(age_range, "pressure"). During touch event analysis, the touch event filtering algorithm filters valid touch events, specifically including:

[0082] (1) Initialization. First, create an empty collection filtered_events to store filtered touch events;

[0083] (2) Basic filtering. Traverse all touch events and obtain the duration (duration), touch distance (distance), and pressure change (pressure_var) of each touch event. If the touch duration is less than the minimum threshold and the touch distance is less than the minimum threshold, skip this event (filtering out short and meaningless touches); if the pressure change is greater than the maximum threshold, skip this event (filtering out touches with unstable pressure); if the above filtering conditions are met, add the event to the filtered event set;

[0084] (3) Context-dependent filtering. Perform context-dependent filtering on the event set after basic filtering and call the ContextualFilter function;

[0085] (4) Output results. Returns a set of touch events filtered by the context.

[0086] The main purpose of the touch event filtering algorithm is to improve the quality of touch events through multi-layer filtering (basic attribute filtering and context-related filtering) and remove invalid or interfering touch inputs.

[0087] Furthermore, to identify children's unique touch patterns, this embodiment designs a special pattern recognition algorithm, including a children's touch pattern recognition and intention inference algorithm. Specifically:

[0088] 1. Children’s touch pattern recognition algorithm

[0089] (1) Input: touch_events - filtered touch event sequence;

[0090] (2) Feature extraction: Extracting features such as touch speed, acceleration, direction change, and pressure change;

[0091] (3) Mode classification: classify touch features into modes such as click, long press, slide, tap, swipe, and multi-finger operation;

[0092] (4) Child-specific pattern recognition: Each pattern is further analyzed. If it is a random scratch, it is marked as a "scribble" type and the confidence is calculated. If it is a tapping operation, it is marked as a "tapping" type and the confidence is calculated. If it is a multi-finger chaos operation, it is marked as a "multi_finger_chaos" type and the confidence is calculated.

[0093] (5) Output: Returns the recognized touch mode.

[0094] 2. Touch Intention Inference Algorithm

[0095] (1) Input: patterns - recognized touch patterns, ui_context - user interface context, and age_range - child age range;

[0096] (2) Intention mapping: Obtain an adaptive intention mapping table based on the child’s age range;

[0097] (3) Intent inference: Initialize an empty intent set, then for each touch mode, obtain the candidate intent corresponding to the mode, and calculate the intent score for each candidate intent (based on the mode and interface context), and then store the score in the intent set;

[0098] (4) Intent fusion: Fusion of the final intent based on the intent set and interface context;

[0099] (5) Output: Returns the final inferred intent.

[0100] In addition, to improve children's touch experience, this embodiment designs a touch feedback optimization strategy:

[0101] (1) Adaptive touch area: Dynamically adjust the effective area of ​​clickable elements according to the child's age and operation accuracy, providing a larger touch tolerance range for young children; (2) Multimodal feedback: Combining visual, auditory and tactile feedback to provide children with multi-channel operation confirmation and enhance operation perception; (3) Progressive guidance: When unstable touch is detected, progressive visual guidance is provided to help children complete fine operations; (4) Intention prediction and correction: Based on context and historical operations, predict the child's possible touch intention and perform intelligent operations in low-confidence situations.

[0102] In view of the fact that children have weak fine motor control ability, this embodiment designs a child touch interaction adaptive algorithm, which specifically includes an adaptive touch area algorithm, a multi-modal feedback mechanism, and a progressive guidance mechanism, among which:

[0103] 1. The adaptive touch area algorithm dynamically adjusts the touch area of ​​UI elements based on the child’s age and touch accuracy history. The algorithm includes:

[0104] (1) Input: interface element set, children's age range, and historical touch accuracy records;

[0105] (2) Calculate the expansion coefficient: a base coefficient based on age (the younger the age, the larger the coefficient), an adjustment coefficient based on touch accuracy (the lower the accuracy, the larger the coefficient), and a contextual coefficient based on the current activity;

[0106] (3) Apply expansion: Apply the final expansion factor to each UI element, and you can further set important elements to get an additional 20% expansion;

[0107] (4) Conflict resolution: dealing with regional conflicts that may arise after expansion;

[0108] (5) Output: Returns the adjusted touch area.

[0109] 2. A multimodal feedback mechanism provides multisensory feedback to enhance children’s touch experience. The mechanism includes:

[0110] (1) Input: touch events, interface context, and child age range;

[0111] (2) Touch analysis: classify touch types and calculate confidence;

[0112] (3) Multimodal feedback: visual feedback (enhanced when low confidence), auditory feedback (age-adaptive volume), and tactile feedback (if supported by the device);

[0113] (4) Feedback coordination: Coordinating feedback from different modes;

[0114] (5) Output: Return feedback results.

[0115] 3. Progressive guidance mechanism can help children complete complex touch operations:

[0116] (1) Input: touch history, target operation, and child age range;

[0117] (2) Difficulty assessment: assess the operational difficulty and user skills, and calculate the required guidance level;

[0118] (3) Guidance strategies: minimal guidance—providing only visual cues, moderate guidance—providing animated demonstrations and corrective feedback, and intensive guidance—providing step-by-step guidance;

[0119] (4) Guided adaptation: updating the user skill model;

[0120] (5) Output: Return the guidance result.

[0121] Through these optimizations, this embodiment can accurately identify children's touch intentions, provide a smooth interactive experience even in cases of unstable operation, multi-touch confusion, and frequent accidental touches, and provide a reliable touch behavior basis for role switching decisions.

[0122] In one embodiment, the performing multimodal fusion on the voice features, visual features, and touch features to obtain the multimodal perception information includes:

[0123] Performing projection mapping on the speech features, visual features, and touch features respectively to obtain corresponding speech feature vectors, visual feature vectors, and touch feature vectors;

[0124] The speech feature vector, the visual feature vector and the touch feature vector are fused based on a cross-modal attention mechanism to obtain the multimodal perception information.

[0125] In this embodiment, a multimodal signal fusion algorithm based on an attention mechanism is used to effectively integrate speech, vision, and touch signals. The fusion process can be expressed as follows: Based on the characteristics of speech, vision, and touch in children's education scenarios, this embodiment designs a multimodal signal fusion algorithm based on an attention mechanism to effectively integrate different modal signals. This algorithm includes four main steps: feature extraction and preprocessing, modal feature projection, cross-modal attention calculation, and feature fusion.

[0126] Specifically, when performing modal feature representation, the feature representations of the three modalities are as follows:

[0127] Speech feature A: Contains GFCC features, speech emotion features and semantic features, expressed as a vector ;

[0128] Visual features V: Contains facial expression features, micro-expression sequences and attention states, represented as vectors ;

[0129] Touch feature T: includes touch pattern feature, intention feature and stability index, expressed as a vector ;

[0130] Among them, dA, dV, and dT represent the dimensions of speech, vision, and touch features, respectively.

[0131] Since features of different modalities have different dimensions and distribution characteristics, this embodiment first maps them to the same feature space through linear projection:

[0132] ;

[0133] in, 、 and are all learnable projection matrices, and d is the dimension of the unified feature space.

[0134] The core of the cross-modal attention mechanism is to calculate the mutual attention weights between different modal features. Given the projected feature X i 、X j and X k (corresponding to W A A.W. V V and W T T), the attention is calculated as follows:

[0135] Input the projected modal feature X i 、X j and X k, The query-key-value calculation is performed through the query matrix, key-value matrix and value matrix, and the modal reliability bias is added to calculate the attention score. Then the attention weight normalization and weighted feature calculation are performed to obtain the attention calculation result.

[0136] It should be noted that the modality reliability bias is one of the innovations of this algorithm. It dynamically adjusts the attention score based on the quality and reliability of each modality feature. The bias is calculated as follows:

[0137] ;

[0138] Where: A , β V and β T is the reliability bias term for different modes, γ A , γ V and γ TAll are learnable modal weight coefficients, SNR(A) represents the speech signal-to-noise ratio, ranging from [0,1], EmotionConf(A) represents the confidence of speech emotion recognition, ranging from [0,1], FaceVisibility(V ) represents facial visibility, ranging from [0,1], ExpressionConf(V ) represents the confidence of expression recognition, ranging from [0,1], TouchStability(T) represents touch stability, ranging from [0,1], and IntentConf(T) represents the confidence of touch intent recognition, ranging from [0,1].

[0139] In order to capture the feature relationships of different subspaces, this embodiment adopts a multi-head attention mechanism. The final multimodal fusion feature is calculated by the following method:

[0140] ;

[0141] in, is a nonlinear activation function, such as ReLU or GELU.

[0142] Furthermore, for different scenarios and tasks, this embodiment designs a dynamic fusion strategy to automatically adjust the fusion weight according to the task type. Specifically, the input includes the original modal features A, V and T and the task type task_type, and the task adaptability weight is λ A ,λ V ,λ T ←GetTaskWeights(task_type). Then perform feature projection:

[0143] X A ←λ A W A A

[0144] X V ←λ V W V V

[0145] X T ←λ T W T T

[0146] Attention Fusion:

[0147] Z←MultiHeadCrossModalAttention(X A , X V , X T , h)

[0148] Residual connection and normalization:

[0149] Z′←LayerNorm(Z + Concat(X A , X V , X T ))

[0150] Feedforward Network:

[0151] F←FFN(Z′)

[0152] F′←LayerNorm(F + Z′)

[0153] Output:

[0154] return F′

[0155] end procedure

[0156] The GetTaskWeights function returns the appropriate modality weight based on the task type, for example:

[0157] Speech understanding task: λ A = 0.6, λ V = 0.3, λ T = 0.1;

[0158] Emotion recognition task: λ A = 0.4, λ V = 0.5, λ T = 0.1;

[0159] Interaction intention recognition task: λ A = 0.3, λ V = 0.2, λ T = 0.5.

[0160] Through these optimizations, this embodiment effectively integrates information from three modalities: speech, vision, and touch, providing a comprehensive and accurate multimodal perception foundation for role-switching decisions. Multimodal fusion significantly improves the system's understanding and adaptability, particularly in situations where children experience significant emotional fluctuations and incomplete expressions.

[0161] In general, the innovation of this embodiment in multimodal perception is mainly reflected in the following aspects:

[0162] (1) In view of the characteristics of children's speech, the GFCC feature extraction algorithm was optimized to improve the capture capability of the high-frequency band (4000-7000Hz) and better adapt to the characteristics of children's voice; (2) In terms of expression recognition, the temporal attention mechanism and micro-expression capture technology are used to accurately identify children's rapidly changing emotional expressions; (3) In view of the unstable characteristics of children's touch behavior, adaptive threshold adjustment and accidental touch filtering algorithms are designed to improve the accuracy of touch intention recognition; (4) At the multimodal fusion level, a new cross-modal attention mechanism is proposed. Through the three steps of feature projection, attention calculation and feature fusion, adaptive weight adjustment of different modal signals is achieved, which effectively solves the information imbalance problem in traditional fusion methods.

[0163] In one embodiment, determining the candidate roles through an intelligent role switching decision mechanism based on the multimodal perception information includes:

[0164] The multimodal perception information is contextually understood based on a multi-head self-attention mechanism, and the multimodal perception information is used in conjunction with a dialogue state tracker to identify intent, thereby obtaining intent information. The intent information includes explicit requests, changes in dialogue topics, and changes in emotional states.

[0165] A multi-factor weighted scoring model is used to assign different weights to the intention information, and the corresponding candidate roles are determined based on the intention information after weight assignment.

[0166] In this embodiment, the role switching decision-making mechanism utilizes contextual understanding based on an attention mechanism, a conversation state tracker, and a smooth transition strategy to achieve intelligent and natural role switching. This embodiment employs a multi-factor fusion decision-making mechanism that comprehensively considers factors such as explicit user requests, changes in conversation topics, emotional state shifts, and changes in learning progress. In practice, explicit user requests are given the highest weight, followed by changes in conversation topics, and emotional state shifts and changes in learning progress are relatively low weighted. Decision-making is achieved through a multi-factor weighted scoring model, enabling automatic and intelligent role switching.

[0167] Specifically, this embodiment designs a contextual understanding algorithm based on a multi-head self-attention mechanism to address the characteristics of children's conversations. This algorithm can effectively capture the long-range dependencies and topic jumping characteristics in children's conversations. The process includes: inputting the conversation history sequence into the embedding layer, and then passing through position encoding, multi-head self-attention, layer normalization, feedforward network, and context representation. Here, this embodiment uses the multi-head self-attention mechanism to achieve a deep understanding of the conversation history. Furthermore, this embodiment optimizes the standard Transformer architecture in the following ways to address the characteristics of children's conversations:

[0168] (1) Positional encoding enhancement: Relative positional encoding is used instead of absolute positional encoding to better handle topic jumping in children's conversations; (2) Attention window adjustment: A larger attention window is set (considering up to the past 20 rounds of conversation) to capture long-distance dependencies; (3) Emotional marker injection: Emotional markers are added to the input embedding to enhance the ability to perceive children's emotional changes; (4) Incomplete expression processing: A special filling mechanism is added to handle incomplete expressions commonly seen in children.

[0169] Through these optimizations, this embodiment can more accurately understand the contextual information in children's conversations, especially in handling topic jumps, incomplete expressions, and emotional changes, providing a reliable contextual basis for role switching decisions.

[0170] This embodiment also implements a state tracker specifically tailored to the characteristics of children's conversations. By continuously monitoring and updating the conversation state, the conversation state tracker provides a key basis for role-switching decisions. This algorithm can capture changes in intent, emotional fluctuations, and topic shifts in children's conversations, enabling precise tracking of the conversation flow. This mathematical representation enables the system to fully understand the conversation context, providing a reliable foundation for intelligent role switching. Specifically, it is expressed as follows:

[0171] ;

[0172] in, , represents the dialogue state vector at time step t, which contains the following key components:

[0173] (1) Intent∈{1, 2, . . . , I}, which represents the currently recognized user intent category, including: query intent, such as seeking information, asking questions, exploring knowledge, etc.; command intent, such as requesting to perform a task, switch roles, adjust the system, etc.; expression intent, i.e. sharing feelings, expressing emotions, telling experiences, etc.; social intent, such as greetings, thanks, goodbyes, and other social interactions; game intent, such as requesting games, participating in interactive activities, etc.

[0174] (2) , indicating the slot filling status, records the key information in the conversation, including role slot: the role information currently requested or mentioned; topic slot: the topic or knowledge field currently discussed; task slot: the specific task currently performed or requested; emotion slot: the emotional needs or status expressed by the user; time slot: the time information related to the conversation;

[0175] (3) , represents the emotional state vector, including: valence ∈ [-1, 1], emotional value, positive and negative polarity; arousal ∈ [0, 1]: emotional intensity, activation level; type ∈ {1, 2, . . . , E}: discrete emotion category; confidence ∈ [0, 1]: confidence of emotion recognition;

[0176] (4) , represents the topic representation vector, including: current_topic ∈ {1, 2, . . . , T}: current topic category; topic_embedding ∈ Rte: topic semantic embedding; topic_duration ∈ R+: current topic duration; topic_shift_prob ∈ [0, 1]: topic transition probability;

[0177] (5) ActiveRole∈{1, 2, . . . , R}, represents the current active role, which can be the main role currently interacting with the user, can be a single role or a combination of multiple roles, and can contain role activity and switching probability information;

[0178] (6) d = 1 + s + e + t + 1: The total dimension of the dialogue state vector.

[0179] Represents the state update function, the input is the state S at the previous moment t-1 、The current user inputs u t and the system response r at the previous moment t-1 , the output is the updated dialogue state S t , implemented by combining multiple specialized functions to handle different aspects of the state:

[0180] Update(S, u, r) = {I(u), Slots(S, u), E(u), T(S, u), R(S, u, r)}.

[0181] It represents the intent recognition function. Its input is the user's current input u, which contains multimodal information. Its output is the recognized intent category. It is implemented as a multimodal intent classifier based on deep learning:

[0182] ;

[0183] where P(i|u) represents the conditional probability of intention i given user input u.

[0184] represents the slot filling function, which takes the previous state S and the current user input u as input and outputs the updated slot state. It is implemented using the sequence labeling and value extraction model:

[0185] ;

[0186] Among them, Extract_Values ​​means extracting slot values ​​from user input, and Update_Slots updates existing slots.

[0187] Represents a sentiment analysis function, whose input is the user's current input u and output is the sentiment state vector, implemented as a multimodal sentiment analysis model:

[0188] ;

[0189] Among them, EmotionClassifier is a sentiment analyzer that integrates speech, vision and text features.

[0190] It represents the topic recognition function, which takes the previous state S and the current user input u as input and outputs the topic representation vector. It is implemented by topic modeling and change detection:

[0191] ;

[0192] Among them, TopicShift means detecting topic changes, NewTopic means creating a new topic, and UpdateTopic means updating an existing topic.

[0193] represents the active role update function, which takes the previous state S, the current user input u, and the previous system response r as input and outputs the updated active role. It is implemented by rule-based and probability-based role selection:

[0194] ;

[0195] Among them, HasRoleRequest indicates whether there is a clear role request, RequestedRole indicates extracting the requested role, and SuggestRole indicates recommending a suitable role based on the status.

[0196] In view of the characteristics of children's conversations, this embodiment optimizes the standard DST as follows:

[0197] (1) Enhanced children's language comprehension: Targeting the incompleteness and non-standardization of children's language, including semantic completion, i.e. automatically completing the semantics of children's incomplete expressions; error-tolerant processing, i.e. handling pronunciation errors and grammatical errors in children's language; and context association, i.e. using context to infer implicit intentions and referents. (2) Low-threshold slot filling: Targeting the characteristics of children's implicit expressions, including lowering the certainty threshold, i.e. accepting slot values ​​with a lower confidence level; multi-round accumulation, i.e. accumulating slot information through multiple rounds of dialogue; and implicit extraction, i.e. extracting implicit slot values ​​from non-direct expressions. (3) Emotion-sensitive tracking: Targeting the characteristics of children's large emotional fluctuations, including micro-expression capture, i.e. identifying subtle emotional change signals of children; emotional trend analysis, i.e. tracking emotional change trends rather than instantaneous states; and multimodal fusion, i.e. comprehensively judging emotional states based on speech, expression, and behavior. (4) Topic jump adaptation: In view of the fact that children frequently switch topics, the following methods are adopted: loose topic boundaries, which allow greater tolerance for topic changes; multi-topic parallelism, which tracks multiple possible active topics at the same time; and topic association graph, which constructs an association network between topics to predict possible jumps.

[0198] Through these optimization strategies, this embodiment can more accurately track the children's dialogue status, especially in handling incomplete expressions, emotional fluctuations and topic jumps, providing a reliable status basis for role switching decisions.

[0199] In one embodiment, when there are multiple candidate roles, selecting one of the candidate roles as the target role through a multi-role conflict resolution mechanism, or combining the candidate roles into the target role, includes:

[0200] The conflict detection algorithm is used to perform semantic similarity analysis and logical consistency analysis on multiple candidate roles. The redundancy and contradiction between multiple candidate roles are determined based on the results of the semantic similarity analysis and the logical consistency analysis, and the role conflict detection results are obtained:

[0201] Weights are assigned to multiple candidate roles according to a preset dynamic weight assignment strategy, and a weight matrix is ​​constructed based on the weights. The dynamic weight assignment strategy includes: comprehensively calculating the response priority of each candidate role based on preset dimensions and assigning weights accordingly.

[0202] According to the following formula, the target role is determined by combining the role conflict detection results and the weight matrix:

[0203] ;

[0204] in, Indicates the target role, represents the candidate role with the highest weight, represents the maximum weight value, represents the second highest weight value, represents the single role selection threshold, represents the synthetic role selection threshold, Synthesize represents the role synthesis function, Represents the input value of the character synthesis function.

[0205] In this embodiment, a multi-role conflict resolution mechanism can resolve response conflicts in multi-role concurrent scenarios. A conflict detection algorithm can identify two main conflict types: redundancy and contradiction. A dynamic weight allocation mechanism is used to calculate a weight matrix. The priority of each role's response is comprehensively calculated based on contextual relevance, emotional adaptability, and educational value. This allows for intelligent selection of the optimal response or synthesis of multiple role responses based on the weight matrix, ensuring interactive coherence and educational value.

[0206] Specifically, the conflict detection algorithm described in this embodiment is used to identify redundancy and contradiction between responses. Its input is the response vector R of the two roles. i and R j The output is the conflict type (redundant, contradictory or non-conflicting). The logical relationship provides a basic judgment for the subsequent weight allocation and response selection. The weight matrix construction is to assign weights to each role response. Its input is the role response vector R i , the output is the response weight W(R i ), the logical relationship is to consider five key factors (basic weight, context relevance, relationship between roles, historical performance and educational value). The response selection and synthesis strategy is used to select or synthesize responses based on conflict type and weight. The input is the weight matrix W(R i ) and conflict type, and the output is the final selected response R*. The logical relationship is to use the weight matrix to determine the response selection and adopt different strategies according to the conflict type.

[0207] In practical applications, conflict detection algorithms can first identify conflict types between responses. A weight matrix construction method then assigns weights to each response. Finally, a final response is selected or synthesized based on the conflict type and weights. For example, "When a child asks, 'Why is the sky blue?' The scientific explorer character might offer a scientific explanation, while the storyteller character might tell a myth about the sky's color. The system then calculates the weights of each response based on the current educational objectives, the child's interests, preferences, and emotional state, and selects the most appropriate response or synthesizes responses from multiple characters."

[0208] The conflict detection algorithm accurately identifies conflicts in multi-role responses by analyzing the semantic similarity and logical consistency between responses. The algorithm can detect two main types of conflicts: redundancy and contradiction, providing a foundation for subsequent conflict resolution, as follows:

[0209] ;

[0210] in, , represents the response vectors of the two characters, which include the following dimensions: (1) semantic representation, i.e., the semantic embedding vector of the response content; (2) emotional characteristics, i.e., the emotional tendency and intensity of the response; and (3) educational value, i.e., the educational significance score of the response.

[0211] , represents the semantic similarity function, which is used to calculate the similarity between two responses in the semantic space, and takes into account vocabulary overlap, semantic association and expression. Its value range is [0,1], where 1 indicates complete similarity.

[0212] , represents the logical consistency function, which is used to evaluate the logical compatibility of two responses and the consistency of analyzing opinions, facts and reasoning. Its value range is [0,1], and 1 indicates complete consistency.

[0213] , represents the conflict determination threshold, where θ r represents the redundancy determination threshold, with a typical value of 0.8-0.9, θ c Indicates the contradiction determination threshold, with a typical value of 0.2-0.3.

[0214] The weight matrix construction method is based on a multi-factor scoring approach. By comprehensively considering five key factors, including the role's basic weight, contextual relevance, inter-role relationships, historical performance, and educational value, it assigns a reasonable weight to each role response. This method effectively balances the strengths of different roles and provides a quantitative basis for conflict resolution, as follows:

[0215] ;

[0216] in, , represents the response R of role i i The comprehensive weight of .

[0217] , represents the role-based weight function, whose input is the response vector R i The output is the basic weight score, with a value range of [0,1], which is calculated based on the role preset weight and the current scene adaptability:

[0218] ;

[0219] Here, ω i ∈ [0,1], represents the preset weight of role i, and SceneCompat is the scene adaptation function.

[0220] , represents the context relevance scoring function, whose input is the response vector R i And the context vector C, the output is the relevance score, the value range is [0,1], calculated based on semantic similarity and topic matching:

[0221] ;

[0222] Here, is the balance coefficient, usually 0.6-0.7.

[0223] , represents the relationship function between roles, whose input is the response vector R i and other role response vectors R j The output is the role relationship score, with a value range of [0,1], and is calculated based on the role collaboration and complementarity:

[0224] ;

[0225] Here, β∈[0,1] is the balance coefficient, usually 0.5, and n is the total number of roles.

[0226] , represents the historical performance scoring function, whose input is the response vector R i The output is the historical performance score, with a value range of [0,1], calculated based on historical interaction effects and user feedback:

[0227] ;

[0228] Here, δ∈[0,1] is the balance coefficient, which is usually taken as 0.7.

[0229] , represents the educational value scoring function, whose input is the response vector R i The output is an educational value score with a range of [0,1], calculated based on knowledge content, inspiration, and age appropriateness:

[0230] Education(R i )=γ1·Knowledge(R i ) +γ2·Inspiration(R i )+γ3·AgeAppropriate(R i );

[0231] Here, γ1,γ2,γ3∈[0,1] and γ1+γ2+γ3=1.

[0232] , all represent weight coefficients and satisfy λ1+λ2+λ3+λ4+λ5= 1.

[0233] In practical applications, it can be dynamically adjusted according to the scenario, for example:

[0234] (1) Knowledge learning scenario: λ1=0.1, λ2=0.3, λ3=0.1, λ4=0.1, λ5=0.4;

[0235] (2) Emotional communication scenario: λ1=0.2, λ2=0.3, λ3=0.2, λ4=0.2, λ5=0.1;

[0236] (3) Game interaction scenario: λ1=0.2, λ2=0.2, λ3=0.3, λ4=0.2, λ5=0.1;

[0237] To ensure the comparability of response weights of different roles, the calculated weights are normalized:

[0238] ;

[0239] Where n is the number of currently active roles, W′(R i ) is the normalized weight, satisfying ∑n i =1,W′(R i ) = 1.

[0240] The response selection and synthesis strategy implements an intelligent response processing mechanism based on the weight matrix and conflict type. This strategy can automatically decide whether to select a single response or synthesize multiple role responses based on the weight difference, providing children with the best interactive experience. Specifically:

[0241] ;

[0242] in, represents the final selected response vector. Indicates the role response with the highest weight. A single role response with the highest weight value can be selected and is suitable for situations where the weight difference is significant. Represents the maximum weight value, that is, the highest weight value among all role responses, and is used to evaluate the weight differences between responses. Indicates the second highest weight value, that is, the second highest weight value among all role responses, which is used to calculate the weight difference to determine whether a synthetic response is needed. It represents the single response selection threshold. When the weight difference exceeds this threshold, a single response is selected. The typical value is 0.2-0.3, which can be adjusted dynamically according to the scenario. Higher δ values ​​tend to synthesize more responses. Indicates the weight threshold of the synthetic response. Only responses with weights exceeding this threshold will be included in the synthesis. The typical value is 0.3-0.4, which can be adjusted dynamically according to the scene. sA value of reduces the number of responses that participate in the synthesis. Represents a response synthesis function, the input is a set of responses that meet the weight threshold , the output is the synthesized response vector, which is achieved by extraction, integration and template:

[0243] .

[0244] Here, Represents the core content extraction function, which takes the response set Rset as input and outputs the core content set of each response. It is implemented based on key sentence extraction and semantic importance scoring: ;

[0245] Among them, the CoreContent function extracts the core content of a single response.

[0246] Represents the role information acquisition function, the input is the response set R set , the output is the corresponding role information set, which is achieved by mapping the response to its corresponding role identifier and characteristics:

[0247] ;

[0248] Among them, ID i is the role identifier, F i is the character feature vector.

[0249] Represents a template generation function. Its input is a core content set and a role information set. Its output is a synthetic response generated based on the template. The implementation method is to select an appropriate template according to the conflict type and fill in the content:

[0250] ;

[0251] Among them, Tcomplementary, Talternative and Tconflicting are complementary, alternative and conflicting template functions respectively.

[0252] In one embodiment, switching the current role of the agent to the target role based on a preset role model library includes:

[0253] According to the following formula, the three-stage Markov process of prediction-transition-confirmation is used to switch roles:

[0254] P(r t |r t-1 ,μ t ,S t )=P(r tannounce |r t-1 ,S t )·P(r t transition |r t announce ,μ t )·P(r t confirm |r t transition );

[0255] in, represents the role response vector in the preview phase, represents the role response vector in the transition phase, Represents the role response vector in the confirmation phase.

[0256] In this embodiment, the sampling progressive role transition technology implements a smooth transition strategy for role switching through a three-stage Markov process of prediction-transition-confirmation, avoiding abrupt role conversions and significantly improving the user experience. Specific examples are as follows:

[0257] P(r t |r t-1 ,μ t ,S t )=P(r t announce |r t-1 ,S t )·P(r t transition |r t announce ,μ t )·P(r t confirm |r t transition );

[0258] in, Represents the role response vector in the preview phase, which is used to switch the intention prompt information, maintain the current role's tone characteristics, and introduce the target role elements. Represents the role response vector in the transition phase, which is used to integrate the characteristics of the current role and the target role, achieve gradual role transition and maintain dialogue coherence. Represents the confirmation stage role response vector, which is used to fully switch to the target role, confirm the new role identity, and establish a new interaction tone.

[0259] In one embodiment, the agent role switching method based on multimodal perception further includes:

[0260] Obtaining a user feature vector and constructing a user profile based on the user feature vector;

[0261] Setting an adaptation function based on a preset character style, and setting a character behavior mode through the adaptation function;

[0262] A role model is constructed by combining the user portrait and the role behavior pattern, and the role model is adaptively adjusted using a crowd portrait adaptation mechanism, thereby constructing a role model library containing multiple role models.

[0263] The role model library in this embodiment not only stores the knowledge graph and behavior patterns of pre-set roles, but also enables adaptive character adjustments. The role knowledge graph uses a hierarchical structure, distinguishing between core knowledge and extended knowledge. Character behavior patterns are finely defined using three sets of parameters: language style, interaction style, and educational style. Character personalization parameters support dynamic adjustment. A crowd-profile adaptation mechanism automatically adjusts character performance based on the age, cognitive level, and learning style of different children, achieving a truly personalized educational experience.

[0264] The role model library in this embodiment utilizes a hierarchical design, enabling a refined representation of role knowledge and behavior. For example, when constructing a role model library for children, a comprehensive user profile is first constructed using the child's feature vectors (age, cognitive level, learning style, and personality traits). Next, adaptation functions are designed across three dimensions: language style, interaction style, and educational style, enabling precise adjustments to the role's behavior. Finally, a weight matrix mechanism is introduced to dynamically adjust the influence of each adaptation function based on different scenarios. This mechanism enables the system to adaptively adjust role behavior based on the child's feature vectors, providing a personalized educational experience.

[0265] Specifically, the role knowledge graph is represented by triples:

[0266] ;

[0267] Among them, E is the entity set (Entities), which represents the concepts, objects, etc. in the knowledge graph. R is the relationship set (Relations), which represents the relationship between entities. , represents a set of triples, each triple (h, r, t)∈T represents the head entity h connected to the tail entity t through the relation r.

[0268] The knowledge graph consists of a core layer and an extension layer:

[0269] ;

[0270] Among them, G core Represents core knowledge, including the basic knowledge of the role, which is fixed and unchanged. extension Represents extended layer knowledge, which can be updated dynamically over time.

[0271] The state of the extended layer knowledge at time t is represented as G t extension , over time, new knowledge is added and outdated knowledge is removed:

[0272] ;

[0273] Where NewKnowledge(t) represents the set of new knowledge triples generated at time t, and ObsoleteKnowledge(t) represents the set of obsolete knowledge triples that need to be removed at time t.

[0274] The character's behavior pattern is represented by three sets of parameter vectors:

[0275] ;

[0276] Among them, the language style parameter vector , l1 is vocabulary complexity (Vocabulary Complexity), l2 is sentence variety (Sentence Variety), l3 is expression (Expression Style) ... l m Other related language features.

[0277] Interaction style parameter vector , i1 is proactiveness, i2 is response speed, i3 is emotional expression...i n For other interactive features.

[0278] Educational style parameter vector , e1 is the guidance style (Guidance Style), e2 is the feedback type (Feedback Type), e3 is the challenge level (Challenge Level) ... e p For other educational style characteristics.

[0279] The character behavior parameters can be optimized and adjusted based on user feedback. The optimization process can be expressed as:

[0280] ;

[0281] in, is the character behavior parameter at time t, The behavior adjustment amount is calculated based on user interaction feedback. It consists of the following three parts:

[0282] ;

[0283] Among them, λ L ,λ I ,λ E are the weight factors for language, interaction and educational style adjustment, 、 、 They respectively represent the adjustment vectors calculated based on the feedback.

[0284] The character's final behavior is adjusted based on the smoothing factor of historical data Perform weighted averaging:

[0285] ;

[0286] Among them, B adjusted is the adjusted behavioral parameter, Controls how much historical behavior influences new behavior.

[0287] The crowd portrait adaptation mechanism, based on the child's feature vector, enables personalized adjustments to the character's behavior. This mechanism uses feature matrix mapping to enable the character to dynamically adjust its behavior pattern based on the child's age, cognitive level, learning style, and personality traits:

[0288] ;

[0289] ;

[0290] Among them, B adapted is the role behavior parameter after adaptation, B role is the basic behavioral parameter of the character, C = (a, c, l, p) is the child feature vector, a is age, c is cognitive level, l is learning style, and p is personality trait.

[0291] f L 、f I 、f E These are the adaptation functions for language style, interaction style, and educational style. The adaptation function is implemented through weighted adjustment. Take the language style adaptation function as an example:

[0292] ;

[0293] ;

[0294] in, is the language style parameter vector, is the language style adjustment amount based on children’s characteristics, and WL is the language style adaptation weight matrix, which is obtained through optimization of experimental data.

[0295] The adaptation functions for interactive and educational styles take a similar form:

[0296] ;

[0297] ;

[0298] Among them, W I and W E are the adaptation weight matrices for interactive style and educational style, respectively.

[0299] Several embodiments are provided below to illustrate the application scenarios of the agent role switching method based on multimodal perception.

[0300] Example 1: Children's education application based on tablet computer

[0301] In this embodiment, the multimodal sensing-based agent role switching method is embedded in a children's education application on a tablet computer. The application contains five preset roles: math teacher, science explorer, storyteller, game partner, and life coach.

[0302] The multimodal perception module collects children's facial expressions using the tablet's front camera, voice data using the built-in microphone, and touch data via the touchscreen. The speech feature recognition unit processes the speech signal using the GFCC feature extraction algorithm, with parameters set as follows: 16kHz sampling rate, 25ms frame length, 10ms frame shift, 0.97 pre-emphasis factor, 7000Hz high-frequency cutoff, 40 Mel filter banks, and 13th-order cepstral coefficients. The visual emotion recognition unit uses an improved VGG-Face network combined with a temporal attention mechanism to capture children's micro-expressions. Emotion classification includes six basic emotions (happiness, sadness, anger, surprise, fear, and disgust) as well as a neutral state, achieving an 87.5% recognition accuracy. The touch intention analysis unit uses an adaptive threshold adjustment algorithm to dynamically adjust touch determination parameters based on the child's age, improving touch intention recognition accuracy by 15.3% for children aged 5-8.

[0303] The conversation state tracking module implements a child conversation optimization strategy. It employs a low-threshold slot-filling technique, lowering the slot certainty threshold from the standard 0.7 to 0.5, and collects information through multiple rounds of accumulation. It also implements a relaxed topic boundary strategy, allowing jumps with topic similarity as low as 0.4 and simultaneously tracking up to three active topics, effectively addressing children's frequent topic switching.

[0304] The role switching module automatically switches between five roles based on multimodal perception results and the conversation state. It implements a multi-factor weighted scoring model with weights configured as follows: explicit user request (0.5), change in conversation topic (0.3), emotional state shift (0.1), and change in learning progress (0.1). For example, if a child asks a math question (a change in topic), the system automatically switches to the math teacher role; if a child appears tired or inattentive (a change in emotional state), the system switches to the playmate role for interactive play; and if a child explicitly requests a story (an explicit user request), the system switches to the storyteller role.

[0305] During the role switching process, a three-stage Markov transition strategy is adopted. For example, when switching from a math teacher to a game partner, the user will first say in the role of the math teacher: "I see you're a little tired. Do you want to play a game with your game partner, Xiao Ming, to relax?" (Preview phase, maintaining the tone characteristics of the current role and introducing elements of the target role). Then, the user enters the transition phase: "Xiao Ming is here! Let's play a number game in this embodiment!" (Integrating the characteristics of the current and target roles), and finally, the user completely switches to the game partner role: "I'm Xiao Ming. Let's play an interesting number chain game in this embodiment!" (Confirmation phase, completely switching to the target role and establishing a new interactive tone). Frequent switching detection is also implemented. When more than two role switching requests are detected within 30 seconds, a 5-10 second transition buffer is added to avoid cognitive burden.

[0306] When children interact with multiple characters simultaneously, the conflict resolution module handles potential conflicting responses. Based on the conflict detection algorithm, a BERT-based semantic similarity model (redundancy threshold θr = 0.85) and a knowledge graph-based logical consistency check (contradiction threshold θc = 0.25) are used. For example, when a child asks, "Why is the sky blue?" the Science Explorer character might offer a scientific explanation: "The sky is blue because the blue light in sunlight is scattered more by air molecules." Meanwhile, the Storyteller character might recount a mythical story about the sky's color: "A long time ago, the sky was white. Then the blue spirits sprinkled magic paint..."

[0307] The weights of each response were calculated using the weight matrix construction method. In the knowledge learning scenario, the weight coefficients were configured as follows: base weight (λ1 = 0.1), contextual relevance (λ2 = 0.3), role-relationship (λ3 = 0.1), historical performance (λ4 = 0.1), and educational value (λ5 = 0.4). The weight matrix was normalized to ensure that the sum of the weights was 1.

[0308] Based on the response selection strategy, when the difference between the highest and second-highest weights exceeds a threshold of δ = 0.25, the response from the character with the highest weight is selected; otherwise, responses from multiple characters are synthesized. In the above example, if the current scenario is science learning, the response weight of the Science Explorer character might be 0.65, and that of the Storyteller might be 0.35. The system will select the Science Explorer's response. If the weights are close (such as 0.55 and 0.45), a synthesized response is generated: "The sky is blue because the blue light in sunlight is scattered more by air molecules. Interestingly, in some ancient myths, people believed that the color of the sky came from magic paint sprinkled by blue spirits."

[0309] The role model library includes role knowledge graphs, behavioral patterns, and crowd portrait adaptation mechanisms. The performance of each role is automatically adjusted based on the child's age, cognitive level, and learning style. For example, for children aged 5-6, the language style parameter vector L of the math teacher role is adjusted to lower vocabulary complexity (l1=0.3) and higher repetition (l2=0.7), using simpler vocabulary and more visual aids; for children aged 7-8, it is adjusted to higher vocabulary complexity (l1=0.6) and lower repetition (l2=0.4), using more complex concepts and more interactive questions.

[0310] Example 2: Intelligent Educational Robot

[0311] In this embodiment, the agent role switching method based on multimodal perception is implemented as an intelligent educational robot. The robot is equipped with a high-definition camera, a microphone array, a touch screen, and a robotic arm, which can provide richer interaction methods.

[0312] The multimodal perception module utilizes a microphone array to achieve sound source localization and speech enhancement, accurately recognizing children's speech even in noisy environments. The speech feature recognition unit uses the GFCC feature extraction algorithm, combined with beamforming technology, to improve the signal-to-noise ratio by 6.5dB, achieving a child speech recognition accuracy of 92.3% in noisy environments. The visual emotion recognition unit incorporates the robot's motion capabilities to implement a temporal attention mechanism that can actively adjust the viewing angle to obtain better facial expression data, improving emotion recognition accuracy to 91.2%. The touch intention analysis unit uses an adaptive threshold adjustment algorithm to not only analyze touchscreen operations but also identify physical interactions between children and the robot's mechanical arms through pressure sensors, achieving a touch intention recognition accuracy of 89.7%.

[0313] The dialogue state tracking module implements a child-optimized dialogue strategy and is enhanced for robot scenarios. It utilizes a low-threshold slot filling technique, with a slot certainty threshold set to 0.45—lower than tablet applications—to accommodate the more flexible dialogue mode of robot interactions. Furthermore, based on a topic-hopping adaptation strategy, the topic similarity threshold is lowered to 0.35, and up to four active topics are tracked simultaneously to accommodate children's more frequent topic switching in physical interaction environments. Emotion-sensitivity tracking has also been enhanced, achieving an emotional state recognition accuracy of 93.5% by integrating voice, facial expressions, and body movements.

[0314] The role switching module enables richer character expressions on the robot platform. It utilizes a multi-factor weighted scoring model with weights assigned to: explicit user request (0.45), change in conversation topic (0.25), change in emotional state (0.15), and change in learning progress (0.15). Compared to tablet applications, the weighting of emotional and learning progress factors is increased to better suit the robot's immersive interaction scenarios. Different roles differ not only in voice and content, but also in robot posture, LED expressions, and interaction methods. For example, the storyteller character speaks at a slower rate (0.8 times the standard speed), adopts a softer tone (lower pitch by 10%), displays a warm expression (orange, 60% brightness), and the robotic arm performs soothing gestures (small movements and slow speed). The partner character, on the other hand, speaks at a faster rate (1.2 times the standard speed), with a more lively tone (higher pitch by 15%), displays an excited expression (blue, 85% brightness), and the robotic arm performs more dynamic gestures (large movements and fast speed).

[0315] During the role switching process, a three-stage Markov transition strategy is employed, enhanced by the robot's multimodal output capabilities. For example, when switching from a scientific explorer to a game partner, the system not only provides a voice prompt during the pre-announcement phase, "This example has done so many experiments, are you a little tired? Do you want to play a game with your game partner, Xiao Ming?", but also gradually transitions the LED expression from a focused state (green) to an active state (blue), and the robotic arm's posture changes from an instructive posture to an interactive posture. During the transition phase, the system integrates the characteristics of the two characters: "Xiao Ming is here! This example can apply the scientific knowledge you just learned in the game!" while the LED expression and robotic arm's posture continue to transition. During the confirmation phase, the system fully switches to the game partner role: "I am Xiao Ming. Let this example use the mechanical principles you just learned to play a marble game!" The LED and robotic arm also fully transform to the characteristics of the game partner. Frequent switching detection is also implemented. If more than two role switch requests are detected within 25 seconds, an 8-12 second transition buffer is added to reduce cognitive burden.

[0316] When dealing with multiple roles’ concurrent responses, the robot’s multimodal output capability is fully utilized. Based on the conflict detection algorithm, a semantic similarity model based on RoBERTa is adopted (redundancy judgment threshold θ r =0.82) and logical consistency detection based on knowledge graph (contradiction judgment threshold θ c =0.22), a stricter threshold than tablet applications to accommodate the higher requirements of robot interaction. Based on the weight matrix construction method, weight coefficients were configured for educational game scenarios: base weight (λ1 = 0.15), contextual relevance (λ2 = 0.25), inter-role relationships (λ3 = 0.25), historical performance (λ4 = 0.15), and educational value (λ5 = 0.2). The weight of inter-role relationships was increased to promote multi-role collaboration.

[0317] When the difference between the highest and second-highest weights exceeds a threshold of δ = 0.2 (lower than the threshold for tablet applications), the character with the highest weight is selected for the response; otherwise, multiple character responses are synthesized. For example, when a child asks, "Why do objects fall?" the science explorer character might explain, "Objects fall because of gravity. Any object with mass attracts another." Meanwhile, the game partner character might say, "In this embodiment, we can play a fun game to see if objects of different weights fall to the ground at the same time!" If the weights are similar, not only will the voice synthesis response be "Objects fall because of gravity. In this embodiment, we can play a game to see if objects of different weights fall to the ground at the same time!" The system will also display a gravity-related diagram on the screen and use the robotic arm to perform corresponding demonstration movements, achieving parallel information transmission and multimodal output.

[0318] The role model library has been optimized for the robot platform, adding definitions for physical interaction modes. Based on the role knowledge graph, knowledge triples related to physical interactions have been expanded for robot scenarios, adding approximately 2,000 knowledge points specific to robot interaction. Character behavior models are parameterized, adding robot-specific behavioral parameter vectors $\vec{P} = (p_1, p_2, ..., p_q)$, including parameters such as robotic arm motion patterns, LED expression patterns, and spatial position adjustment. For example, the math teacher role uses the robotic arm to point to math problems on the screen (p1=0.8, a high proportion of indicative gestures) and guides children to answer through touch or gestures (p2=0.6, a medium level of interactive guidance). The science explorer role encourages children to pick up objects for observation (p1=0.4, a low proportion of indicative gestures) and uses the robot's camera for image analysis and explanation (p3=0.9, a high level of visual analysis).

[0319] The crowd portrait adaptation mechanism has been enhanced for robot interaction scenarios. The robot's physical interaction parameters are dynamically adjusted based on the child's age, cognitive level, learning style, and mobility. For example, for highly mobile children, the robot's range of motion and frequency of interaction are increased; for children with inattention, the frequency of LED expression changes and brightness contrast are increased to attract attention. The adaptation function is expanded to: $B_{\text{adapted}} = \text{Adapt}(B_{\text{role}}, \vec{C}, E)$, where $E$ represents environmental factors, including the size of the activity space, ambient noise level, and lighting conditions. This allows the system to adjust the robot's behavior based on the actual environment.

[0320] Example 3: Intelligent Early Childhood Education Robot (No Camera Version)

[0321] In this embodiment, the agent role switching method based on multimodal perception is implemented as an intelligent early childhood education robot for home use. The robot is not equipped with a camera and mainly interacts through a microphone array, touch screen, LED lights and speakers. It is particularly suitable for home environments with high requirements for privacy.

[0322] The multimodal perception module has been optimized and enhanced for use without visual input. The speech feature recognition unit utilizes the GFCC feature extraction algorithm, combined with an enhanced acoustic model, specifically optimized for the characteristics of children's speech. Parameter settings include: 24kHz sampling rate (higher than the standard version), 20ms frame length, 8ms frame shift, 0.98 pre-emphasis factor, 7500Hz high-frequency cutoff, 48 Mel filter banks, and 16th-order cepstral coefficients. Voiceprint recognition is also included, capable of distinguishing the voices of different family members, achieving a 90.8% accuracy rate for child speech recognition. The touch intent analysis unit utilizes an adaptive threshold adjustment algorithm and enhances pressure sensitivity and gesture recognition capabilities. It infers a child's emotional state through touchscreen operation patterns and pressure changes, partially compensating for the lack of visual input. The accuracy rate for touch intent recognition reaches 92.1%.

[0323] The dialogue state tracking module implements a child dialogue optimization strategy and is specifically enhanced for scenarios without visual input. Using child language understanding enhancement technology, the parameters of the semantic completion model are optimized for pure voice input, improving the ability to understand children's incomplete expressions. Furthermore, a low-threshold slot filling technology is implemented, with the slot certainty threshold set to 0.42, lower than the standard version, to accommodate dialogue understanding without visual assistance. Emotion-sensitive tracking has also been enhanced. Through the fusion analysis of speech prosodic features (pitch, volume, speaking rate, pauses) and touch behavior patterns, a dual-modal emotion recognition model has been constructed, achieving an emotional state recognition accuracy of 85.7%. While lower than the version with visual input, it is significantly higher than single-modal emotion recognition.

[0324] The role switching module is optimized for scenarios without visual interaction. It uses a multi-factor weighted scoring model with the following weights: explicit user request (0.55), change in conversation topic (0.25), emotional state transition (0.12), and change in learning progress (0.08). Compared to the standard version, the weight of explicit user request is increased, while the weights of emotional state and learning progress are reduced to accommodate lower emotion recognition accuracy. Different roles are primarily distinguished by voice characteristics and LED lighting effects. For example, the storyteller character speaks at a slower rate (0.85 times the standard speed), in a softer tone (pitch reduced by 8%), and displays a warm, pulsating LED light effect (orange, 55% brightness, 0.5Hz pulsation frequency). The partner character, on the other hand, speaks at a faster rate (1.15 times the standard speed), in a more lively tone (pitch increased by 12%), and displays a vibrant LED light effect (blue, 80% brightness, 1.2Hz pulsation frequency).

[0325] During the role switching process, a three-stage Markov transition strategy is employed, and sound and lighting effects enhance the transition experience. For example, when switching from a math teacher to a game partner, the system provides a voice prompt in the pre-announcement phase: "This embodiment has been studying math for a while. Do you want to take a break? Your game partner, Xiao Ming, has an interesting math game to share with you!" Simultaneously, the LED lighting gradually transitions from a steady green (65% brightness) to a vibrant blue (75% brightness). During the transition phase, the system integrates the voice features of the two characters: "Xiao Ming is here! This embodiment can use the math knowledge you just learned to play a game!" The speech speed and pitch are intermediate between the two characters, and the LED lighting continues to transition. During the confirmation phase, the system fully switches to the game partner's voice and lighting features: "Hi, I'm Xiao Ming! Let's play a number game with this embodiment!" Frequent switch detection is also implemented. If more than two role switch requests are detected within 20 seconds, a 10-15 second transition buffer is added, longer than the standard version, to ensure that children can clearly perceive the role change without visual feedback.

[0326] The conflict resolution module has been adjusted for scenarios without visual output. Based on the conflict detection algorithm, a lightweight semantic similarity model based on DistilBERT (redundancy determination threshold θr=0.88) and a simplified logical consistency detection (contradiction determination threshold θc=0.28) are used. The thresholds have been adjusted compared to the standard version to adapt to the characteristics of pure voice and light output. Based on the weight matrix construction method, the weight coefficient configuration is adopted in the family education scenario: basic weight (λ1=0.15), contextual relevance (λ2=0.35), role relationship (λ3=0.15), historical performance (λ4=0.15) and educational value (λ5=0.2). The weight of contextual relevance is increased to ensure the coherence of responses in the absence of visual assistance.

[0327] When the difference between the highest and second-highest weights exceeds a threshold of δ = 0.3 (higher than the threshold in the standard version), the response from the character with the highest weight is selected; otherwise, multiple character responses are synthesized. This increased threshold is intended to reduce character confusion that can occur in the absence of visual distinction. For example, when a child asks, "What is addition?" the math teacher character might explain, "Addition is combining two or more quantities to get a total," while the playmate character might say, "Addition is like collecting treasure! If you have two candies and get three more, you have five!" If the weight difference is significant, a single character response is selected. If the weights are similar, a speech synthesis response is used: "Addition is combining quantities. Just like collecting treasure, if you have two candies and get three more, you have five!" LED lighting changes (such as flashing when numbers appear) enhance the expressiveness and partially compensate for the lack of visual output.

[0328] The character model library has been optimized for the non-camera version. A character knowledge graph has been implemented, adding more knowledge triplets related to daily life and family interactions for family education scenarios, totaling approximately 1,500 family-specific knowledge points. Character behavior patterns are implemented using a parameterized approach, focusing primarily on the language style parameter vector $\vec{L}$ and the voice feature parameter vector $\vec{S} = (s_1, s_2, ..., s_k)$. $\vec{S}$ is specifically designed for the non-visual version and includes parameters such as timbre, rhythmic patterns, and emotional expression. For example, the voice feature parameters for the storyteller character are set to a warm timbre (s1=0.8), a rhythmic rhythm (s2=0.7), and rich emotional variation (s3=0.9); while the math teacher character's voice features are set to a clear timbre (s1=0.6), a well-paced rhythm (s2=0.4), and moderate emotional variation (s3=0.5).

[0329] The crowd portrait adaptation mechanism has been adjusted for scenarios without visual interaction. Voice and lighting parameters are dynamically adjusted based on the child's age, cognitive level, learning style, and auditory sensitivity. For example, for children with high auditory sensitivity, the volume is lowered, speech clarity is increased, and background sound effects are reduced. For children with attention deficit disorder, the rhythmic changes in speech and the frequency of light interaction are increased to maintain attention. The adaptation function is modified to: $B_{\text{adapted}} = \text{Adapt}(B_{\text{role}}, \vec{C}, A)$, where $A$ represents the auditory feature vector, including auditory sensitivity, pitch preference, and rhythm perception ability. This enables the system to optimize the interactive experience based on the child's auditory characteristics and compensate for the lack of visual interaction.

[0330] This embodiment is particularly suitable for home environments with high privacy requirements. By optimizing voice recognition, touch analysis, and lighting feedback, it can provide a personalized and intelligent children's educational experience without the use of a camera. Special enhancements have been made in voice interaction and emotional expression, demonstrating the adaptability and scalability of this embodiment's technical solution, enabling the implementation of core functions on different hardware configurations.

[0331] Figure 2 A schematic block diagram of an agent role switching system 200 based on multimodal perception provided by an embodiment of the present invention, the system 200 includes:

[0332] The information collection unit 201 is used to collect the user's multimodal perception information through the multimodal perception module;

[0333] a role determination unit 202 configured to determine a candidate role through an intelligent role switching decision mechanism based on the multimodal perception information;

[0334] A first setting unit 203 is configured to set the candidate role as a target role when the number of the candidate role is one;

[0335] A second setting unit 204 is configured to select one of the candidate roles as a target role through a multi-role conflict resolution mechanism when there are multiple candidate roles, or to combine the candidate roles into a target role;

[0336] The role switching unit 205 is used to switch the current role of the agent to the target role based on a preset role model library.

[0337] In one embodiment, the information collection unit 201 includes:

[0338] A speech extraction unit, configured to extract features from the speech information using a GFCC algorithm to obtain speech features;

[0339] A visual extraction unit, configured to capture the user's facial expression information from the visual information based on a temporal attention mechanism to obtain visual features;

[0340] a touch extraction unit, configured to perform touch intention recognition on the touch information by using an adaptive interference filtering algorithm to obtain touch features;

[0341] The multimodal fusion unit is used to perform multimodal fusion on the voice features, visual features and touch features to obtain the multimodal perception information.

[0342] In one embodiment, the multimodal fusion unit includes:

[0343] A projection mapping unit, configured to perform projection mapping on the speech features, visual features, and touch features, respectively, to obtain corresponding speech feature vectors, visual feature vectors, and touch feature vectors;

[0344] A feature fusion unit is used to fuse the speech feature vector, the visual feature vector and the touch feature vector based on a cross-modal attention mechanism to obtain the multimodal perception information.

[0345] In one embodiment, the role determination unit 202 includes:

[0346] An intent recognition unit, configured to perform contextual understanding of the multimodal perception information based on a multi-head self-attention mechanism and, in conjunction with a dialogue state tracker, perform intent recognition on the multimodal perception information to obtain intent information; wherein the intent information includes explicit requests, changes in dialogue topics, and changes in emotional states;

[0347] The weight selection unit is used to assign different weights to the intention information using a multi-factor weighted scoring model, and determine the corresponding candidate role based on the intention information after weight assignment.

[0348] In one embodiment, the second setting unit 204 includes:

[0349] The conflict detection unit is used to perform semantic similarity analysis and logical consistency analysis on multiple candidate roles using a conflict detection algorithm, and determine the redundancy and contradiction between the multiple candidate roles based on the results of the semantic similarity analysis and the results of the logical consistency analysis, and obtain the role conflict detection result:

[0350] A weight allocation unit is used to allocate weights to multiple candidate roles according to a preset dynamic weight allocation strategy and construct a weight matrix based on the weights. The dynamic weight allocation strategy includes: comprehensively calculating the response priority of each candidate role according to preset dimensions and allocating weights accordingly.

[0351] The role determination unit is used to determine the target role by combining the role conflict detection result and the weight matrix according to the following formula:

[0352] ;

[0353] in, Indicates the target role, represents the candidate role with the highest weight, represents the maximum weight value, represents the second highest weight value, represents the single role selection threshold, represents the synthetic role selection threshold, Synthesize represents the role synthesis function, Represents the input value of the character synthesis function.

[0354] In one embodiment, the role switching unit 205 includes:

[0355] The role transition unit is used to perform role switching using the three-stage Markov process of prediction-transition-confirmation according to the following formula:

[0356] P(r t |r t-1 ,μ t ,S t )=P(r t announce |r t-1 ,S t )·P(r t transition |r t announce ,μ t )·P(r t confirm |r t transition );

[0357] in, represents the role response vector in the preview phase, represents the role response vector in the transition phase, Represents the role response vector in the confirmation phase.

[0358] In one embodiment, the multimodal perception-based agent role switching system 200 further includes:

[0359] A portrait construction unit, configured to obtain a user feature vector and construct a user portrait based on the user feature vector;

[0360] A behavior setting unit, configured to set an adaptation function based on a preset character style, and set a character behavior mode through the adaptation function;

[0361] The model construction unit is used to construct a role model by combining the user portrait and the role behavior pattern, and to adaptively adjust the role model by using the crowd portrait adaptation mechanism, so as to construct a role model library containing multiple role models.

[0362] Since the embodiments of the system part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the system part, and will not be repeated here.

[0363] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When executed, the computer program can implement the steps provided in the above embodiment. The storage medium may include a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, or other medium capable of storing program code.

[0364] The present invention also provides an intelligent terminal that may include a memory and a processor. The memory stores a computer program, and when the processor calls the computer program in the memory, the steps provided in the above embodiments may be implemented. Of course, the intelligent terminal may also include various network interfaces, a power supply, and other components.

[0365] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.

[0366] It should also be noted that, in this specification, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.

Claims

1. A method for agent role switching based on multimodal perception, characterized in that: include: Collecting the user's multimodal perception information through the multimodal perception module; Based on the multimodal perception information, determining candidate roles through an intelligent role switching decision mechanism; When the number of the candidate roles is one, setting the candidate role as the target role; When there are multiple candidate roles, one of the candidate roles is selected as the target role through a multi-role conflict resolution mechanism, or the candidate roles are combined into the target role; Switch the current role of the agent to the target role based on the preset role model library; When there are multiple candidate roles, selecting one of the candidate roles as the target role through a multi-role conflict resolution mechanism, or synthesizing the candidate roles into the target role, includes: The conflict detection algorithm is used to perform semantic similarity analysis and logical consistency analysis on multiple candidate roles. The redundancy and contradiction between multiple candidate roles are determined based on the results of the semantic similarity analysis and the logical consistency analysis, and the role conflict detection results are obtained: Weights are assigned to multiple candidate roles according to a preset dynamic weight assignment strategy, and a weight matrix is ​​constructed based on the weights. The dynamic weight assignment strategy includes: comprehensively calculating the response priority of each candidate role based on preset dimensions and assigning weights accordingly. According to the following formula, the target role is determined by combining the role conflict detection results and the weight matrix: ; in, Indicates the target role, represents the candidate role with the highest weight, represents the maximum weight value, represents the second highest weight value, represents the single role selection threshold, represents the synthetic role selection threshold, represents the role synthesis function, Represents the input value of the character synthesis function.

2. The agent role switching method based on multimodal perception according to claim 1 is characterized in that: The multimodal perception information includes voice information, visual information and touch information; The collecting of the user's multimodal perception information by the multimodal perception module includes: Extracting features from the speech information using a GFCC algorithm to obtain speech features; Capturing the user's facial expression information from the visual information based on a temporal attention mechanism to obtain visual features; performing touch intention recognition on the touch information by using an adaptive interference filtering algorithm to obtain touch features; Multimodal fusion is performed on the voice features, visual features, and touch features to obtain the multimodal perception information.

3. The agent role switching method based on multimodal perception according to claim 2 is characterized in that: The multimodal fusion of the voice features, visual features, and touch features to obtain the multimodal perception information includes: Performing projection mapping on the speech features, visual features, and touch features respectively to obtain corresponding speech feature vectors, visual feature vectors, and touch feature vectors; The speech feature vector, the visual feature vector and the touch feature vector are fused based on a cross-modal attention mechanism to obtain the multimodal perception information.

4. The agent role switching method based on multimodal perception according to claim 1 is characterized in that: The determining of candidate roles through an intelligent role switching decision mechanism based on the multimodal perception information includes: The multimodal perception information is contextually understood based on a multi-head self-attention mechanism, and the multimodal perception information is used in conjunction with a dialogue state tracker to identify intent, thereby obtaining intent information. The intent information includes explicit requests, changes in dialogue topics, and changes in emotional states. A multi-factor weighted scoring model is used to assign different weights to the intention information, and the corresponding candidate roles are determined based on the intention information after weight assignment.

5. The agent role switching method based on multimodal perception according to claim 1 is characterized in that: The step of switching the current role of the agent to the target role based on the preset role model library includes: According to the following formula, the three-stage Markov process of prediction-transition-confirmation is used to switch roles: P(r t |r t-1 ,μ t ,S t )=P(r t announce |r t-1 ,S t )·P(r t transition |r t announce ,μ t )·P(r t confirm |r t transition ); in, represents the role response vector in the preview phase, represents the role response vector in the transition phase, Represents the role response vector in the confirmation phase.

6. The agent role switching method based on multimodal perception according to claim 1 is characterized in that: Also includes: Obtaining a user feature vector and constructing a user profile based on the user feature vector; Setting an adaptation function based on a preset character style, and setting a character behavior mode through the adaptation function; A role model is constructed by combining the user portrait and the role behavior pattern, and the role model is adaptively adjusted using a crowd portrait adaptation mechanism, thereby constructing a role model library containing multiple role models.

7. An agent role switching system based on multimodal perception, characterized in that: include: An information collection unit, configured to collect the user's multimodal perception information through a multimodal perception module; a role determination unit, configured to determine a candidate role through an intelligent role switching decision mechanism based on the multimodal perception information; a first setting unit, configured to set the candidate role as a target role when the number of the candidate role is one; A second setting unit is configured to select one of the candidate roles as a target role through a multi-role conflict resolution mechanism when there are multiple candidate roles, or to combine the candidate roles into a target role; A role switching unit, used to switch the current role of the agent to the target role based on a preset role model library; The second setting unit includes: The conflict detection unit is used to perform semantic similarity analysis and logical consistency analysis on multiple candidate roles using a conflict detection algorithm, and determine the redundancy and contradiction between the multiple candidate roles based on the results of the semantic similarity analysis and the results of the logical consistency analysis, and obtain the role conflict detection result: A weight allocation unit is used to allocate weights to multiple candidate roles according to a preset dynamic weight allocation strategy and construct a weight matrix based on the weights. The dynamic weight allocation strategy includes: comprehensively calculating the response priority of each candidate role according to preset dimensions and allocating weights accordingly. The role determination unit is used to determine the target role by combining the role conflict detection result and the weight matrix according to the following formula: ; in, Indicates the target role, represents the candidate role with the highest weight, represents the maximum weight value, represents the second highest weight value, represents the single role selection threshold, represents the synthetic role selection threshold, represents the role synthesis function, Represents the input value of the character synthesis function.

8. An intelligent terminal, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for switching intelligent agent roles based on multimodal perception as described in any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the agent role switching method based on multimodal perception as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Speech synthesis method and device, electronic equipment and storage medium

    CN116913245A

  • Model dialogue method and device based on style label, equipment and storage medium

    CN119961412A