Child interactive star key system for AI multi-mode everything identification

By using a multimodal perception and adaptive strategy decision-making module, combined with a visual encoder and speech recognition to generate intent vectors, the shortcomings of children's interactive systems in multimodal understanding and personalized adaptation are addressed. This enables dynamic matching of interactive content and emotional resonance, thereby improving the educational effectiveness of the interactive system.

CN121879575APending Publication Date: 2026-04-17SHENZHEN RONGYIN TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN RONGYIN TECHNOLOGY CO LTD
Filing Date
2026-01-05
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing interactive systems for children are insufficient in terms of multimodal understanding depth and personalized adaptability. They cannot accurately interpret the combination of visual scenes and voice commands, and the interactive content cannot dynamically evolve with the user's abilities and emotions, resulting in a lack of emotional resonance and educational effect in the interaction process.

Method used

A multimodal perception module is used to acquire image and audio streams. A visual encoder and speech recognition are used to generate intent vectors and real-time emotion state vectors. An adaptive strategy decision-making module is used to determine the cognitive age stage and match the optimal interaction strategy. Adaptive content is generated through a context orchestration and retrieval module. A large language model is used for security filtering and multimodal expression execution.

Benefits of technology

It improves the accuracy of interaction recognition, achieves dynamic matching of content difficulty, ensures that the output is synchronized with the user's emotions, and maintains the user's long-term interaction interest and educational effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121879575A_ABST
    Figure CN121879575A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and man-machine interaction, and discloses an AI multi-mode everything identification child interaction star key system, which comprises a multi-mode perception module, a data storage module, a self-adaptive strategy decision module, a situation arrangement and retrieval module, a content generation and safety control module, a multi-mode expression execution module and an asynchronous capability evaluation module. The system maps visual features to a semantic space through an attention mechanism to analyze a multi-modal anaphora intention, determines a cognitive age stage based on a user dynamic portrait matrix and matches an interaction strategy, generates reply content by using a large language model, and dynamically adjusts voice and action output parameters according to a real-time emotion vector. And finally updating the portrait by using a moving average algorithm. The problems that in the prior art, visual semantic alignment is lacked, and the content difficulty cannot follow dynamic evolution of the user ability are solved, and accurate intention understanding and personalized growth interaction are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and human-computer interaction technology, specifically to a children's interactive key system for AI multimodal object recognition. Background Technology

[0002] With the popularization of artificial intelligence technology, intelligent interactive devices have been widely used in children's education and companionship scenarios. Existing intelligent toys, early education robots, and smart speakers typically have basic voice recognition and dialogue capabilities, enabling them to interact with users through pre-built question-and-answer databases or large cloud models to complete tasks such as story playback and encyclopedic Q&A.

[0003] However, in practical applications, such systems still have limitations in terms of multimodal understanding depth and personalized adaptability. Traditional interactive systems mostly focus on single voice interaction or treat image recognition as an independent trigger function, lacking a mechanism for deep alignment of visual scene information with the speech semantic space. When users combine gestures or visual scenes to issue commands containing demonstrative pronouns during interaction, the system often fails to accurately interpret the referent in the speech, leading to misunderstandings of intent.

[0004] Furthermore, existing technologies typically employ generic processing models for content generation and interaction strategies, making it difficult to finely adapt to the specific cognitive development stages of children. System-generated responses often overlook users' language comprehension and logical thinking abilities, easily resulting in overly obscure vocabulary or content that is too simplistic for young children. Simultaneously, most devices lack dynamic perception of users' real-time emotional states and long-term tracking mechanisms for user development, leading to a lack of emotional resonance in the interaction process. They also fail to automatically adjust the difficulty of teaching or dialogue strategies as users' abilities improve, making it difficult to maintain long-term user engagement and educational effectiveness. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides an AI-powered multimodal object recognition system for children's interactive systems. This system solves the technical problems of existing children's interactive systems, which lack visual semantic alignment and cognitive age-appropriate adaptation mechanisms, resulting in the inability to accurately identify multimodal referential intentions and the inability of interactive content to dynamically evolve with the user's abilities and emotions.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solution: an AI multimodal object recognition child interactive star key system, comprising: a multimodal perception module, which acquires images and audio streams, and generates intent vectors and real-time emotion state vectors using a built-in visual encoder and speech recognition unit; The data storage module stores a dynamic user profile matrix representing the user's physical and mental characteristics, an age segmentation strategy, and a two-layer knowledge base. The adaptive strategy decision-making module determines the cognitive age stage based on the real-time emotional state vector and the user dynamic profile matrix, and matches the optimal interaction strategy accordingly. The context arrangement and retrieval module generates structured prompts based on the optimal interaction strategy and performs vector retrieval on the two-layer knowledge base; The content generation and security control module uses a large language model to generate response content based on structured prompts and performs security filtering. The multimodal expression execution module converts the generated response content into voice output and visual output; The asynchronous capability assessment module analyzes interaction data to quantitatively assess user capabilities and update the user dynamic profile matrix.

[0007] Preferably, the user dynamic profile matrix includes physiological age, interest feature vector, five-dimensional ability score vector, real-time emotional state vector, dynamic personality feature vector, and short-term dialogue context vector. The number of dimensions of the interest feature vector is equal to the total number of entity categories in the pre-set knowledge graph, and the value of each dimension represents the user's attention weight to that entity category; the five-dimensional ability scoring vector contains normalized values ​​corresponding to the five dimensions of language expression, logical cognition, emotion management, social interaction and self-awareness.

[0008] Preferably, when generating the intent vector, the multimodal perception module performs the following operations: High-dimensional visual features of images are extracted using a visual encoder to identify a set of objects in a scene and generate visual semantic description text based on the relative positional relationships between the objects. The visual semantic description text is concatenated with the user input text sequence obtained through speech transcription; The concatenated text is input into a natural language processing encoder, and the visual scene information is mapped to the semantic space through an attention mechanism, thereby parsing out the intent vector containing visual referential information.

[0009] Preferably, when determining the cognitive age stage, the adaptive strategy decision-making module performs the following operations: Extract language complexity features, cognitive pattern features, and attention features from the data output by the multimodal perception module; The language complexity feature is calculated based on the sentence length and vocabulary ratio of the input text; the cognitive pattern feature is calculated based on the density of logical connectives and the frequency of abstract nouns. The attention features are calculated based on gaze duration and response delay; The language complexity features, cognitive pattern features, and attention features are weighted, summed, and normalized to obtain a probability distribution vector. The dimension index with the largest probability value in the probability distribution vector is selected as the current user's cognitive age stage.

[0010] Preferably, when the adaptive strategy decision-making module retrieves and matches the optimal interaction strategy, it calculates the score of each strategy in the candidate strategy set based on the objective function. The objective function is configured as follows: calculate the similarity between the strategy attribute vector and the current context vector as the context matching degree, and calculate the similarity between the strategy attribute vector and the interest feature vector and dynamic personality feature vector in the user dynamic profile matrix as the personalization fit degree. The context matching degree and the personalized fit degree are weighted and summed to obtain the final score of the strategy, wherein the weighting coefficient is dynamically adjusted according to the entropy value of the real-time emotion state vector.

[0011] Preferably, after performing vector retrieval on the two-layer knowledge base, the context arrangement and retrieval module also performs reordering and filtering operations based on cognitive adaptability. This operation specifically includes: obtaining the target difficulty value obtained by mapping the cognitive age stage; Calculate the comprehensive matching score for each candidate knowledge fragment. The comprehensive matching score is obtained by calculating the similarity between the query vector and the knowledge fragment vector, and subtracting the absolute value of the difference between the weighted calculation of the difficulty level of the knowledge fragment and the target difficulty value. Candidate knowledge fragments are sorted according to the comprehensive matching score, and the fragments with the highest ranking are selected as related knowledge fragments.

[0012] Preferably, the structured prompts generated by the context arrangement and retrieval module include a role setting field, a context information field, a pedagogical constraint field, and a user input field; The role setting field is determined by the prompt word template determined by the optimal interaction strategy; the context information field is composed of short-term dialogue context vectors, user history interaction fragments, and related knowledge fragments; the pedagogical constraint field contains instruction text used to restrict the output sentence structure and rhetorical devices of the large language model.

[0013] Preferably, the content generation and security control module and the multimodal expression execution module perform dynamic parameter adjustments based on real-time emotion state vectors when generating output: The content generation and security control module maps the probability values ​​of specific emotion tags in the real-time emotion state vector to the sampling temperature parameter of the large language model in order to control the randomness of the generated text. When performing speech synthesis, the multimodal expression execution module multiplies the baseline pitch parameter in the age-based strategy by a correction coefficient based on the cognitive age stage, and dynamically adjusts the baseline speech rate parameter according to the probability difference between positive and negative emotion labels in the real-time emotion state vector.

[0014] Preferably, when the multimodal expression execution module (500) converts the generated response content into visual output, it performs an audiovisual collaborative driving operation, specifically including: The acoustic features of the speech output are extracted and phoneme decoded to obtain a phoneme sequence. Based on the pre-configured phoneme-visual mapping relationship, the phoneme sequence is converted into a mouth shape visual weight sequence aligned with the speech time axis. The lip shape pixel weight sequence is linearly weighted and fused with the facial expression hybrid deformation weight set generated based on the real-time emotion state vector to obtain target driving frame data. If the output terminal is a 3D virtual image, the target driving frame data is used to drive the rendering engine to perform real-time displacement updates on the mesh vertex coordinates of the pre-configured 3D model to achieve facial animation rendering. If the output terminal is a physical robot component, the target drive frame data is mapped to the target control angle of the joint servo motor of the physical robot component through a pre-configured degree of freedom mapping matrix.

[0015] Preferably, the asynchronous capability assessment module uses a moving average algorithm to update capability features when updating the user dynamic profile matrix; Specifically, this involves reading the historical capability feature vector stored in the data storage module and combining it with the comprehensive cognitive performance score vector obtained from this interaction. The updated ability feature vector is obtained by weighted summing of the historical ability feature vector and the comprehensive cognitive performance feature vector using a memory forgetting factor; the value of the memory forgetting factor is negatively correlated with the time interval between the user's last interaction.

[0016] This invention provides an AI-powered multimodal object recognition system for children's interactive key systems. It offers the following advantages: 1. This invention uses a multimodal perception module and an attention mechanism to map scene features extracted by the visual encoder to the semantic space and concatenate them with the speech text sequence. This technique enables the system to parse intent vectors containing visual pronouns such as "look at this" and "what is that," overcoming the shortcomings of traditional single-modal voice interaction in understanding user commands by combining visual scenes, thereby improving the accuracy of interactive recognition in complex audiovisual environments.

[0017] 2. This invention utilizes an adaptive strategy decision-making module to extract the user's language complexity and attention characteristics, calculates and determines the cognitive age stage, and calculates the absolute value of the difference between the difficulty of the knowledge segment and the target difficulty in the context arrangement stage, using this as the basis for retrieval and ranking. This logic ensures that the knowledge base content called by the system will not cause users to have difficulty understanding due to excessive difficulty, nor will it cause users to lose interest due to excessive difficulty, thereby realizing the dynamic matching of the difficulty of knowledge supply with the user's cognitive ability.

[0018] 3. At the execution end, this invention maps the real-time emotional state vector to the sampling temperature of the large language model and the pitch and speech rate parameters of the speech synthesis, so that the output speech and actions can respond to the user's current emotional fluctuations. At the same time, the asynchronous ability assessment module uses a moving average algorithm that introduces a memory forgetting factor to update the user's dynamic profile, so that the system's evaluation of the user's ability can objectively reflect the memory decay or ability solidification caused by the passage of time, ensuring that the data source on which the subsequent interaction strategy is based has timeliness and accuracy. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the overall system architecture of the present invention; Figure 2 This is a schematic diagram of the end-to-end adaptive interaction process of the present invention.

[0020] Among them, 100 is the multimodal perception module; 200 is the context orchestration and retrieval module; 300 is the adaptive strategy decision-making module; 400 is the content generation and security control module; 500 is the multimodal expression execution module; 600 is the asynchronous capability assessment module; and 700 is the data storage module. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] See attached document Figures 1-2 This invention provides a children's all-age adaptive interaction and ability assessment system based on a multimodal large model and dynamic knowledge graph. It mainly includes a multimodal perception module 100, a context arrangement and retrieval module 200, an adaptive strategy decision-making module 300, a content generation and security control module 400, a multimodal expression execution module 500, an asynchronous ability assessment module 600, and a data storage module 700. These modules are connected via a data bus or network communication protocol to achieve real-time data transmission and command coordination.

[0023] The multimodal perception module 100 is configured as the system's input terminal, used to collect and process the user's multimodal interaction data. This multimodal perception module 100 is connected to an image acquisition device and an audio acquisition device, used to acquire real-time image stream input and audio stream input, respectively. The multimodal perception module 100 integrates a visual semantic encoder and an automatic speech recognition unit, used to convert unstructured video and audio signals into structured feature vectors that can be processed by a computer. Specifically, the multimodal perception module 100 extracts features from the input data, generating a text sequence containing the user's intent and an emotional state vector representing the user's current psychological state. This emotional state vector is calculated by fusing facial expression features and speech prosody features, and is used for subsequent strategy adjustments.

[0024] The data storage module 700 is used to maintain the system's core data assets, including a user dynamic profile library, an age-based strategy configuration library, and a two-layer knowledge base. The user dynamic profile library stores the state space model of each user, which is defined as a user dynamic profile matrix. User dynamic profile matrix Including physiological age Interest feature vector Five-dimensional ability scoring vector Real-time emotional state vector Dynamic personality trait vector and short-term dialogue context vectors The age-segmentation strategy configuration library stores preset interaction parameters for different age groups, including prompt word templates, speech synthesis parameters, and interaction waiting thresholds. The two-layer knowledge base includes a memory vector library that records users' historical interactions and an authoritative knowledge base containing educational psychology content.

[0025] The adaptive strategy decision module 300 is the core of the system's decision-making process. It is used to base decisions on the features output by the multimodal perception module 100 and the user dynamic profile matrix in the data storage module 700. The adaptive strategy decision module 300 dynamically determines the current interaction strategy. First, based on the user's language complexity and interaction behavior characteristics, it uses a classification algorithm to determine the user's cognitive age stage. Then, combining the current context and the user's personalized characteristics, the adaptive strategy decision module 300 retrieves the optimal interaction strategy from the age-appropriate strategy configuration library. This interaction strategy determines the system's response language style, guidance method, and multimodal output parameter configuration.

[0026] The context orchestration and retrieval module 200 connects the multimodal perception module 100 and the data storage module 700 to execute the retrieval enhancement generation process. The context orchestration and retrieval module 200 uses user-inputted text content and visual semantic tags as query conditions and performs vector retrieval on the memory vector library and authoritative knowledge base in the data storage module 700 to obtain the Top-K relevant knowledge fragments. The context orchestration and retrieval module 200 assembles the retrieved knowledge fragments, the current dialogue context, and the interaction strategy parameters determined by the adaptive strategy decision module 300 to construct structured prompts that include role settings, contextual information, and pedagogical constraints.

[0027] The content generation and security control module 400 connects to the context arrangement and retrieval module 200, and is used to generate response content based on a large language model and perform multi-level guardrail filtering. The content generation and security control module 400 receives structured prompts and uses the large language model to generate initial response text. The content generation and security control module 400 has a built-in rewriting model with fine-tuned instructions. When it identifies that the user's age is at a specific logical developmental stage and the initial response is a declarative sentence structure, the rewriting model transforms the initial response into a heuristic question structure to conform to the guided teaching principles in educational psychology. In addition, the content generation and security control module 400 also includes a semantic classification-based security filtering unit to block generated content containing inappropriate information.

[0028] The multimodal expression execution module 500, connected to the content generation and security control module 400, is used to convert the final generated response content into audiovisual output suitable for the user's age. The multimodal expression execution module 500 receives text responses and strategy parameters, generates speech output through speech synthesis technology, and dynamically adjusts the speech rate, tone, and prosodic features according to the strategy parameters. Simultaneously, the multimodal expression execution module 500 generates static images or dynamic visual content to aid understanding, based on the semantic complexity of the response content.

[0029] The asynchronous ability assessment module 600 operates independently of the main interaction process described above, and is used to analyze and quantify the interaction data in the background. The asynchronous ability assessment module 600 acquires complete interaction records containing multiple rounds of dialogue, and uses thought chain reasoning technology to analyze the user's logical thinking trajectory, emotional change curve, and language expression ability. Based on the analysis results, the asynchronous ability assessment module 600 calculates scores for each ability and updates the five-dimensional ability score vector in the data storage module 700 using an exponentially weighted moving average algorithm. The update process follows the mathematical relationship described below: ; in, For the updated capability vector, This is the currently stored capability vector. The ability score observed in this interaction is given. The forgetting factor ranges from 0 to 1. This mechanism ensures the dynamic and timely nature of user ability assessment, enabling the system's adaptive evolution.

[0030] See attached document Figure 1 In one embodiment of the present invention, the data storage module 700 is used to store and manage the user dynamic profile matrix and the age-based strategy configuration table. The underlying architecture of the data storage module 700 can adopt a hybrid storage architecture combining vector databases, relational databases, and graph databases.

[0031] To achieve accurate quantification and tracking of user status, the system constructs a dynamic user profile matrix. This matrix is ​​a time step... A dynamically updated set of multidimensional states is used to comprehensively describe the user's physiological and psychological characteristics at the current moment. Specifically, this is the user dynamic profile matrix. In mathematics, it is defined as an ordered set containing six key components: ; The specific technical definitions of each component are as follows: Quantity Representing a user's physiological age, measured in months, it forms the basis for the initial age segmentation strategy of the system.

[0032] Quantity The interest feature vector is a high-dimensional sparse vector. The system uses entity linking technology to map specific things mentioned in the user's historical interactions to pre-defined knowledge graph entity nodes. The number of dimensions of this vector is equal to the total number of pre-defined entity categories in the knowledge graph. Each dimension in the vector corresponds to an entity category or specific entity in the knowledge graph, and the value on that dimension represents the user's attention weight to that entity, with a value ranging from 0 to 1 in a closed interval.

[0033] Quantity This represents a five-dimensional capability scoring vector, used to quantify various development indicators of a user. This vector... It contains five floating-point values, corresponding to language expression ability, logical cognition ability, emotional management ability, social interaction ability, and self-awareness ability, respectively. The values ​​of these five dimensions are all normalized to a closed interval between 0 and 1 to represent the differences in the user's development level in different dimensions, so as to support the system to provide targeted guidance for weaknesses.

[0034] Quantity This represents a real-time emotional state vector, which is a probability distribution vector. The number of its dimensions corresponds to the number of preset emotional labels, including labels such as excitement, curiosity, anxiety, and frustration. Each element in the vector represents the confidence probability of the user currently being in that emotional state, and the sum of all elements is 1. This vector is calculated by feature fusion from the multimodal perception module and serves as an important input for query context triggering rules.

[0035] Quantity This represents a dynamic personality trait vector, used to describe a user's relatively stable psychological and behavioral characteristics. The vector includes four sub-dimensions: activity level, focus level, social tendency, and exploratory drive. Unlike real-time emotional state vectors, dynamic personality trait vectors have a time-cumulative characteristic; their values ​​are weighted and updated based on the statistical patterns of long-term interaction data, used to characterize a user's personalized behavioral patterns.

[0036] Quantity Represents the short-term dialogue context vector, obtained by analyzing the most recent... The dialogue text is obtained by embedding encoding, which is used to preserve the continuous semantic information of the interaction and provide semantic input for the system to generate coherent responses.

[0037] In addition to the user dynamic profile matrix, the data storage module 700 is also configured with an age-based strategy configuration table. This configuration table is a hierarchical parameter mapping structure that stores key-value pairs indexed by age stage and interaction control parameters. This configuration table contains at least three levels of preset strategies: The first-level strategy is designed for users aged 1 to 3 years in the sensorimotor stage. Its configuration parameters include: the speech synthesis speed coefficient is set to 0.8 times the standard speech speed, the pitch base frequency is increased by 20%, the interaction waiting threshold is set to 5 seconds, and the constraint condition for generating prompt words is to focus on the use of onomatopoeia and repetition of simple concepts.

[0038] The second-level strategy is designed for pre-computational users aged 3 to 6. Its configuration parameters include: standard speech rate and pitch for speech synthesis, an interaction waiting threshold of 4 seconds, and prompt constraints that focus on complete sentence expression and causal logic guidance.

[0039] The third-level strategy is designed for users aged 6 to 8 who are in the specific computational stage. Its configuration parameters include: speech synthesis speed close to natural adult conversation, interaction waiting threshold set to 3 seconds, and prompt word constraints focusing on the construction of rhetorical questions and complex logical deduction.

[0040] The data storage module 700 also stores a context-triggered rule base, which defines the mapping logic from specific combinations of user states to system response patterns. This mapping logic can be represented as a conditional lookup table or a decision tree structure, including emotional context triggers, cognitive context triggers, social context triggers, and exploratory context triggers.

[0041] In specific implementation, the emotional situation trigger is set as follows: when the real-time emotional state vector... When the dimension value of the corresponding frustration label exceeds the preset threshold, the system forcibly switches to the encouragement strategy, calling the preset soothing language data and soft voice parameters.

[0042] The cognitive context trigger is set as follows: when the system detects that the user's interaction response delay exceeds the waiting threshold for the current age group, or when there are multiple instances of input with unclear intent, it is determined to be cognitive overload, and the system automatically reduces the complexity level of content generation.

[0043] The social context trigger is set as follows: when the intent recognition result indicates that the user shows a willingness to share or inquire, the system activates the collaborative interaction mode to guide the user to conduct multiple rounds of interactive communication.

[0044] Explore the context trigger settings as follows: when the real-time emotion state vector When the dimension value of the corresponding curiosity tag exceeds the preset high threshold, the system calls the deep exploration content interface to provide extended information on related knowledge.

[0045] The aforementioned user dynamic profile matrix, age-based strategy configuration table, and context-triggered rule base together constitute the system's decision data foundation, supporting subsequent modules in implementing data-driven adaptive interactive control. The specific storage format, index construction method, and database read / write optimization techniques for each of the aforementioned vectors can be implemented using existing data storage and retrieval schemes, and will not be elaborated upon here.

[0046] See attached document Figure 2 In one embodiment of the present invention, the multimodal perception and feature state space construction in step S1 can be specifically implemented by the multimodal perception module 100 performing the following sub-steps: S101, Synchronous acquisition and preprocessing of multimodal data streams.

[0047] The multimodal perception module 100 connects to the camera device and microphone array via a data acquisition interface, and acquires video stream data containing the user's face and body posture, and audio stream data containing the user's voice. To ensure the temporal alignment of subsequent feature fusion, the multimodal perception module 100 uses a time synchronization algorithm to synchronously mark the video stream frame sequence and the audio stream sampling segments, establishing a unified time reference.

[0048] For video stream data, noise reduction, keyframe extraction, and normalization are performed; for audio stream data, echo cancellation and pre-emphasis processing are performed. These basic image and audio signal processing methods can be implemented by those skilled in the art using existing digital signal processing techniques, and are well-known in the field, so they will not be elaborated upon here.

[0049] S102, Visual Semantic Understanding and Scene Graph Construction.

[0050] The multimodal perception module 100 inputs the processed image frames into the visual encoder. This visual encoder is configured as a trained convolutional neural network model to extract high-dimensional visual features from the image. Based on these high-dimensional visual features, an object detection task is performed to identify a set of objects present in the current scene. This set of objects contains several detected objects, and the system records the semantic label text, the positional coordinates in the image, and the detection confidence score for each detected object.

[0051] Furthermore, the multimodal perception module 100 constructs a scene graph based on the detected set of objects and the relative positional relationships between them. Subsequently, the system converts the spatial relationships of each object node and its connecting edges in the scene graph into natural language text descriptions according to preset natural language templates. For example, when the coordinate overlap rate of two objects meets preset conditions, text describing the contact relationship between the two objects is generated.

[0052] S103, speech transcription and multimodal intent recognition.

[0053] The multimodal perception module 100 uses an automatic speech recognition engine to convert audio stream data into a sequence of user-input text. Subsequently, the intent recognition unit will process the user's input text sequence. The text sequence is then fused with the visual semantic description text generated in step S102 at the feature level or decision level. Specifically, the module concatenates the text sequence with the visual semantic description text and inputs it into the natural language processing encoder. Through the encoder's attention mechanism, the visual scene information is mapped into the semantic space, thereby parsing out the intent vector containing visual referential information. .

[0054] S104, Multimodal Emotion Feature Fusion and State Vector Generation.

[0055] The multimodal perception module 100 extracts emotional features within the same time window from both the audio stream and the image stream. First, it analyzes the prosodic features of the audio stream using a speech emotion recognition model. These prosodic features include at least the fundamental frequency, intensity, and speech rate, generating an audio emotion feature vector. .

[0056] Simultaneously, facial expression recognition models are used to analyze the geometric changes and micro-expression features of key facial points in the image stream, generating visual emotion feature vectors. .

[0057] The system performs a weighted fusion operation on the feature vectors of the two modalities. Specifically, a fully connected network layer maps the audio and visual emotion feature vectors to a unified emotion space. After adding a bias term, a normalized exponential function is used to process the output, resulting in a real-time emotion state vector in the form of a probability distribution. The values ​​of each dimension in this vector represent the user's confidence level in the corresponding preset emotion label.

[0058] S105, User state space update.

[0059] The multimodal perception module 100 calculates the intent vector obtained from the above steps. and real-time emotional state vector The existing user profile data in the data storage module is combined into user state space data for the current time step, and the user state space data is transmitted to the adaptive strategy decision module 300.

[0060] See attached document Figure 2 In one embodiment of the present invention, the age-based profile-based strategy decision-making and retrieval enhancement in step S2 can be specifically implemented by the adaptive strategy decision-making module 300 executing the following sub-steps: S201, Multidimensional Feature Extraction and Cognitive Age Stage Recognition. The adaptive policy decision module 300 receives user state space data output by the multimodal perception module 100 and extracts three feature vectors representing the user's cognitive development level: language complexity feature. Cognitive pattern characteristics and attention characteristics .

[0061] Specifically, language complexity features It is a numerical vector whose dimensions include the average sentence length of the user input text sequence, the proportion of low-frequency words, and the maximum depth of the syntax tree.

[0062] Cognitive pattern characteristics It is a numerical vector that is calculated by statistically analyzing the density of logical connectors (including causal and adversative conjunctions) and the frequency of use of abstract nouns in the user's historical N rounds of dialogue. It is used to quantify the user's logical thinking and abstract generalization ability.

[0063] attention characteristics It is a numerical vector whose dimensions include the average duration of the user's gaze during the interaction and the reciprocal of the response delay time of the interaction command.

[0064] Based on the three feature vectors mentioned above, the adaptive strategy decision module 300 uses a pre-set classification algorithm to calculate the user's current cognitive age stage. Specifically, the module assigns preset weight matrices to the language complexity feature, cognitive pattern feature, and attention feature, calculates the weighted sum, and superimposes the bias vector. Finally, it outputs a probability distribution vector through a normalization function. The module selects the dimension index with the highest probability value in this probability distribution vector as the current user's determined cognitive age stage. .

[0065] S202, Construction of candidate policy set and context parsing.

[0066] The adaptive strategy decision-making module 300 determines the cognitive age stage. The corresponding candidate strategy set is retrieved from the age-based strategy configuration table of the data storage module 700. The set of candidate strategies It contains several predefined interaction strategies, each of which... Each is associated with a strategy attribute vector, which describes the preset values ​​of the strategy in three dimensions: emotional support, knowledge density, and guidance intensity.

[0067] At the same time, the module combines the current dialogue history and real-time emotion state vectors. Generate context vectors Context vectors It is achieved by using real-time emotion state vectors Intent vector of the current dialogue The result is obtained after weighted summation and normalization.

[0068] S203, Optimal policy calculation based on the objective function. The adaptive policy decision module 300 calculates the candidate policy set. Each strategy The strategy with the highest score is selected as the optimal strategy. The calculation process is based on the following objective function: ; in: This represents the set of candidate strategies determined by the preceding steps; Indicates in set The current strategy variable during traversal calculation; The operator representing the maximum value is used to extract values ​​from a set. Choose the strategy that maximizes the value of the weighted sum function within the parentheses. ; Configure the context matching function as the computation strategy. 2. Policy attribute vector and context vector Cosine similarity between them; Configure the personalized adaptation function as a calculation strategy. Strategy attribute vectors and user dynamic profile matrix Interest feature vectors in and dynamic personality trait vector Weighted cosine similarity; This is the adjustment coefficient for context matching weights, with a value ranging from 0 to 1; The adjustment coefficient for personalized adaptation weights, this coefficient is related to The sum of these values ​​is 1, which is used to form a complementary weighted balance mechanism between contextual matching and personalized adaptation.

[0069] The adaptive strategy decision-making module 300 bases its decisions on real-time emotion state vectors. Dynamic adjustment of entropy value The size of the entropy value is increased when the entropy value is below a preset threshold, indicating that the emotion is clear and strong. The value is used to increase the weight of context matching.

[0070] S204, Strategy parameter fine-tuning and retrieval instruction generation.

[0071] Determining the optimal strategy Then, the adaptive strategy decision-making module 300 determines the appropriate strategy based on the real-time emotion state vector. The specific parameters of the strategy are numerically adjusted. Specifically, the optimal strategy... The preset base speech rate parameter is multiplied by a speech rate correction coefficient, which is then compared with the real-time emotion state vector. The probability values ​​of the corresponding anxiety labels are negatively correlated.

[0072] The revised strategy parameters include: prompt word template ID, pedagogical constraints, speech synthesis parameters, and visual output complexity level.

[0073] The module integrates the aforementioned strategy parameters with the user-input text content and visual semantic tags to generate a search instruction. This search instruction includes a filtering field to restrict the search to only return results related to determining cognitive age stage. The knowledge base content matches the difficulty level.

[0074] See attached document Figure 2In one embodiment of the present invention, the dynamic context assembly and structured prompt word construction in step S3 can be specifically implemented by the context arrangement and retrieval module 200 by performing the following sub-steps: S301, Search command parsing and dual-channel parallel retrieval.

[0075] The context orchestration and retrieval module 200 receives a retrieval instruction from the adaptive strategy decision module 300. This instruction includes strategy-weighted query text, visual semantic tags, and filtering condition fields. The module uses a pre-built vector embedding model to encode the features of the query text and visual semantic tags, generating a high-dimensional query vector. .

[0076] Subsequently, the module is based on the query vector Parallel retrieval is performed on the memory vector library and authoritative knowledge base in the data storage module 700. For the memory vector library, the module retrieves the user's historical interaction fragments that are most relevant to the current intent, and denoted as the user's historical interaction fragment set. .

[0077] For authoritative knowledge bases, the module retrieves candidate knowledge fragments containing educational knowledge points or factual information, which are denoted as the candidate knowledge fragment set. Each entry in the authoritative knowledge base is pre-associated with a difficulty level label, which indicates the cognitive age group to which the knowledge point is applicable.

[0078] The above retrieval process uses cosine similarity as the basic metric. For indexing and nearest neighbor search techniques for large-scale vector data, those skilled in the art can use existing approximate nearest neighbor search algorithms, which are well-known techniques in the field and will not be elaborated here.

[0079] S302, Knowledge reordering and filtering based on cognitive adaptability.

[0080] To ensure that the retrieved knowledge content is suitable for the user's current cognitive development level, the module analyzes the candidate knowledge fragment set obtained in S301. Perform reordering and filtering operations. This operation is based on the filtering criteria field in the search command, which specifies the current user's cognitive age stage. Corresponding target difficulty value .

[0081] Module calculates the set of candidate knowledge fragments Each knowledge fragment Overall matching score The calculation follows the formula: ; in: Represents the set of candidate knowledge fragments Any fragment of knowledge in it; The query vector generated by the preceding steps; For knowledge fragments The corresponding vector representation; The cosine similarity between the query vector and the knowledge fragment vector is used to measure the relevance of the content. For knowledge fragments The pre-marked difficulty level value; To determine the cognitive age stage The target difficulty value obtained through mapping; This represents the absolute difference between the difficulty of a knowledge segment and the target difficulty, and is used to penalize age-inappropriate content. and These are the relevance weighting coefficient and the difficulty penalty coefficient, respectively, and satisfy the following conditions: .

[0082] The module is based on the overall matching score. Sort the candidate knowledge fragments in descending order and extract the first few fragments. These fragments form the final set of related knowledge fragments. .

[0083] S303, Dynamic assembly of structured prompts.

[0084] The context arrangement and retrieval module 200 assembles the above retrieval results with the constraints of the current interaction based on the predefined prompt word template, and constructs structured prompt words for input into the large language model. The structured prompt is configured to contain at least four logical data fields: a role setting field, a contextual information field, a pedagogical constraint field, and a user input field.

[0085] The specific definitions and content population logic for each data field are as follows: The fields are set for the role, and the content is determined by the prompt word template ID determined by the adaptive strategy decision module 300, such as setting it as a patient kindergarten teacher or a rigorous science guide.

[0086] For the contextual information field, the module extracts short-term dialogue context vectors from the user dynamic profile matrix. The corresponding original text list, and the set of user history interaction fragments obtained in step S301. and the set of related knowledge fragments obtained in step S302 The text is then pieced together to form a text block containing background information.

[0087] This is the pedagogical constraint field, populated with content derived from the pedagogical constraints generated by the adaptive strategy decision module 300. This field contains instruction text to restrict the model's output format, such as prohibiting complex sentences exceeding 15 characters, requiring the use of personification, or using Socratic rhetoric rather than direct correction for incorrect user responses.

[0088] The user input field is filled with the original user text input and visual semantic description text for this interaction.

[0089] S304, Prompt word optimization and token truncation handling.

[0090] In generating structured prompt words Then, the module calculates its total token count. If the token count exceeds the context window limit of the large language model, the module executes a priority-based truncation strategy. This strategy prioritizes preserving... and Secondly, retain Finally, for The code is trimmed based on time proximity until the length constraint is met. The final structured prompt is then transmitted to the content generation and security control module 400.

[0091] See attached document Figure 2 In one embodiment of the present invention, the multimodal content generation and security closed-loop control in step S4 can be implemented by the content generation and security control module 400 executing the following sub-steps: S401, content generation based on dynamic decoding strategy.

[0092] The content generation and security control module 400 receives structured prompts transmitted by the context arrangement and retrieval module 200. This data is then input into a pre-built generative large language model. The module is configured to use the real-time emotion state vector output by the adaptive policy decision module 300. The model's decoding hyperparameters are calculated using a mapping function to control the randomness of the generated content.

[0093] Specifically, the module calculates the sampling temperature parameters. This calculation will use real-time emotion state vectors. The probability value of the corresponding curiosity tag is mapped to a preset temperature range [0.1, 1.0]. The higher the probability value, the higher the generated temperature range. The larger the value, the better. The generation process utilizes the above parameters, combined with a kernel sampling strategy, to output an initial response text sequence from a large language model. .

[0094] S402, dual security verification and risk blocking.

[0095] The content generation and security control module has 400 pairs of initial response text sequences. Perform dual security checks.

[0096] The first layer of verification is sensitive word filtering based on deterministic automata. The module utilizes the Aho-Corasick multi-pattern matching algorithm to... The code compares the word against a pre-set blacklist of keywords. If a sensitive word is found, the code will be blocked.

[0097] The second layer of verification is a semantic-based risk assessment. The module will... Input is fed into a pre-trained text classification model, and the output is a risk probability vector. This vector contains probability values ​​corresponding to four dimensions: violence, pornography, hate speech, and harmful inducement. The module weights this risk probability vector with a preset risk weight vector and takes the maximum value as the comprehensive risk score. .

[0098] like Exceeding the preset security threshold Module discard It retrieves a pre-set security response matching the current intent from the security corpus of the data storage module 700 as the final text. If the threshold is not exceeded, then let .

[0099] S403, multimodal speech synthesis and virtual avatar driven.

[0100] Content generation and security control module 400 based on the final text Based on the current user's age-based strategy parameters, voice stream data and virtual avatar-driven data are generated.

[0101] In the speech synthesis stage, the module utilizes an end-to-end speech synthesis model to generate waveform data. During synthesis, the module applies prosodic control parameters based on user state. Specifically, the system multiplies the baseline pitch parameter in the age-appropriate strategy by a correction coefficient based on cognitive age stage, while dynamically adjusting the baseline speech rate parameter according to the probability difference between positive and negative emotion labels in the real-time emotion state vector. These adjusted parameters are input into the speech synthesis model to generate an audio stream with specific emotional coloring and age adaptability. .

[0102] Meanwhile, the module extracts phoneme sequences and timestamps from the audio stream using a phoneme alignment algorithm, maps them onto the pre-set hybrid deformation weights of the 3D virtual image, generates lip-sync data and facial expression frame sequences synchronized with the speech, and completes the construction of multimodal output data.

[0103] See attached document Figure 1 In one embodiment of the present invention, the multimodal expression execution and feedback in step S5 can be specifically implemented by the multimodal expression execution module 500 performing the following sub-steps: S501, buffering and clock synchronization for multimodal data streams.

[0104] The multimodal expression execution module 500 receives audio stream data from the content generation and security control module 400. The module generates lip-sync data and facial expression frame sequences. It constructs two independent circular buffers in memory: an audio buffer and a digital buffer. With video instruction buffer .

[0105] The module sets a system master clock. This is used for unified control of playback progress. At the beginning of each frame rendering cycle, the module reads the timestamp corresponding to the current playback sample point in the audio buffer. Based on this, the timestamp is retrieved from the video instruction buffer. The visual frame data. The module calculates the synchronization deviation value. .like If the preset synchronization threshold is exceeded, the module will either drop a frame or repeat the previous frame to force lip-sync. The read / write control of the aforementioned circular buffer and the basic clock synchronization logic can be implemented using existing multimedia player underlying technologies, which are well-known in the field and will not be elaborated upon here.

[0106] S502, real-time rendering and mesh deformation of virtual avatars.

[0107] The multimodal expression execution module 500 performs mesh deformation operations on a pre-set 3D virtual avatar model based on a synchronized sequence of facial expression frames. This 3D virtual avatar model consists of a basic mesh and multiple deformation targets.

[0108] The module reads the set of hybrid deformation weights contained in the current frame and performs a linear superposition calculation on each vertex in the model. Specifically, the final position coordinates of a vertex are equal to the base stationary coordinates of that vertex plus the sum of the weighted displacement vectors of all deformable targets, where the weight elements are the corresponding real-time weight values ​​in the set of hybrid deformation weights.

[0109] After updating the coordinates of all vertices, the module calls the rendering pipeline of the graphics processing unit, and combines preset lighting parameters and material textures to rasterize the 3D model into two-dimensional image frames. .

[0110] S503, hardware interface mapping and signal output.

[0111] The multimodal expression execution module 500, based on the system's hardware deployment configuration, renders the generated image frames. With audio stream It is converted into the corresponding physical output signal.

[0112] For display terminals, the module transmits image frames through the display interface. The data is transmitted to the screen for a refresh display.

[0113] For audio terminals, the module converts the audio stream using a digital-to-analog converter. It is converted into an analog electrical signal to drive the speaker to produce sound.

[0114] If the system contains physical robot components, the module performs motion redirection mapping. The module uses a pre-stored mapping configuration table and a set of hybrid deformation weights. The specific weight subset corresponding to the physical actuator is extracted and mapped to the target angle of the servo motor. This mapping process usually adopts a linear transformation method, which multiplies the motion weight by a preset mechanical transmission ratio coefficient and adds the initial zero offset of the motor to obtain the control signal of the target motor.

[0115] S504, rendering load monitoring and adaptive degradation.

[0116] The multimodal expression execution module runs 500 performance monitoring threads to collect the current frame rate and computational unit load rate in real time. When the frame rate is detected to be continuously lower than the preset smoothness threshold, the module triggers a rendering degradation strategy. This strategy is configured to perform operations including but not limited to: switching to a low-polygon-count LOD model, disabling dynamic shadow calculation, or reducing texture sampling resolution. By reducing computational complexity, the module maintains the continuity of audio output and the real-time nature of user interaction under hardware resource constraints.

[0117] See attached document Figure 2 In one embodiment of the present invention, the non-real-time capability assessment and profile closed-loop update in step S6 can be implemented by the asynchronous capability assessment module 600 executing the following sub-steps: S601, Interactive Session Segmentation and Multidimensional Data Aggregation.

[0118] When the asynchronous capability assessment module 600 detects an interaction end signal or an interaction silence duration exceeding a preset threshold, it triggers an asynchronous assessment task. The module retrieves complete log data for this interaction cycle from the data storage module 700. This log data includes: the text sequence input by the user, the history of strategy parameters generated by the adaptive strategy decision module 300, and the emotional state change curve recorded by the multimodal perception module 100.

[0119] The module cleans and aggregates the unstructured data to construct a session feature matrix. This matrix contains the question-and-answer pairs for each round of dialogue, the time series of user response delays, and the variance of sentiment fluctuations. For noise removal and missing value imputation during the data cleaning process, those skilled in the art can employ commonly used signal processing and interpolation algorithms, which are well-known techniques in the field and will not be elaborated upon here.

[0120] S602, Quantitative assessment of knowledge mastery and cognitive ability.

[0121] The module is based on the session feature matrix. The user's performance score in this interaction is calculated from two dimensions: breadth and depth of knowledge.

[0122] First, the module identifies all key knowledge points involved in the dialogue and extracts corresponding standard answer vectors from a structured knowledge base. The module then calculates the user's knowledge mastery index. The calculation logic for this index is as follows: Iterate through all valid question-answer pairs in the current session, calculate the cosine similarity between the user's answer vector and the standard answer vector, and then utilize the difficulty level labels associated with the knowledge points. The difficulty weight coefficient obtained by normalization is used to weight the similarity, and finally the average of the weighted similarity of all question-answer pairs is calculated.

[0123] Secondly, the module combines emotional stability and logical coherence to calculate the score vector of comprehensive cognitive performance. The dimension of this vector corresponds to the aforementioned cognitive pattern features. The dimensions remain consistent. The module statistically analyzes the user's emotional state vector during the session. The entropy variance is used as the first dimension value, and the success rate of dereference resolution in the user input text is used as the second dimension value. These values ​​are combined to form... .

[0124] S603, Long-term memory update of user profiles.

[0125] To achieve a dynamic user profile matrix The asynchronous capability assessment module 600 uses a moving average algorithm to update the user's historical capability characteristics through iterative updates. The module reads the currently stored cognitive pattern characteristics. And combined with the comprehensive cognitive performance score vector calculated in this session Generate new feature vectors. The update process is represented as: ; in: The updated cognitive pattern feature vector will be stored in data storage module 700 for use in the next interaction; The feature vector of historical cognitive patterns stored in the database before the update; The comprehensive cognitive performance score vector calculated for this session; This is a memory forgetting factor, ranging from 0 to 1. This factor controls the weight of historical experience on current evaluation. Its value is negatively correlated with the time interval since the user's last interaction; the longer the time interval, the greater the influence. The smaller the value.

[0126] S604, Personalized Learning Path Replanning and Report Generation.

[0127] After updating the user profile, the module maps the updated feature vectors to a pre-defined ability assessment coordinate system, with knowledge breadth, logical depth, and focus as axes. The module identifies the user's weak knowledge areas and strong cognitive areas.

[0128] The module generates a learning suggestion report based on the recognition results. This report includes a list of knowledge points to be reviewed and recommended types of subsequent interaction strategies. This learning suggestion report is converted into a structured data format and stored, and then pushed to the adaptive strategy decision module 300 as contextual information at the start of the next session.

[0129] See attached document Figure 1 In one embodiment of the present invention, the above-mentioned adaptive cognitive interaction system is deployed in a distributed hardware architecture that includes a cloud server cluster and an edge interaction terminal.

[0130] The hardware architecture consists of three parts at the physical level: edge interaction terminals, cloud computing platform, and communication network.

[0131] The edge interaction terminal is configured as a front-end device located at the user end, performing real-time acquisition, preprocessing, and final feedback presentation of multimodal data. In specific implementations, the edge interaction terminal includes a visual acquisition unit, an audio acquisition array, a local computing unit, and a multimodal output unit.

[0132] The visual acquisition unit is configured as at least one high-resolution optical camera, or a combined module consisting of an RGB camera and a depth sensor (such as a ToF sensor or a structured light sensor). This unit is used to capture the user's facial images and body movements, providing raw visual data to the multimodal perception module 100.

[0133] The audio acquisition array consists of microphone arrays distributed at different locations on the terminal housing. It supports beamforming and echo cancellation technology and is used to pick up the user's voice commands and environmental background noise, providing raw audio data for the multimodal sensing module 100.

[0134] The local computing unit integrates a central processing unit and a neural network processing unit. This unit is configured to run front-end models and perform tasks including face detection, keyword wake-up, and encoding and compression of audiovisual data to reduce the bandwidth consumption of data uploaded to the cloud.

[0135] The multimodal output unit includes a display screen, a speaker, and an electromechanical control subsystem. The display screen is used to present virtual images and visualized teaching content; the speaker is used to play synthesized speech; the electromechanical control subsystem includes a microcontroller, a motor drive circuit, and several servo motors. The microcontroller receives digital control signals from the local computing unit, converts them into pulse-width modulated signals, and controls the servo motors to perform physical movements through the motor drive circuit.

[0136] The cloud computing platform serves as the core processing center of the system, connecting to edge interactive terminals via a communication network. This platform deploys the aforementioned adaptive policy decision-making module 300, context orchestration and retrieval module 200, and content generation and security control module 400.

[0137] The cloud computing platform physically consists of a cluster of high-performance graphics processing units and a large-scale storage array.

[0138] The GPU cluster is configured to load and run large language models and visual encoding models in parallel, performing intent understanding, policy generation, and text / image generation computations.

[0139] The large-scale storage array contains vector databases, relational databases, and structured knowledge bases.

[0140] Vector databases are used to store knowledge fragments that have been vectorized and encoded. It supports approximate nearest neighbor retrieval for high-dimensional vectors based on user historical memory fragments.

[0141] The relational database is used to store the user's structured profile data, system logs, user authentication information, as well as the pre-built blacklist word library and security corpus required by the aforementioned content generation and security control module 400.

[0142] A structured knowledge base is used to store subject-specific knowledge entries in the field of education, and maintains each entry with its difficulty level label. The mapping relationship between them provides the data basis for the filter condition fields in the aforementioned search instructions.

[0143] The communication network is used to transmit encrypted data packets between edge interactive terminals and the cloud computing platform. This network is based on the TCP / IP protocol suite, and the physical layer employs 5G mobile communication networks, Wi-Fi 6 wireless LANs, or fiber optic wired networks. The system is configured with a data transmission mechanism based on the QUIC protocol at the network layer to reduce latency caused by handshake delays and packet loss retransmissions.

[0144] Based on the aforementioned hardware facilities, the adaptive cognitive interaction method proposed in this invention can be applied to a variety of specific implementation scenarios, including but not limited to family companion robot scenarios and smart education auxiliary terminal scenarios.

[0145] In the scenario of home companion robots, the edge interaction terminal is implemented as a robot entity with autonomous mobility.

[0146] In this scenario, the hardware interface mapping step in the aforementioned multimodal expression execution module 500 is configured as follows: the hybrid deformation weight set... A specific weight subset in the model is mapped to the target angle of the robot's servo motors. The robot uses a chassis wheel assembly for spatial movement and a head servo motor for line-of-sight tracking.

[0147] For example, when the system determines the user's real-time emotional state vector Corresponding to the loneliness label, and determining the cognitive age stage. During the pre-computation phase, the robot's control chassis moves to a user-preset distance range, and the head servo motor's pitch motion simulates a listening posture, while simultaneously displaying preset, friendly virtual avatar expressions on the screen. This interaction method, combining physical and virtual elements, can provide young children with emotional support and cognitive guidance through physical interaction.

[0148] In the context of smart education auxiliary terminals, edge interaction terminals are implemented as desktop smart tablets or smart podiums in classrooms.

[0149] In this scenario, the visual acquisition unit is configured with an overhead view to capture workbooks or teaching aids on the desktop in real time.

[0150] The system is configured to use a multimodal perception module 100 to perform text recognition and gesture trajectory analysis on overhead images to determine the user's problem-solving progress and knowledge mastery. When the system detects that the user spends more than a preset threshold time on a specific knowledge point and their facial expression corresponds to a confusion tag, the adaptive strategy decision module 300 adjusts the teaching strategy, pushing a visual breakdown animation of the knowledge point on the display screen and playing a metaphorical voice explanation through the speaker, thereby achieving personalized auxiliary teaching.

[0151] Furthermore, to ensure the security and privacy of user data, the edge interaction terminal integrates a trusted platform module. This module generates and stores hardware-level encryption keys. All collected user biometric data undergoes feature extraction and one-way hashing processing in the local computing unit before being transmitted to the cloud via a communication network. The cloud computing platform completes identity authentication and data association without decrypting the original biometric data, thereby reducing the risk of sensitive data leakage. For the specific implementation of the aforementioned hardware encryption and secure transmission, those skilled in the art can use existing AES encryption standards and SSL / TLS transmission protocols, which are well-known technologies in the field and will not be elaborated upon here.

[0152] The method in this embodiment can be used to execute the above system embodiment, and its principle and technical effect are similar, so it will not be described again here.

Claims

1. An AI multi-modal everything-recognizing child interactive star key system, characterized in that, include: The multimodal perception module (100) acquires images and audio streams, and generates intent vectors and real-time emotion state vectors using the built-in visual encoder and speech recognition unit; Data storage module (700) stores a dynamic user profile matrix representing the user's physical and mental characteristics, an age segmentation strategy, and a two-layer knowledge base; The adaptive strategy decision module (300) determines the cognitive age stage based on the real-time emotional state vector and the user dynamic profile matrix, and matches the optimal interaction strategy from the age-segmented strategies based on the determined cognitive age stage. The context arrangement and retrieval module (200) generates structured prompt words based on the optimal interaction strategy and performs vector retrieval on the two-layer knowledge base; The content generation and security control module (400) uses a large language model to generate response content based on the structured prompt words and performs security filtering; The multimodal expression execution module (500) converts the generated response content into voice output and visual output; The asynchronous capability assessment module (600) analyzes interaction data to quantitatively assess user capabilities and update the user dynamic profile matrix.

2. The AI multi-modal everything-recognizing child interaction star key system of claim 1, wherein, The user dynamic profile matrix includes physiological age, interest feature vector, five-dimensional ability score vector, real-time emotional state vector, dynamic personality feature vector, and short-term dialogue context vector. The number of dimensions of the interest feature vector is equal to the total number of entity categories in the pre-set knowledge graph, and the value of each dimension represents the user's attention weight to the corresponding entity category; the five-dimensional ability scoring vector contains normalized values ​​corresponding to the five dimensions of language expression, logical cognition, emotion management, social interaction and self-awareness.

3. The AI ​​multimodal object recognition child interactive key system according to claim 1, characterized in that, When generating the intent vector, the multimodal perception module (100) performs the following operations: The visual encoder is used to extract high-dimensional visual features from the image, identify the set of objects in the scene, and generate visual semantic description text based on the relative positional relationship between the objects. The speech recognition unit is used to transcribe the audio stream into a user input text sequence, and the visual semantic description text is then concatenated with the user input text sequence. The concatenated text is input into a natural language processing encoder, and the visual scene information is mapped to the semantic space through an attention mechanism, thereby parsing out the intent vector containing visual referential information.

4. The AI ​​multimodal object recognition child interactive key system according to claim 1, characterized in that, When determining the cognitive age stage, the adaptive strategy decision-making module (300) performs the following operations: Extract language complexity features, cognitive pattern features, and attention features from the user dynamic profile matrix and the real-time emotional state vector; The language complexity feature is calculated based on the sentence length and vocabulary ratio of the input text output by the multimodal perception module (100); The cognitive pattern features are calculated based on the density of logical connectives and the frequency of abstract nouns; The attention features are calculated based on the duration of gaze retention at the visual focus point and the response delay. The language complexity features, cognitive pattern features, and attention features are weighted, summed, and normalized to obtain a probability distribution vector. The dimension index with the largest probability value in the probability distribution vector is selected as the current user's cognitive age stage.

5. The AI ​​multimodal object recognition child interactive key system according to claim 1, characterized in that, When matching the optimal interaction strategy, the adaptive strategy decision module (300) first selects a set of candidate strategies from the age-based strategies based on the determined cognitive age stage, and then calculates the score of each strategy in the set of candidate strategies based on the objective function. The objective function is configured as follows: calculate the similarity between the strategy attribute vector and the current context vector as the context matching degree, and calculate the similarity between the strategy attribute vector and the interest feature vector and dynamic personality feature vector in the user dynamic profile matrix as the personalization adaptation degree. The context matching degree and the personalized fit degree are weighted and summed to obtain the final score of the corresponding strategy, wherein the weighting coefficient is dynamically adjusted according to the entropy value of the real-time emotional state vector.

6. The AI ​​multimodal object recognition child interactive key system according to claim 1, characterized in that, After performing vector retrieval on the two-layer knowledge base, the context orchestration and retrieval module (200) also performs reordering and filtering operations based on cognitive adaptability, specifically including: Obtain the target difficulty value mapped from the determined cognitive age stage; Calculate the comprehensive matching score for each candidate knowledge fragment. The comprehensive matching score is obtained by calculating the similarity between the query vector and the knowledge fragment vector, and subtracting the absolute value of the difference between the weighted calculation of the knowledge fragment difficulty level and the target difficulty value. The candidate knowledge fragments are sorted according to the comprehensive matching score, and the fragments with the highest ranking are selected as the associated knowledge fragments.

7. The AI ​​multimodal object recognition child interactive key system according to claim 1, characterized in that, The structured prompts generated by the context arrangement and retrieval module (200) include a role setting field, a context information field, a pedagogical constraint field, and a user input field; The character setting field is determined by the prompt word template determined by the optimal interaction strategy; The contextual information field is composed of short-term dialogue context vectors, user history interaction fragments, and related knowledge fragments; The pedagogical constraint field contains instruction text for limiting the output sentence structure and rhetorical devices of the large language model.

8. The AI ​​multimodal object recognition child interactive key system according to claim 1, characterized in that, When generating output, the content generation and security control module (400) and the multimodal expression execution module (500) perform dynamic parameter adjustments based on the real-time emotion state vector: The content generation and security control module (400) maps the probability value of the corresponding emotion tag in the real-time emotion state vector to the sampling temperature parameter of the large language model in order to control the randomness of the generated text. When performing speech synthesis, the multimodal expression execution module (500) multiplies the baseline pitch parameter in the age-classification strategy by a correction coefficient based on the determined cognitive age stage, and dynamically adjusts the baseline speech rate parameter according to the probability difference between positive and negative emotion labels in the real-time emotion state vector.

9. The AI ​​multimodal object recognition child interactive key system according to claim 1, characterized in that, When the multimodal expression execution module (500) converts the generated response content into visual output, it performs an audiovisual collaborative driving operation, specifically including: The acoustic features of the speech output are extracted and phoneme decoded to obtain a phoneme sequence. Based on the pre-configured phoneme-visual mapping relationship, the phoneme sequence is converted into a mouth shape visual weight sequence aligned with the speech time axis. The lip shape pixel weight sequence is linearly weighted and fused with the facial expression hybrid deformation weight set generated based on the real-time emotion state vector to obtain target driving frame data. If the output terminal is a 3D virtual image, the target driving frame data is used to drive the rendering engine to perform real-time displacement updates on the mesh vertex coordinates of the pre-configured 3D model to achieve facial animation rendering. If the output terminal is a physical robot component, the target drive frame data is mapped to the target control angle of the joint servo motor of the physical robot component through a pre-configured degree of freedom mapping matrix.

10. The AI ​​multimodal object recognition child interactive key system according to claim 1, characterized in that, When updating the user dynamic profile matrix, the asynchronous capability assessment module (600) uses a moving average algorithm to update capability features. The specific steps for updating capability features are as follows: Read the historical ability feature vector stored in the data storage module (700) and combine it with the comprehensive cognitive performance score vector obtained from this interaction calculation; The historical ability feature vector and the comprehensive cognitive performance subvector are weighted and summed using a memory forgetting factor to obtain an updated ability feature vector. The value of the memory forgetting factor is negatively correlated with the time interval between the user's last interaction.

Citation Information

Cited By

  • A child positive guidance prompt word generation method and system based on personality-emotion interaction

    CN122153012A