Intelligent interaction method and language capability classification model training method and device

By acquiring the target user's language knowledge graph and behavioral information, and formulating personalized interaction strategies, this technology solves the problem that existing intelligent companion robots cannot effectively guide individual language development, and achieves intelligent and personalized human-computer interaction that matches the user's language ability.

CN116127006BActive Publication Date: 2026-04-28MASHANG CONSUMER FINANCE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
MASHANG CONSUMER FINANCE CO LTD
Filing Date
2022-10-26
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing intelligent companion robots lack effective intelligent means to guide individual language development, resulting in poor human-computer interaction, especially the inability to formulate personalized interaction strategies based on the user's language ability category.

Method used

By acquiring the target user's target language knowledge graph, matching target knowledge information according to the user's language ability category, and combining user behavior information to formulate interaction strategies, personalized and customized human-computer interaction can be achieved.

Benefits of technology

Ensure that the interaction process matches the user's language ability category, avoid interaction barriers, improve the intelligence and personalization of human-computer interaction, and provide accurate language ability guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116127006B_ABST
    Figure CN116127006B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose an intelligent interaction method, a language ability classification model training method and device. The intelligent interaction method comprises: obtaining a target language knowledge graph corresponding to a target user; the target language knowledge graph is created based on a language ability category of the target user, and the language ability category of the target user is determined based on a language ability feature of the target user; determining target knowledge information matched with the language ability category of the target user according to the target language knowledge graph; obtaining behavior information of the target user; the behavior information comprises at least one of the following: action information, language information and emotion information; determining an interaction strategy corresponding to the target user according to the target knowledge information and the behavior information, and interacting with the target user based on the interaction strategy. The technical solution can realize personalized human-computer interaction effect matched with the language ability category of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence technology, and in particular to an intelligent interaction method, a language ability classification model training method and device. Background Technology

[0002] Individual language development guidance is typically provided by guardians or teachers based on their experience. To achieve intelligent language development guidance, related technologies include intelligent companion robots and systems, such as children's companion robots and systems, which mainly include a voice input component, a touch input component, a sound output component, and a display component. The voice input component extracts the user's voice information from ambient sounds; the touch input component receives the user's touch operations through an interactive interface; the sound output component outputs a response voice or a feedback sound in response to touch operations; and the display component outputs the interactive interface or facial expressions matching the voice information and response. It is evident that intelligent companion robots possess the function of intelligently accompanying users, saving users (such as parents and guardians) the effort and time spent on accompanying them. However, existing intelligent companion robots are limited to saving the effort and time spent on human companionship; they have not yet proposed effective solutions for more professionally guiding the development of individual language abilities. Summary of the Invention

[0003] The purpose of this application is to provide an intelligent interaction method, a language ability classification model training method and apparatus to solve the problem of the lack of intelligent guidance for individual language development in the prior art.

[0004] To solve the above-mentioned technical problems, the embodiments of this application are implemented as follows:

[0005] On one hand, embodiments of this application provide an intelligent interaction method, including:

[0006] Obtain a target language knowledge graph corresponding to the target user; the target language knowledge graph is created based on the target user's language ability category, and the target user's language ability category is determined based on the target user's language ability characteristics;

[0007] Based on the target language knowledge graph, target knowledge information matching the target user's language ability category is determined;

[0008] Obtain behavioral information of the target user; the behavioral information includes at least one of the following: action information, language information, and emotional information;

[0009] Based on the target knowledge information and the behavioral information, an interaction strategy corresponding to the target user is determined, and interaction with the target user is carried out based on the interaction strategy.

[0010] The technical solution of this application involves acquiring a target language knowledge graph corresponding to the target user, determining target knowledge information matching the target user's language ability category based on the target language knowledge graph, acquiring the target user's behavioral information, and then determining an interaction strategy corresponding to the target user based on the target knowledge information and the target user's behavioral information. Since the interaction strategy is jointly determined based on the target knowledge information matching the target user's language ability category and the target user's behavioral information, and the target knowledge information is determined based on the target language knowledge graph created based on the target user's language ability category, when interacting with the target user based on the interaction strategy, it not only ensures that the interaction process matches the target user's language ability category, avoiding interaction obstacles (such as difficulty understanding language or actions during the interaction process) caused by a mismatch between the interaction strategy and the target user's language ability category, thus achieving accurate guidance for the target user during human-computer interaction, but also comprehensively considers the target user's action information, language information, and / or emotional information to formulate the interaction strategy, improving the intelligence and personalization of human-computer interaction. Furthermore, since there is a correspondence between the target language knowledge graph and the target user, that is, each user has its own corresponding language knowledge graph, this technical solution can not only achieve intelligent interaction with users, but also achieve customized language knowledge graphs for different users, thereby further realizing personalized and customized human-computer interaction effects.

[0011] On the other hand, embodiments of this application provide a method for training a language ability classification model, including:

[0012] Obtain the sample language ability characteristics and sample language ability categories of sample users; the sample language ability characteristics include object response characteristics and / or sound response characteristics;

[0013] The language ability features of the sample users are input into the language ability classification model to be trained, and the language ability of the sample users is classified to obtain the classification result.

[0014] Based on the classification results and the language ability categories of the samples, the model parameters of the language ability classification model to be trained are adjusted.

[0015] The technical solution of this application involves acquiring sample language ability features and categories of sample users, inputting these features into a language ability classification model to be trained, classifying the language abilities of the sample users, obtaining classification results, and then adjusting the model parameters of the language ability classification model to be trained based on the classification results and sample language ability categories, thereby obtaining a trained language ability classification model. Since the sample language ability features of the sample users include object response features and / or vocal response features, the language ability classification model trained based on these features not only has the ability to analyze the user's object response features and / or vocal response features, but also the ability to classify the user's language abilities based on these features. Furthermore, during intelligent human-computer interaction, the language ability classification model can accurately analyze the target user's language ability category, providing strong data support for the subsequent intelligent human-computer interaction process.

[0016] Furthermore, embodiments of this application provide an intelligent interactive device, including:

[0017] The first acquisition module is used to acquire a target language knowledge graph corresponding to the target user; the target language knowledge graph is created based on the target user's language ability category, and the target user's language ability category is determined based on the target user's language ability characteristics;

[0018] The first determining module is used to determine target knowledge information that matches the language ability category of the target user based on the target language knowledge graph;

[0019] The second acquisition module is used to acquire the behavioral information of the target user; the behavioral information includes at least one of the following: action information, language information, and emotional information;

[0020] The second determining module is used to determine an interaction strategy corresponding to the target user based on the target knowledge information and the behavioral information, and to interact with the target user based on the interaction strategy.

[0021] In another aspect, embodiments of this application provide an electronic device, including a processor and a memory electrically connected to the processor. The memory stores a computer program, and the processor is configured to call and execute the computer program from the memory to implement the intelligent interaction method described above, or the processor is configured to call and execute the computer program from the memory to implement the language ability classification model training method described above.

[0022] In another aspect, embodiments of this application provide a storage medium for storing a computer program that can be executed by a processor to implement the intelligent interaction method described above, or the computer program can be executed by a processor to implement the language ability classification model training method described above. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in one or more embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a schematic flowchart of an intelligent interaction method according to an embodiment of this application;

[0025] Figure 2 This is a schematic structural diagram of an initial language knowledge graph according to an embodiment of this application;

[0026] Figure 3 This is a schematic structural diagram of a target language knowledge graph according to an embodiment of this application;

[0027] Figure 4 This is a schematic diagram illustrating the application principle of a language ability classification model according to an embodiment of this application;

[0028] Figure 5 This is a schematic flowchart illustrating an intelligent interaction method according to an embodiment of this application;

[0029] Figure 6 This is a schematic flowchart illustrating an intelligent interaction method according to another embodiment of this application;

[0030] Figure 7 This is a schematic flowchart of a language ability classification model training method according to an embodiment of this application;

[0031] Figure 8 This is a schematic diagram illustrating a language ability classification model training method according to an embodiment of this application;

[0032] Figure 9 This is a schematic block diagram of an intelligent interactive device according to an embodiment of this application;

[0033] Figure 10 A schematic block diagram of a language ability classification model training device according to an embodiment of this application;

[0034] Figure 11 A schematic block diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0035] This application provides an intelligent interaction method, a language ability classification model training method and apparatus to solve the problem of the lack of intelligent guidance for individual language development in the prior art.

[0036] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.

[0037] In guiding individual language development, related technologies offer intelligent companion robots and systems, such as child companion robots and systems, primarily comprising voice input components, touch input components, sound output components, and display components. The voice input component extracts the user's voice information from ambient sounds; the touch input component receives touch inputs through an interactive interface; the sound output component outputs response voice messages or feedback sounds in response to touch inputs; and the display component outputs the interactive interface or facial expressions matching the voice messages and responses. Taking a child companion robot as an example, during human-computer interaction, the child needs to actively make sounds or perform interactive operations through the interface, such as triggering designated buttons. Only then will the robot collect the child's voice messages or commands issued through the interface and execute corresponding actions based on the collected information or commands. It is evident that while child companion robots offer intelligent companionship, saving parents time and effort, they are limited to saving time and effort compared to human companionship and do not focus on more professionally guiding the development of individual language abilities. The intelligent interaction method provided in this application obtains a target language knowledge graph corresponding to the target user, determines target knowledge information matching the target user's language ability category based on the target language knowledge graph, and obtains the target user's behavioral information. Then, based on the target knowledge information and the target user's behavioral information, an interaction strategy corresponding to the target user is determined. Since the interaction strategy is jointly determined based on the target knowledge information matching the target user's language ability category and the target user's behavioral information, and the target knowledge information is determined based on the target language knowledge graph created based on the target user's language ability category, when interacting with the target user based on the interaction strategy, it not only ensures that the interaction process matches the target user's language ability category, but also avoids interaction obstacles (such as difficulty understanding language or actions during the interaction process) caused by a mismatch between the interaction strategy and the target user's language ability category, thereby achieving accurate guidance for the target user during human-computer interaction.

[0038] Furthermore, the aforementioned child companion robots interact with children by executing actions based on collected information (such as the child's voice or commands issued through the interface) and a built-in universal interaction method. This universal interaction method refers to using the same interaction method for all users. For example, the robot has a built-in dialogue algorithm that collects the child's voice information and performs speech and semantic analysis to engage in dialogue with the child. Clearly, the child companion robot does not consider the child's current language development level during human-computer interaction, resulting in poor interaction effectiveness. For instance, if a child's language development is still at the simple sentence stage, meaning they can only converse in simple sentences, if the child companion robot issues complex, long sentences based on the universal interaction method, the child will have difficulty understanding the robot's meaning and will be unable to proceed with further interaction. The intelligent interaction method provided in this application can not only customize a personalized target language knowledge graph for the target user, but also comprehensively consider the target knowledge information that matches the target user's language ability category, as well as the target user's action information, language information and / or emotional information to formulate an interaction strategy. Therefore, the interaction process fully considers the target user's current language development ability, making the human-computer interaction process more personalized and customized.

[0039] The intelligent interaction method and language ability classification model training method provided in this application can be executed by an electronic device or by software installed in an electronic device. Specifically, the electronic device can be a terminal device or a server device. In this application embodiment, the electronic device can be an intelligent robot with interactive functions.

[0040] Figure 1 This is a schematic flowchart of an intelligent interaction method according to an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps S102-S108:

[0041] S102, Obtain the target language knowledge graph corresponding to the target user; the target language knowledge graph is created based on the target user's language ability category, and the target user's language ability category is determined based on the target user's language ability characteristics.

[0042] The language ability category is used to represent the user's current language ability. Optionally, according to the developmental stages of language ability, it can be divided into the following categories: language preparation stage, language completion stage, language intelligence stage, dialogue stage, writing stage, echo stage, addressing stage, speaking center, writing center, visual language center, language center, etc. Furthermore, the developmental stages of language ability can be further subdivided. For example, the language preparation stage includes the single-word sentence stage, the two-word sentence stage, and the simple sentence stage; the language completion stage includes the compound sentence stage and the interrogative sentence stage; and so on.

[0043] A target language knowledge graph comprises multiple knowledge nodes and the corresponding knowledge information for each node. Each knowledge node corresponds to a specific language ability category. In other words, the target language knowledge graph contains a knowledge node corresponding to each language ability category. For example, the target language knowledge graph includes a knowledge node corresponding to the language preparation stage, and the next-level knowledge nodes for that stage are the knowledge nodes corresponding to the single-word sentence stage, the two-word sentence stage, and the simple sentence stage, respectively.

[0044] Optionally, the target user's language ability category can be determined based on a pre-trained language ability classification model, which is trained based on sample language ability features and categories from multiple sample users. The specific training method for the language ability classification model will be described in detail in the following embodiments and will not be repeated here. Alternatively, the target user's language ability category can also be determined in other ways, not limited to the pre-trained language ability classification model. For example, different correspondences between language ability features and language ability categories can be pre-defined, thereby determining the language ability category corresponding to the target user's language ability features based on these correspondences.

[0045] S104. Based on the target language knowledge graph, determine the target knowledge information that matches the target user's language ability category.

[0046] Among them, the target knowledge information that matches the target user's language ability category can be understood as the language information that the target user can learn at the stage corresponding to their language ability category, including phonetics (such as the phonetics of characters, words, and sentences), semantics, grammar, pragmatic skills (such as single-word sentences and complex sentences), etc.

[0047] S106, Obtain the target user's behavioral information; the behavioral information includes at least one of the following: action information, language information, and emotional information.

[0048] Among them, action information refers to the actions performed by the target user, such as raising a hand or picking up food; language information refers to the sounds made by the target user; and emotional information refers to information that can represent the target user's emotions, including emotions (such as crying or laughing) and facial expressions (such as squinting or grinning).

[0049] S108. Based on the target knowledge information and behavioral information, determine the interaction strategy corresponding to the target user, and interact with the target user based on the interaction strategy.

[0050] The interaction strategy may include the interaction method and the interaction content corresponding to each interaction method. The interaction method may include language interaction and / or action interaction.

[0051] Optionally, the interaction method for the target user can be determined based on the target user's behavioral information, and the interaction content corresponding to the interaction method can be further determined based on the target knowledge information that matches the target user's language ability category.

[0052] For example, in the target knowledge information that matches the target user's language ability category, the sentence-building method that the target user can learn is a single-word sentence, and the target user's behavioral information is: looking at the apple in the room, then the corresponding interaction strategy for the target user can be determined as follows: the interaction methods include language interaction and action interaction. The interaction content corresponding to the language interaction method is to utter the single-word sentence "apple", and the interaction content corresponding to the action interaction method is: to get the apple for the target user and display the image "apple" on the screen that is pre-configured and visible to the target user.

[0053] The technical solution of this application involves acquiring a target language knowledge graph corresponding to the target user, determining target knowledge information matching the target user's language ability category based on the target language knowledge graph, acquiring the target user's behavioral information, and then determining an interaction strategy corresponding to the target user based on the target knowledge information and the target user's behavioral information. Since the interaction strategy is jointly determined based on the target knowledge information matching the target user's language ability category and the target user's behavioral information, and the target knowledge information is determined based on the target language knowledge graph created based on the target user's language ability category, when interacting with the target user based on the interaction strategy, it not only ensures that the interaction process matches the target user's language ability category, avoiding interaction obstacles (such as difficulty understanding language or actions during the interaction process) caused by a mismatch between the interaction strategy and the target user's language ability category, thus achieving accurate guidance for the target user during human-computer interaction, but also comprehensively considers the target user's action information, language information, and / or emotional information to formulate the interaction strategy, improving the intelligence and personalization of human-computer interaction. Furthermore, since there is a correspondence between the target language knowledge graph and the target user, that is, each user has its own corresponding language knowledge graph, this technical solution can not only achieve intelligent interaction with users, but also achieve customized language knowledge graphs for different users, thereby further realizing personalized and customized human-computer interaction effects.

[0054] In one embodiment, the acquisition (or generation) of a target language knowledge graph may include the following steps A1-A5:

[0055] Step A1: Obtain the pre-created initial language knowledge graph; the initial language knowledge graph includes: multiple knowledge nodes, and user information and knowledge information corresponding to each knowledge node.

[0056] The initial language knowledge map can be generated based on classical language development theories. Classical language development theories, also known as "language acquisition theories," explain children's acquisition of speaking and listening abilities in their native language. Language consists of three components: phonetics, grammar, and semantics. As an effective communication tool, language requires both the speaker and the listener to possess a series of skills and rules—pragmatic skills. Children gradually master these four components—phonetics, grammar, semantics, and pragmatic skills—in order to acquire the ability to learn and understand their native language.

[0057] Optionally, based on classical language development theory, entity recognition technology is used to acquire all entities within the theory, including the phonetic, grammatical, semantic, and pragmatic skills to be learned at each stage of language development. Relationship extraction and attribute extraction techniques are then used to extract the relationships between all entities. Based on the acquired entities and their relationships, each stage of language development is treated as a knowledge node, and the knowledge to be learned at each stage (including phonetic, grammatical, semantic, and pragmatic skills) is associated with that knowledge node, thus constructing a complete initial language knowledge graph. Entity recognition, relationship extraction, and attribute extraction techniques are all existing technologies and will not be elaborated upon further.

[0058] Optionally, the arrows between two adjacent knowledge nodes are used to indicate the developmental direction of the language development stages corresponding to the two adjacent knowledge nodes. For different language ability classification dimensions, a parent-child node approach can be used to construct the graph. The child nodes of the knowledge node "Language Preparation Period" include the knowledge nodes corresponding to the single-word sentence stage, the two-word sentence stage, and the simple sentence stage, respectively.

[0059] Each knowledge node corresponds to its own user information. User information may include at least one of the following: user age range, user identity information, etc. User age range can be represented by a specific age range, such as 0-1 years old, 1-2 years old, etc. User identity information may include infants, preschool children, primary school students, etc.

[0060] Figure 2 This is a schematic structural diagram of an initial language knowledge graph according to an embodiment of this application. Due to limitations in graph size, Figure 2Only a portion of the knowledge nodes in the initial language knowledge graph are shown, with ellipses indicating other knowledge nodes and sub-nodes not shown. Arrows between adjacent knowledge nodes indicate the direction of language development at the corresponding stages. For example, an arrow between the knowledge nodes "single-word stage" and "two-word stage," pointing towards the "two-word stage," indicates that the language development direction is from the single-word stage to the two-word stage. It should be noted that arrows may or may not exist between adjacent knowledge nodes at the same level, or between multiple sub-nodes of the same knowledge node. If no arrow exists between two knowledge nodes, it means there is no connection between them.

[0061] Step A2: Based on the initial language knowledge graph and the target user's user information, determine the knowledge information that matches the target user's user information.

[0062] The target user's information may include at least one of the following: the target user's age, user identity information, etc. By comparing the target user's information with the user information corresponding to each knowledge node in the initial language knowledge graph, knowledge nodes matching the target user's information can be identified. The knowledge information corresponding to these matching knowledge nodes is then determined; this is the knowledge information that matches the target user's information. The knowledge information that matches the target user's information is the knowledge information that the target user should acquire at this stage, also known as prior knowledge information.

[0063] Step A3: Based on knowledge information that matches the user information of the target user, collect the language ability characteristics of the target user; the language ability characteristics include object response characteristics and / or sound response characteristics.

[0064] Optionally, when collecting the target user's object-related reaction characteristics, an object-related action matching the target user's prior knowledge information can be performed, such as moving the food "apple" from one location to the target user's eyes, and then collecting the target user's reaction information to this action. When collecting the target user's voice-related reaction characteristics, voice information matching the target user's prior knowledge information can be emitted to the user, such as emitting the voice "apple" to the target user if the target user should currently learn single-word sentences, and then collecting the target user's reaction information to this voice.

[0065] Optionally, the response characteristics to objects can be characterized by the sensitivity to those objects, and the response characteristics to sound can be characterized by the sensitivity to those sounds. When collecting information on the target user's response to actions or speech, existing brain signal acquisition technologies can be used to collect the target user's brain signals, including at least one of electroencephalogram (EEG) data and the brightness of the brain's language areas. By analyzing the brain signals, the target user's response information, such as response sensitivity, can be determined. EEG data represents the changes in electrical waves during brain activity; therefore, the larger and faster the amplitude of these changes, the higher the target user's response sensitivity, and vice versa. The brightness of the brain's language areas reflects the activity level of those areas; therefore, the higher the brightness of the language areas, the higher the target user's response sensitivity, and vice versa.

[0066] Step A4: Determine the target user's language ability category based on language ability characteristics.

[0067] Step A5: Generate a target language knowledge graph based on the initial language knowledge graph and the target user's language ability category.

[0068] In this embodiment, each knowledge node corresponds to a language ability category. Based on this, when generating a target language knowledge graph from the initial language knowledge graph and the target user's language ability categories, the target user's language ability categories are first matched with each knowledge node. The first knowledge node matching the target user's language ability category is then determined based on the matching results. The target language knowledge graph is then generated based on the first knowledge node. The target language knowledge graph includes the first knowledge node and the corresponding knowledge information. Therefore, when determining the target knowledge information matching the target user's language ability category based on the target language knowledge graph, the knowledge information corresponding to the first knowledge node can be identified as the target knowledge information.

[0069] Furthermore, to enable the target language knowledge graph to provide language guidance, it can be generated based on a first knowledge node and a second knowledge node. The target language knowledge graph includes: a first knowledge node, a second knowledge node, knowledge information corresponding to the first knowledge node, and knowledge information corresponding to the second knowledge node. The second knowledge node is at least one knowledge node adjacent to the first knowledge node. Optionally, the second knowledge node is the next knowledge node adjacent to the first knowledge node. Thus, based on the target language knowledge graph, the target user's next stage of language ability category and the knowledge information the target user should acquire in the next stage can be determined. Therefore, when determining target knowledge information matching the target user's language ability category based on the target language knowledge graph, the target knowledge information can be determined based on the knowledge information corresponding to the first knowledge node and / or the second knowledge node. For example, if the first knowledge node includes multiple child nodes, and each pair of adjacent child nodes has an arrow indicating the direction of stage development, and the target user currently corresponds to the first child node of the first knowledge node (i.e., the earliest stage child node), then the knowledge information corresponding to the first knowledge node can be determined as the target knowledge information, or the knowledge information corresponding to the first child node of the first knowledge node can be determined as the target knowledge information. If the target user is currently in the last child node of the first knowledge node (i.e., the latest stage child node), since the next stage of the last child node will be the second knowledge node, in order to guide the target user's language ability development step by step, the knowledge information corresponding to the first knowledge node and the knowledge information corresponding to the second knowledge node can be jointly determined as the target knowledge information, or the knowledge information corresponding to the last child node of the first knowledge node and the knowledge information corresponding to the second knowledge node can be jointly determined as the target knowledge information.

[0070] In one embodiment, the knowledge nodes included in the target language knowledge graph can be understood as knowledge nodes displayed on the front-end interface, while other knowledge nodes included in the initial language knowledge graph but not belonging to the target language knowledge graph can be understood as knowledge nodes hidden on the front-end interface. Specifically, in response to a display command for the target language knowledge graph, the knowledge nodes included in the target language knowledge graph, such as the first knowledge node, or the first knowledge node and the second knowledge node, are displayed on the front-end interface, while other knowledge nodes included in the initial language knowledge graph but not belonging to the target language knowledge graph are hidden. Furthermore, the front-end interface can provide an input field for displaying other knowledge nodes; when a display command is received through the input field, the hidden other knowledge nodes are displayed.

[0071] Figure 3This is a schematic structural diagram of a target language knowledge graph according to an embodiment of this application. In this embodiment, it is assumed that the target user's current language ability category is the single-word sentence stage, and the next stage's corresponding language ability category is the two-word sentence stage. Figure 3 The target language knowledge graph shown only includes knowledge nodes corresponding to the single-word sentence stage and knowledge nodes corresponding to the two-word sentence stage.

[0072] In this embodiment, by hiding other knowledge nodes included in the initial language knowledge graph but not belonging to the target language knowledge graph, and only displaying the knowledge nodes included in the target language knowledge graph, users can clearly understand the target user's current language ability category at a glance. Furthermore, if the target language knowledge graph includes a second knowledge node, users can also clearly understand the target user's next stage of language ability category. In addition, during human-computer interaction with the target user, interaction can be based on the knowledge information corresponding to the second knowledge node included in the target language knowledge graph. This allows for guidance on the next stage of learning information for the target user through human-computer interaction, not only improving the target user's human-computer interaction experience but also accelerating the target user's language ability development.

[0073] In one embodiment, when performing step A4 above, the target user's language ability category can be determined based on the target user's language ability characteristics and a pre-trained language ability classification model. Specifically, the target user's language ability characteristics can be input into the pre-trained language ability classification model to obtain the target user's language ability category, wherein the language ability classification model is trained based on the sample language ability characteristics and sample language ability categories of sample users.

[0074] Language ability features may include language ability feature values. When there are multiple language ability features, each language ability feature corresponds to its own weight. Optionally, language ability features include object response features and sound response features. Object response features are characterized by sensitivity to object responses, and sound response features are characterized by sensitivity to sound responses. Therefore, sensitivity to object responses can be used as object response feature values, and sensitivity to sound responses can be used as sound response feature values.

[0075] Optionally, to facilitate model calculation, the language ability feature values ​​can be normalized to obtain normalized language ability feature values. The normalized language ability feature values ​​are distributed between 0 and 1. These normalized language ability feature values ​​are then input into the language ability classification model.

[0076] In one embodiment, language proficiency features may include language proficiency feature values, and the pre-trained language proficiency classification model includes a score calculation layer and a classification layer. When determining the language proficiency category of a target user based on the language proficiency features and the pre-trained language proficiency classification model, the language proficiency feature values ​​are first input into the pre-trained language proficiency classification model, and then the following steps B1-B2 are performed using the language proficiency classification model:

[0077] Step B1: Through the score calculation layer, the language ability score of the target user is calculated based on the language ability feature value corresponding to each language ability feature and the weight corresponding to each language ability feature.

[0078] In this step, the language ability feature values ​​are weighted and summed according to the weights corresponding to each language ability feature to obtain the target user's language ability score. For example, if the weight corresponding to the object response feature value is 0.6 and the weight corresponding to the voice response feature value is 0.4, and the target user's object response feature value is determined to be 0.9 and the voice response feature value is 0.1, then the target user's language ability score is: 0.9*0.6 + 0.1*0.4 = 0.58.

[0079] Step B2: Through the classification layer, the target user's language ability category is determined based on the target user's language ability score and the preset mapping relationship between language ability scores and language ability categories.

[0080] In the preset mapping relationship between language proficiency scores and language proficiency categories, the higher the language proficiency score, the higher the language development stage of the corresponding language proficiency category.

[0081] Figure 4 This is a schematic diagram illustrating the application principle of a language ability classification model according to an embodiment of this application. Figure 4 As shown, the language ability feature values ​​are input into the language ability classification model. Specifically, they are input into the score calculation layer of the language ability classification model. The score calculation layer calculates the language ability score of the target user, and then the language ability score is input into the classification layer of the language ability classification model. The classification layer classifies the target user's language ability according to the preset mapping relationship between the language ability score and the language ability category, and thus outputs the language ability category of the target user.

[0082] For example, the target user's language proficiency score is calculated to be 0.58. In the preset mapping relationship between language proficiency scores and language proficiency categories, 0.58 corresponds to the language proficiency category of the "language completion stage". Therefore, the target user's language proficiency category can be determined to be the language completion stage.

[0083] In one embodiment, during human-computer interaction with a target user, interaction feature information of the target user during the interaction process can be obtained. This interaction feature information includes object response features and / or voice response features. Then, it is determined whether the interaction feature information matches the target user's language ability category. If they do not match, the target user's language ability features are redefined based on the obtained interaction feature information, resulting in updated language ability features. Then, the target user's language ability category is redefined based on the updated language ability features, resulting in a language ability category to be updated. Finally, the target language knowledge graph is updated based on the language ability category to be updated.

[0084] Since a target language knowledge graph can include a first knowledge node, a second knowledge node, the knowledge information corresponding to the first knowledge node, and the knowledge information corresponding to the second knowledge node, when guiding the development of a target user's language ability through the target language knowledge graph, the knowledge node corresponding to the second knowledge node (i.e., as the target knowledge node) can be used for human-computer interaction with the target user. If, based on the interaction characteristics information during the interaction process, it is determined that the target user's language ability category matches the knowledge node corresponding to the second knowledge node, then the target user's language ability is considered to have developed to the second knowledge node. In this case, the second knowledge node can be updated to a new first knowledge node in the target language knowledge graph.

[0085] In this embodiment, when redetermining the target user's language ability category based on interaction feature information, the method is similar to that used in the above embodiment to determine the target user's language ability category through steps A3-A4, and will not be repeated here.

[0086] In one embodiment, the human-computer interaction process can be actively triggered by an electronic device, or the electronic device can monitor the behavioral information of the target user and trigger the human-computer interaction process when the behavioral information is detected.

[0087] If the human-computer interaction process can be actively triggered by an electronic device, then when acquiring the target user's behavioral information (i.e., executing S106), the target user can be triggered to perform a behavioral event based on target knowledge information that matches the target user's language ability category, and then the behavioral information corresponding to the behavioral event can be acquired. The behavioral information corresponding to the behavioral event may include at least one of action information, language information, and emotional information.

[0088] Optionally, when obtaining behavioral information corresponding to a behavioral event, multimedia data such as audio and video data of the target user performing the behavioral event can be obtained first. Then, the multimedia data can be processed to obtain audio data and image data of the target user performing the behavioral event. Finally, by analyzing the audio data and image data, the behavioral information corresponding to the behavioral event can be obtained.

[0089] If the electronic device monitors the target user's behavior information and triggers a human-computer interaction process when the behavior information is detected, the target user's behavior information can be monitored in real time or according to a preset period. The behavior information may include at least one of action information, language information, and emotional information. When the behavior information is detected, the human-computer interaction process is triggered, and S108 in the above embodiment is executed, that is, an interaction strategy corresponding to the target user is determined, and interaction with the target user is carried out based on the interaction strategy.

[0090] The following specific examples illustrate in detail how to conduct human-computer interaction with a target user under two different triggering methods for the human-computer interaction process. Figure 5 and Figure 6 In the specific embodiment shown, a target language knowledge graph pre-created for the target user is deployed in an intelligent robot. The intelligent robot has a display interface that can display the target user's target language knowledge graph and display text and image information during the human-computer interaction process for the target user.

[0091] Optionally, at least one set of login information can be pre-deployed in the intelligent robot. Each set of login information includes a login account and a corresponding login password. Each set of login information corresponds to a target user, and for each set of login information, a target language knowledge graph associated with the target user is stored. Before using the intelligent robot for human-computer interaction, users need to log in to the intelligent robot using the login information. The intelligent robot determines the target language knowledge graph associated with the login information based on the currently logged-in login information and performs human-computer interaction based on the determined target language knowledge graph.

[0092] By deploying at least one set of login information in the intelligent robot, it is possible not only to use the same intelligent robot to accompany multiple target users, but also to ensure the security of the target user's target language knowledge graph. Specifically, if the login information provided when logging into the intelligent robot is incorrect (such as a mismatch between the login account and password), the intelligent robot will not display the internally stored target language knowledge graph, thereby preventing others from obtaining the target user's target language knowledge graph and ensuring its security.

[0093] Figure 5 This is a schematic flowchart illustrating an intelligent interaction method according to an embodiment of this application. In this embodiment, the human-computer interaction process is actively triggered by an intelligent robot, such as... Figure 5 As shown, the intelligent interaction method includes the following steps S501-S505:

[0094] S501, based on the target user's target language knowledge graph, determine the target knowledge information that matches the target user's language ability category.

[0095] Since the target language knowledge graph includes knowledge nodes corresponding to the target user's current language ability category, as well as the knowledge information corresponding to those knowledge nodes, the knowledge information corresponding to the knowledge nodes in the target language knowledge graph can be identified as target knowledge information that matches the target user's language ability category. Target knowledge information may include phonetics (such as the phonetics of characters, words, and sentences), semantics, grammar, pragmatic skills (such as single-word sentences and complex sentences), etc.

[0096] S502, based on the target knowledge information, trigger the target user to perform an action event.

[0097] Behavioral events may include voice events and / or action events.

[0098] Optionally, the electronic device can trigger a behavior event by performing a behavior-triggered event on the target user. The behavior-triggered event may include a voice output event and / or an action execution event, whereby the voice output event is the output of voice information matching the target knowledge information to the target user. Assuming the target user is currently at the single-word / sentence stage, the target knowledge information corresponding to the target user includes the pragmatic skills of single-word / sentence phrases. Therefore, single-word / sentence voice information can be output to the target user, for example, using a voice familiar to the target user (such as a family member's voice) to output the voice word "apple".

[0099] An action execution event is an action performed on the target user that matches the target knowledge information. Assuming the target user is currently at the single-word sentence stage, the target knowledge information for the target user includes the pragmatic skills of single-word sentences. This means that the target user can only understand simple single-word objects at this stage. Single-word objects can be understood as objects whose names are single words. Therefore, actions related to single-word objects can be performed on the user, such as moving an apple from a certain location to in front of the target user.

[0100] S503 captures audio and video data of the target user's behavioral events using a pre-installed camera device; and collects brain signals from the target user.

[0101] In this embodiment, the audio and video data includes audio data and / or image data. Camera devices can be pre-installed around the target user to record their behavior during human-computer interaction. Optionally, the camera device activates and records the target user's behavior after the human-computer interaction is triggered. Brain signals may include electroencephalogram (EEG) data, brightness of language regions of the brain, etc.

[0102] S504 analyzes the captured audio and video data and brain signals to obtain behavioral information corresponding to the behavioral events.

[0103] The behavioral information corresponding to the behavioral event may include at least one of the following: action information, language information, and emotional information.

[0104] The camera device communicates with the intelligent robot, transmitting captured audio and video data to the robot via this connection. Optionally, if the target user's behavior is a voice event, the intelligent robot can obtain the user's voice information directly from its built-in voice acquisition device, without relying on the camera device.

[0105] Optionally, a pre-set waiting time can be configured after the intelligent robot executes a behavior trigger event, during which time the robot can acquire behavior information corresponding to the behavior event executed by the target user. If the target user's behavior information is not acquired within the waiting time, the current human-computer interaction is considered to have failed. In the event of failure, the process can return to S502 and execute the behavior trigger event to the target user again.

[0106] When analyzing audio and video data to obtain behavioral information corresponding to behavioral events, the audio and image data can be separated first, and then analyzed separately. For audio data, existing audio recognition algorithms can be used to identify the audio data, which can then be converted into text content with semantic information. For image data, computer vision algorithms (such as OpenCV) can be used to segment the image data into frames, and then image recognition can be performed on the resulting multi-frame images to identify the target user's behavioral information, such as whether the target user reacted, what actions the target user performed, what facial expressions the target user made, etc.

[0107] Brain signals can be analyzed using EEG analysis algorithms, including analyzing EEG data and / or the brightness of the language areas of the brain, to determine the target user's reaction sensitivity. Specifically, EEG data represents the electrical wave changes during brain activity. Therefore, the larger and faster the amplitude of these changes in EEG data, the higher the target user's reaction sensitivity; conversely, the smaller and slower the amplitude of these changes, the lower the target user's reaction sensitivity. The brightness of the language areas of the brain reflects the activity level of those areas. Therefore, the higher the brightness of the language areas, the higher the target user's reaction sensitivity; conversely, the lower the brightness, the lower the target user's reaction sensitivity.

[0108] Typically, behavioral information of a target user can be comprehensively analyzed by combining audio / video data and brain signals. Audio / video data is mainly used to analyze whether the target user reacts, and if so, what specific actions or speech they perform. Brain signals are mainly used to analyze the target user's responsiveness. For example, analyzing audio / video data can determine that the target user performed a head-turning action, while analyzing brain signals can determine the target user's sensitivity in performing that action.

[0109] The S505 analyzes the captured audio and video data and brain signals to obtain the interaction characteristics of the target user during the interaction process, and stores the interaction characteristics of the target user during the interaction process locally.

[0110] The locally stored interaction feature information can be used for subsequent updates to the target language knowledge graph. For example, if an update cycle is preset, when the update cycle arrives, the target user's current language development category is analyzed based on the locally stored interaction feature information, and then the target language knowledge graph is updated according to the target user's current language development category.

[0111] exist Figure 5 In the illustrated embodiment, the target user can be any user who needs to use the intelligent robot, such as an infant who needs companionship or an elderly person with limited mobility.

[0112] Figure 6 This is a schematic flowchart illustrating an intelligent interaction method according to another embodiment of this application. In this embodiment, an intelligent robot monitors the behavioral information of the target user to trigger the human-computer interaction process, such as... Figure 6 As shown, the intelligent interaction method includes the following steps S601-S605:

[0113] S601, based on the target user's target language knowledge graph, determine the target knowledge information that matches the target user's language ability category.

[0114] Since the target language knowledge graph includes knowledge nodes corresponding to the target user's current language ability category, as well as the knowledge information corresponding to those knowledge nodes, the knowledge information corresponding to the knowledge nodes in the target language knowledge graph can be identified as target knowledge information that matches the target user's language ability category. Target knowledge information may include phonetics (such as the phonetics of characters, words, and sentences), semantics, grammar, pragmatic skills (such as single-word sentences and complex sentences), etc.

[0115] S602, monitor the behavioral information of the target user, which may include at least one of action information, language information, and emotional information.

[0116] In this embodiment, cameras can be pre-installed around the target user to monitor their behavior. Alternatively, the baby's language information can be monitored using a voice acquisition device built into the intelligent robot.

[0117] S603, when behavioral information is detected, determines the corresponding interaction strategy for the target user based on the target knowledge information and behavioral information that match the target user's language ability category, and interacts with the target user based on the interaction strategy.

[0118] S604, during the interaction, captures audio and video data of the target user's behavioral events through a pre-installed camera device; and collects brain signals from the target user.

[0119] The audio and video data includes audio data and / or image data. Cameras can be pre-installed around the target user to record their behavior during human-computer interaction. Optionally, the camera is activated and records the target user's behavior after the human-computer interaction is triggered. Brain signals may include electroencephalogram (EEG) data, brightness of language regions of the brain, etc.

[0120] The S605 analyzes the captured audio and video data and brain signals to obtain the interaction characteristics information of the target user during the interaction process, and stores the interaction characteristics information of the target user during the interaction process locally.

[0121] The analysis methods for audio / video data and brain signals are similar to those in the above embodiments and will not be repeated here. The locally stored interaction feature information can be used for subsequent updates to the target language knowledge graph. For example, if an update cycle is preset, when the update cycle arrives, the target user's current language development category is analyzed based on the locally stored interaction feature information, and then the target language knowledge graph is updated according to the target user's current language development category.

[0122] exist Figure 6 In the illustrated embodiment, the target user can be any user who needs to use the intelligent robot, such as an infant who needs companionship or an elderly person with limited mobility.

[0123] As can be seen from the above embodiments, whether the human-computer interaction process is actively triggered by an electronic device (such as an intelligent robot) or by an electronic device monitoring the target user's behavioral information and triggering the human-computer interaction process upon detecting behavioral information, a personalized and customized interaction strategy can be determined for the target user based on target knowledge information and the target user's behavioral information that match the target user's language ability category. Therefore, when interacting with the target user based on the interaction strategy, it can not only ensure that the interaction process matches the target user's language ability category, avoiding interaction obstacles (such as difficulty in understanding language or actions during the interaction process) caused by a mismatch between the interaction strategy and the target user's language ability category, but also achieve accurate guidance for the target user during the human-computer interaction process. In addition, it can comprehensively consider the target user's action information, language information, and / or emotional information to formulate the interaction strategy, improving the intelligence and personalization of human-computer interaction.

[0124] Figure 7 This is a schematic flowchart illustrating a language ability classification model training method according to an embodiment of this application, such as... Figure 7 As shown, the method includes:

[0125] S702, Obtain the sample language ability characteristics and sample language ability categories of sample users; the sample language ability characteristics include object response characteristics and / or voice response characteristics.

[0126] The sample language ability categories are used to characterize the current language abilities of the sample users. Optionally, according to the developmental stages of language ability, it can be divided into the following categories: language preparation stage, language completion stage, language intelligence stage, dialogue stage, writing stage, echo stage, addressing stage, speaking center, writing center, visual language center, language center, etc. Furthermore, the developmental stages of language ability can be further subdivided. For example, the language preparation stage includes the single-word sentence stage, the two-word sentence stage, and the simple sentence stage; the language completion stage includes the compound sentence stage and the interrogative sentence stage; and so on.

[0127] Optionally, sample language ability features of sample users can be collected in advance. For each sample user, when collecting the sample user's object-related reaction features, an object-related action matching their prior knowledge information can be performed on the sample user, such as moving the food "apple" from one location to the sample user's eyes, and then collecting the sample user's reaction information to the action. When collecting the sample user's vocal reaction features, vocal information matching their prior knowledge information can be emitted to the sample user, such as emitting the sound "apple" to the sample user if the sample user should currently be learning single-word sentences, and then collecting the sample user's reaction information to the sound.

[0128] The characteristics of responses to objects can be characterized by sensitivity to those objects, and the characteristics of responses to sounds can be characterized by sensitivity to those sounds. When collecting information on the responses of sample users to actions or speech, existing brain signal acquisition technologies can be used to collect brain signals, including at least one of the following: electroencephalogram (EEG) data and the brightness of the language areas of the brain. By analyzing these brain signals, the sample user's response information, such as response sensitivity, can be determined. EEG data represents the changes in electrical waves during brain activity; therefore, the greater and faster the amplitude of these changes, the higher the sample user's response sensitivity, and vice versa. The brightness of the language areas of the brain reflects the activity level of those areas; therefore, the higher the brightness of the language areas, the higher the sample user's response sensitivity, and vice versa.

[0129] Optionally, if the sample user's language ability category is known, the sample user's language ability characteristics can also be estimated based on the sample user's language ability category. For example, if the sample user is an adult user with mature language development, then the sample user's response sensitivity can be considered to be high. If the response sensitivity is represented by a value of 0 to 1, then the response sensitivity amplitude of the sample user can be 0.9 (or other higher values).

[0130] S704. Input the language ability features of the samples into the language ability classification model to be trained, classify the language ability of the sample users, and obtain the classification results.

[0131] S706, adjust the model parameters of the language ability classification model to be trained based on the classification results and the language ability category of the samples.

[0132] Among them, the sample language ability category serves as the label data, i.e., the standard output data, for the language ability classification model to be trained. By comparing the classification result with the sample language ability category, the difference between the classification result and the sample language ability category can be determined. This difference can be used to determine whether to continue adjusting the model parameters.

[0133] The technical solution of this application involves acquiring sample language ability features and categories of sample users, inputting these features into a language ability classification model to be trained, classifying the language abilities of the sample users, obtaining classification results, and then adjusting the model parameters of the language ability classification model to be trained based on the classification results and sample language ability categories, thereby obtaining a trained language ability classification model. Since the sample language ability features of the sample users include object response features and / or vocal response features, the language ability classification model trained based on these features not only has the ability to analyze the user's object response features and / or vocal response features, but also the ability to classify the user's language abilities based on these features. Furthermore, during intelligent human-computer interaction, the language ability classification model can accurately analyze the target user's language ability category, providing strong data support for the subsequent intelligent human-computer interaction process.

[0134] In one embodiment, the sample language ability features include sample language ability feature values. When there are multiple sample language ability features, each sample language ability feature corresponds to its own weight. Optionally, the sample language ability features include object response features and sound response features. Object response features are characterized by sensitivity to object responses, and sound response features are characterized by sensitivity to sound responses. Therefore, sensitivity to object responses can be used as object response feature values, and sensitivity to sound responses can be used as sound response feature values.

[0135] Optionally, to facilitate model calculation, the sample language ability feature values ​​can be normalized to obtain normalized sample language ability feature values. The normalized sample language ability feature values ​​are distributed between 0 and 1. These normalized sample language ability feature values ​​are then input into the language ability classification model.

[0136] like Figure 8 As shown, the language ability classification model to be trained includes a parameter adjustment layer, a score calculation layer, a classification layer, and a fully connected layer. When inputting sample language ability features into the language ability classification model to classify the language ability of sample users, specifically, the sample language ability feature values ​​are input into the language ability classification model to be trained, and then the following steps C1-C3 are performed using the language ability classification model to be trained:

[0137] Step C1: Determine the weight of each language ability feature in the current iteration through the parameter adjustment layer.

[0138] Step C2 involves calculating the sample language ability score of the sample user through the score calculation layer, based on the sample language ability feature value corresponding to each language ability feature and the weight corresponding to each language ability feature.

[0139] In this step, the sample language ability feature values ​​are weighted and summed according to the weights corresponding to each language ability feature, thus obtaining the sample user's sample language ability score. For example, the weight corresponding to the object response feature value is 0.6, and the weight corresponding to the voice response feature value is 0.4. If the sample user's object response feature value is determined to be 0.9 and the voice response feature value is 0.1, then the sample user's sample language ability score is: 0.9*0.6 + 0.1*0.4 = 0.58.

[0140] Step C3 involves classifying the language abilities of sample users through a classification layer, based on the sample language ability scores and the pre-defined mapping relationship between language ability scores and language ability categories.

[0141] After obtaining the language ability classification results of the sample users, the classification results are compared with the sample language ability categories through the fully connected layer of the language ability classification model to be trained, to determine whether the model parameters need to be adjusted further, i.e., whether the preset iteration termination condition is met. The iteration termination condition may include at least one of the following: the accuracy of the classification result is greater than or equal to a preset accuracy threshold, the probability value of the sample user belonging to the sample language ability category is greater than or equal to a preset probability threshold, or the number of iterations reaches a preset number threshold.

[0142] Optionally, the classification result may include the probability value of the sample user belonging to each language ability category. By comparing the classification result with the sample language ability category, the probability value of the sample user belonging to the sample language ability category after this iteration can be determined. If the probability value is greater than or equal to a preset probability threshold, the iteration terminates, and the trained language ability classification model is obtained; if the probability value is less than the preset probability threshold, the fully connected layer feeds the classification result back to the parameter adjustment layer, so that the parameter adjustment layer updates the model parameters based on the classification result of this iteration, and then performs the next iteration based on the updated model parameters.

[0143] Optionally, the classification result may include the first language ability category of the sample user. By comparing the first language ability category with the sample language ability category, it can be determined whether the sample user's ability classification is correct after this iteration. When there are multiple sample users, each sample user has its own classification result. Based on this, the accuracy of the classification results for multiple sample users can be calculated. If the accuracy is greater than or equal to a preset accuracy threshold, the iteration terminates, and the trained language ability classification model is obtained. If the accuracy is less than the preset accuracy threshold, the fully connected layer forwards the classification result to the parameter adjustment layer, allowing the parameter adjustment layer to update the model parameters based on the classification result of this iteration. Then, the next iteration is performed based on the updated model parameters.

[0144] In summary, specific embodiments of this subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing can be advantageous.

[0145] The above describes the intelligent interaction method and language ability classification model training method provided in the embodiments of this application. Based on the same idea, the embodiments of this application also provide an intelligent interaction device and a language ability classification model training device.

[0146] Figure 9 A schematic block diagram of an intelligent interactive device according to an embodiment of this application, such as... Figure 9 The device includes:

[0147] The first acquisition module 91 is used to acquire a target language knowledge graph corresponding to the target user; the target language knowledge graph is created based on the target user's language ability category, and the target user's language ability category is determined based on the target user's language ability characteristics;

[0148] The first determining module 92 is used to determine target knowledge information that matches the language ability category of the target user based on the target language knowledge graph;

[0149] The second acquisition module 93 is used to acquire the behavioral information of the target user; the behavioral information includes at least one of the following: action information, language information, and emotional information;

[0150] The second determining module 94 is used to determine an interaction strategy corresponding to the target user based on the target knowledge information and the behavior information, and to interact with the target user based on the interaction strategy.

[0151] In one embodiment, the first acquisition module 91 includes:

[0152] The first acquisition unit is used to acquire a pre-created initial language knowledge graph; the initial language knowledge graph includes: multiple knowledge nodes, and user information and knowledge information corresponding to each knowledge node;

[0153] The first determining unit is used to determine knowledge information that matches the user information of the target user based on the initial language knowledge graph and the user information of the target user;

[0154] The acquisition unit is used to acquire the language ability characteristics of the target user based on knowledge information that matches the user information of the target user; the language ability characteristics include object response characteristics and / or vocal response characteristics.

[0155] The second determining unit is used to determine the language ability category of the target user based on the language ability characteristics;

[0156] The generation unit is configured to generate the target language knowledge graph based on the initial language knowledge graph and the language ability category of the target user.

[0157] In one embodiment, each knowledge node corresponds to a language ability category;

[0158] The generation unit is used for:

[0159] The target user's language ability category is matched with each knowledge node, and the first knowledge node that matches the target user's language ability category is determined based on the matching results.

[0160] The target language knowledge graph is generated based on the first knowledge node and the second knowledge node; the target language knowledge graph includes: the first knowledge node, the second knowledge node, the knowledge information corresponding to the first knowledge node, and the knowledge information corresponding to the second knowledge node; the second knowledge node is at least one knowledge node adjacent to the first knowledge node.

[0161] In one embodiment, the second determining unit is used to:

[0162] The language ability features are input into a pre-trained language ability classification model to obtain the language ability category of the target user; wherein, the language ability classification model is trained based on the sample language ability features and sample language ability categories of the sample users.

[0163] In one embodiment, the language ability feature includes language ability feature values; the language ability classification model includes a score calculation layer and a classification layer;

[0164] The second determining unit is used for:

[0165] The language proficiency score of the target user is calculated through the score calculation layer based on the language proficiency feature value corresponding to each language proficiency feature and the weight corresponding to each language proficiency feature.

[0166] The classification layer determines the target user's language ability category based on the target user's language ability score and the preset mapping relationship between language ability scores and language ability categories.

[0167] In one embodiment, the second acquisition module 93 includes:

[0168] An execution unit is configured to trigger a target user action event based on the target knowledge information.

[0169] The second acquisition unit is used to acquire the behavior information corresponding to the behavior event.

[0170] In one embodiment, the second acquisition unit is further configured to:

[0171] Acquire multimedia data of the target user's execution of the behavioral event;

[0172] The multimedia data is processed to obtain audio data and / or image data of the target user performing the behavioral event;

[0173] Analyze the audio data and / or the image data to obtain the behavioral information corresponding to the behavioral event.

[0174] In one embodiment, the apparatus further includes:

[0175] The second acquisition module is used to acquire the interaction feature information of the target user during the interaction process; the interaction feature information includes object response features and / or sound response features;

[0176] The judgment module is used to determine whether the interaction feature information matches the language ability features of the target user;

[0177] The third determining module is used to, if not, re-determine the target user's language ability characteristics based on the interaction feature information to obtain updated language ability characteristics;

[0178] The fourth determining module is used to redetermine the target user's language ability category based on the updated language ability characteristics, thereby obtaining the language ability category to be updated;

[0179] The update module is used to update the target language knowledge graph according to the language ability category to be updated.

[0180] The apparatus employing embodiments of this application acquires a target language knowledge graph corresponding to a target user, determines target knowledge information matching the target user's language ability category based on the target language knowledge graph, acquires the target user's behavioral information, and then determines an interaction strategy corresponding to the target user based on the target knowledge information and the target user's behavioral information. Since the interaction strategy is jointly determined based on the target knowledge information matching the target user's language ability category and the target user's behavioral information, and the target knowledge information is determined based on the target language knowledge graph created based on the target user's language ability category, when interacting with the target user based on the interaction strategy, it not only ensures that the interaction process matches the target user's language ability category, avoiding interaction obstacles (such as difficulty understanding language or actions during the interaction process) caused by a mismatch between the interaction strategy and the target user's language ability category, thus achieving accurate guidance for the target user during human-computer interaction, but also comprehensively considers the target user's action information, language information, and / or emotional information to formulate the interaction strategy, improving the intelligence and personalization of human-computer interaction. Furthermore, since there is a correspondence between the target language knowledge graph and the target user, that is, each user has its own corresponding language knowledge graph, the device can not only achieve intelligent interaction with the user, but also realize customized language knowledge graphs for different users, thereby achieving personalized and customized human-computer interaction effects.

[0181] Those skilled in the art will understand that Figure 9 The intelligent interactive device in the text can be used to implement the intelligent interactive method described above. The details described therein should be similar to those in the method section above. To avoid being too complicated, they will not be repeated here.

[0182] Figure 10 A schematic block diagram of a language ability classification model training device according to an embodiment of this application, as shown below. Figure 10 As shown, the device includes:

[0183] The third acquisition module 101 is used to acquire the sample language ability characteristics and sample language ability categories of sample users; the sample language ability characteristics include object response characteristics and / or sound response characteristics.

[0184] The classification module 102 is used to input the language ability features of the sample into the language ability classification model to be trained, classify the language ability of the sample user, and obtain the classification result.

[0185] The parameter adjustment module 103 is used to adjust the model parameters of the language ability classification model to be trained according to the classification results and the sample language ability category.

[0186] In one embodiment, the sample language ability features include sample language ability feature values; the language ability classification model to be trained includes a parameter adjustment layer, a score calculation layer, and a classification layer;

[0187] The parameter adjustment layer is used to determine the weight of each language ability feature in the current iteration.

[0188] The score calculation layer is used to calculate the sample language ability score of the sample user based on the sample language ability feature value corresponding to each language ability feature and the weight corresponding to each language ability feature.

[0189] The classification layer is used to classify the language ability of the sample users based on the sample language ability score and the preset mapping relationship between language ability scores and language ability categories.

[0190] The apparatus of this application acquires sample language ability features and categories of sample users, inputs these features into a language ability classification model to be trained, classifies the language abilities of the sample users, obtains classification results, and then adjusts the model parameters of the language ability classification model to be trained based on the classification results and the sample language ability categories, thereby obtaining a trained language ability classification model. Since the sample language ability features of the sample users include object response features and / or vocal response features, the language ability classification model trained based on these features not only has the ability to analyze the user's object response features and / or vocal response features, but also has the ability to classify the user's language abilities based on these features. Furthermore, during intelligent human-computer interaction, the language ability classification model can accurately analyze the target user's language ability category, providing strong data support for the subsequent intelligent human-computer interaction process.

[0191] Those skilled in the art will understand that Figure 10 The language ability classification model training device in the document can be used to implement the language ability classification model training method described above. The details of the method should be similar to those described in the previous section. To avoid being too complicated, they will not be repeated here.

[0192] Following the same line of thought, embodiments of this application also provide an electronic device, such as... Figure 11As shown. Electronic devices can vary considerably due to differences in configuration or performance, and may include one or more processors 1101 and memory 1102. Memory 1102 may store one or more application programs or data. Memory 1102 may be temporary or persistent storage. The application programs stored in memory 1102 may include one or more modules (not shown), each module may include a series of computer-executable instructions for the electronic device. Furthermore, processor 1101 may be configured to communicate with memory 1102 and execute the series of computer-executable instructions in memory 1102 on the electronic device. The electronic device may also include one or more power supplies 1103, one or more wired or wireless network interfaces 1104, one or more input / output interfaces 1105, and one or more keyboards 1106.

[0193] Specifically, in this embodiment, the electronic device includes a memory and one or more programs, wherein one or more programs are stored in the memory, and one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for use in the electronic device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following:

[0194] Obtain a target language knowledge graph corresponding to the target user; the target language knowledge graph is created based on the target user's language ability category, and the target user's language ability category is determined based on the target user's language ability characteristics;

[0195] Based on the target language knowledge graph, target knowledge information matching the target user's language ability category is determined;

[0196] Obtain the behavioral information of the target user; the behavioral information includes at least one of the following: action information, language information, and emotional information;

[0197] Based on the target knowledge information and the behavioral information, an interaction strategy corresponding to the target user is determined, and interaction with the target user is carried out based on the interaction strategy.

[0198] In another embodiment, the electronic device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for use in the electronic device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following:

[0199] Obtain the sample language ability characteristics and sample language ability categories of sample users; the sample language ability characteristics include object response characteristics and / or sound response characteristics;

[0200] The language ability features of the sample users are input into the language ability classification model to be trained, and the language ability of the sample users is classified to obtain the classification result.

[0201] Based on the classification results and the language ability categories of the samples, the model parameters of the language ability classification model to be trained are adjusted.

[0202] This application also proposes a storage medium that stores one or more computer programs, each including instructions that, when executed by an electronic device comprising multiple applications, enable the electronic device to perform various processes of the above-described intelligent interaction method embodiments, specifically for executing:

[0203] Obtain a target language knowledge graph corresponding to the target user; the target language knowledge graph is created based on the target user's language ability category, and the target user's language ability category is determined based on the target user's language ability characteristics;

[0204] Obtain behavioral information of the target user; the behavioral information includes at least one of the following: action information, language information, and emotional information;

[0205] Based on the target knowledge information and the behavioral information, an interaction strategy corresponding to the target user is determined, and interaction with the target user is carried out based on the interaction strategy.

[0206] This application also proposes a storage medium that stores one or more computer programs, the computer programs including instructions that, when executed by an electronic device including multiple applications, enable the electronic device to perform various processes of the above-described language ability classification model training method embodiments, specifically for executing:

[0207] Obtain the sample language ability characteristics and sample language ability categories of sample users; the sample language ability characteristics include object response characteristics and / or sound response characteristics;

[0208] The language ability features of the sample users are input into the language ability classification model to be trained, and the language ability of the sample users is classified to obtain the classification result.

[0209] Based on the classification results and the language ability categories of the samples, the model parameters of the language ability classification model to be trained are adjusted.

[0210] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0211] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0212] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0213] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0214] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0215] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0216] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0217] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0218] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0219] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0220] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0221] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0222] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. An intelligent interaction method, characterized in that, include: Obtain the target language knowledge graph corresponding to the target user; The target language knowledge graph is created based on the target user's language ability category, which is determined based on the target user's language ability characteristics. The target language knowledge graph includes multiple knowledge nodes and knowledge information associated with each knowledge node. The knowledge information associated with each knowledge node is used to characterize the knowledge that the target user should learn at each language development node. Based on the target language knowledge graph, target knowledge information matching the language ability category of the target user is determined. The target knowledge information is used to represent the knowledge that the target user should learn at the current stage of language development. Obtain the behavioral information of the target user; the behavioral information includes at least one of the following: action information, language information, and emotional information; Based on the target knowledge information and the behavioral information, an interaction strategy corresponding to the target user is determined, and interaction with the target user is carried out based on the interaction strategy.

2. The method according to claim 1, characterized in that, The acquisition of the target language knowledge graph corresponding to the target user includes: Obtain a pre-created initial language knowledge graph; the initial language knowledge graph includes: multiple knowledge nodes, and user information and knowledge information corresponding to each knowledge node; Based on the initial language knowledge graph and the user information of the target user, determine the knowledge information that matches the user information of the target user; Based on knowledge information that matches the user information of the target user, the language ability characteristics of the target user are collected; the language ability characteristics include object response characteristics and / or sound response characteristics. Based on the language ability characteristics, determine the language ability category of the target user; The target language knowledge graph is generated based on the initial language knowledge graph and the target user's language ability category.

3. The method according to claim 2, characterized in that, Each of the aforementioned knowledge nodes corresponds to a language ability category; The step of generating the target language knowledge graph based on the initial language knowledge graph and the target user's language ability category includes: The target user's language ability category is matched with each knowledge node, and the first knowledge node that matches the target user's language ability category is determined based on the matching results. The target language knowledge graph is generated based on the first knowledge node and the second knowledge node; the target language knowledge graph includes: the first knowledge node, the second knowledge node, the knowledge information corresponding to the first knowledge node, and the knowledge information corresponding to the second knowledge node; the second knowledge node is at least one knowledge node adjacent to the first knowledge node.

4. The method according to claim 2, characterized in that, Determining the target user's language ability category based on the language ability characteristics includes: The language ability features are input into a pre-trained language ability classification model to obtain the language ability category of the target user; wherein, the language ability classification model is trained based on the sample language ability features and sample language ability categories of the sample users.

5. The method according to claim 4, characterized in that, The language ability features include language ability feature values; the language ability classification model includes a score calculation layer and a classification layer; The step of inputting the language ability features into a pre-trained language ability classification model to obtain the language ability category of the target user includes: The language proficiency score of the target user is calculated through the score calculation layer based on the language proficiency feature value corresponding to each language proficiency feature and the weight corresponding to each language proficiency feature. The classification layer determines the target user's language ability category based on the target user's language ability score and the preset mapping relationship between language ability scores and language ability categories.

6. The method according to claim 1, characterized in that, The acquisition of the target user's behavioral information includes: Based on the target knowledge information, the target user is triggered to perform a behavioral event; Obtain the behavior information corresponding to the behavior event.

7. The method according to claim 6, characterized in that, The step of obtaining the behavioral information corresponding to the behavioral event includes: Acquire multimedia data of the target user's execution of the behavioral event; The multimedia data is processed to obtain audio data and / or image data of the target user performing the behavioral event; Analyze the audio data and / or the image data to obtain the behavioral information corresponding to the behavioral event.

8. The method according to claim 1, characterized in that, The method further includes: Acquire the interaction feature information of the target user during the interaction process; the interaction feature information includes object response features and / or sound response features; Determine whether the interaction feature information matches the language ability features of the target user; If not, the language proficiency characteristics of the target user are re-determined based on the interaction feature information to obtain the updated language proficiency characteristics; Based on the updated language ability characteristics, the language ability category of the target user is redefined to obtain the language ability category to be updated; The target language knowledge graph is updated based on the language ability category to be updated.

9. An intelligent interactive device, characterized in that, include: The first acquisition module is used to acquire the target language knowledge graph corresponding to the target user; The target language knowledge graph is created based on the target user's language ability category, which is determined based on the target user's language ability characteristics. The target language knowledge graph includes multiple knowledge nodes and knowledge information associated with each knowledge node. The knowledge information associated with each knowledge node is used to characterize the knowledge that the target user should learn at each language development node. The first determining module is used to determine target knowledge information that matches the language ability category of the target user based on the target language knowledge graph. The target knowledge information is used to represent the knowledge that the target user should learn at the current stage of language development. The second acquisition module is used to acquire the behavioral information of the target user; the behavioral information includes at least one of the following: action information, language information, and emotional information; The second determining module is used to determine an interaction strategy corresponding to the target user based on the target knowledge information and the behavioral information, and to interact with the target user based on the interaction strategy.

10. An electronic device, characterized in that, The system includes a processor and a memory electrically connected to the processor, the memory storing a computer program, and the processor being configured to call and execute the computer program from the memory to implement the intelligent interaction method as described in any one of claims 1-8.

11. A storage medium, characterized in that, The storage medium is used to store a computer program that can be executed by a processor to implement the intelligent interaction method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Interaction mode determination method and device, electronic equipment and storage medium

    CN112306238A

  • Cloud language ability evaluation system and wearable recording terminal

    CN112750465A

  • Automated assistants that accommodate multiple age groups and / or vocabulary levels

    IN202027047813A